Jira Software Status · History · Incident #105960

RESOLVED

Disrupted JIRA availability in US East

Critical · Started Sep 3, 2026 · 3:57 PM

  • Duration

    1h 4m

  • Severity

    Critical

  • Detection lead

  • User reports

Summary

Disrupted JIRA availability in US East

### Summary On September 3, 2026, between 14:52 and 17:02 UTC, Atlassian customers using Jira and Jira Service Management in the US-East region experienced intermittent errors and degraded access to core Jira experiences, including Rovo, viewing issues, boards, navigation, search-backed experiences, workflow-related actions. The event was triggered by a progressive rollout of a caching change in Atlassian's authorization service. The change increased cache traffic and caused saturation in the cache layer used for authorization checks, which then resulted in timeouts and service errors for some customers in the region. The incident was detected by our monitoring within approximately 12 minutes. We mitigated by deploying the affected cache service and disabling the feature flags associated with the new cache path. The total time to resolution was 2 hours and 10 minutes. ### IMPACT The overall impact occurred on September 3, 2026 between 14:52 and 17:02 UTC. During this period, some customers in the US-East region using Jira and Jira Service Management experienced intermittent failures or degraded performance when accessing experiences that depend on authorization checks. These included issue view, Rovo, board view, navigation, search-backed experiences, API, and workflow-related actions. ### ROOT CAUSE The issue was caused by a progressive rollout of a new cache path in Atlassian’s authorization service. The change was intended to improve scalability for customers with large authorization datasets by changing how some authorization information was cached. During the rollout, both the existing and new cache paths were active for consistency checks, which increased the number of cache entries and cache operations. In the US-East region, where traffic volume was highest, this additional load saturated parts of the shared cache layer. As cache requests timed out, the authorization service made more fallback requests to upstream systems, further increasing load and causing circuit breakers and timeouts. Dependent Jira experiences that require authorization checks then returned intermittent errors or failed to load for affected users. The initial rollout monitoring did not surface the risk early enough because lower-traffic regions did not show the same level of degradation and some alerts were not scoped or routed in a way that highlighted the regional cache saturation before customer impact grew. ### REMEDIAL ACTIONS PLAN & NEXT STEPS We know that outages impact customer productivity, and we apologize to customers who were affected during this incident. We are prioritizing the following improvement actions to help reduce the likelihood and impact of similar incidents: - **Enhance deployment safeguards:** Improve rollout practices for high-risk cache and authorization-service changes by using smaller, region-aware rollout increments and validating behavior in higher-traffic regions before broader enablement. - **Implement graceful degradation:** Evaluate graceful-degradation options for authorization checks so dependent experiences can continue working, if a cache dependency is unavailable. - **Strengthen regional telemetry and capacity monitoring:** Strengthen cache capacity and health monitoring, including regional detectors for cache CPU, memory pressure, hit rate, eviction rate, timeout rate, and connection behavior. - **Expand production-scale load testing:** Review cache design and capacity planning for authorization data so changes that increase key cardinality or cache operations are load-tested against representative production-scale traffic before rollout. - **Harden client resilience:** Review cache client timeout and reconnect behavior to reduce the risk of connection storms when cache nodes become slow or saturated. We will continue to validate these actions through the incident review process and track them to completion with the responsible engineering teams. Thanks, Atlassian Customer Support


  • Started

    Sep 3, 2026 · 3:57 PM

  • Resolved

    Sep 3, 2026 · 5:02 PM

  • Duration

    1h 4m

  • Severity

    Critical

Event timeline

How this incident unfolded

  • Investigating

    Sep 3 · 3:57 PM Jira Software

    We are actively investigating reports of a service disruption affecting JIRA customers in US East. We will share updates here as more information is available.

  • Monitoring

    Sep 3 · 4:35 PM Jira Software

    Our teams have taken mitigation steps to reduce the impact and recovery among affected customers began at 16:15 UTC. We continue to investigate the root cause of the issue, and will share additional updates here as more information is available.

  • Resolved

    Sep 3 · 5:02 PM Jira Software

    On September 3, 2026, JIRA experienced a disruption, and services were unavailable to affected users. The issue has now been resolved, and the service is operating normally for all affected customers.

  • Postmortem

    Sep 10 · 6:15 PM Jira Software

    ### Summary On September 3, 2026, between 14:52 and 17:02 UTC, Atlassian customers using Jira and Jira Service Management in the US-East region experienced intermittent errors and degraded access to core Jira experiences, including Rovo, viewing issues, boards, navigation, search-backed experiences, workflow-related actions. The event was triggered by a progressive rollout of a caching change in Atlassian's authorization service. The change increased cache traffic and caused saturation in the cache layer used for authorization checks, which then resulted in timeouts and service errors for some customers in the region. The incident was detected by our monitoring within approximately 12 minutes. We mitigated by deploying the affected cache service and disabling the feature flags associated with the new cache path. The total time to resolution was 2 hours and 10 minutes. ### IMPACT The overall impact occurred on September 3, 2026 between 14:52 and 17:02 UTC. During this period, some customers in the US-East region using Jira and Jira Service Management experienced intermittent failures or degraded performance when accessing experiences that depend on authorization checks. These included issue view, Rovo, board view, navigation, search-backed experiences, API, and workflow-related actions. ### ROOT CAUSE The issue was caused by a progressive rollout of a new cache path in Atlassian’s authorization service. The change was intended to improve scalability for customers with large authorization datasets by changing how some authorization information was cached. During the rollout, both the existing and new cache paths were active for consistency checks, which increased the number of cache entries and cache operations. In the US-East region, where traffic volume was highest, this additional load saturated parts of the shared cache layer. As cache requests timed out, the authorization service made more fallback requests to upstream systems, further increasing load and causing circuit breakers and timeouts. Dependent Jira experiences that require authorization checks then returned intermittent errors or failed to load for affected users. The initial rollout monitoring did not surface the risk early enough because lower-traffic regions did not show the same level of degradation and some alerts were not scoped or routed in a way that highlighted the regional cache saturation before customer impact grew. ### REMEDIAL ACTIONS PLAN & NEXT STEPS We know that outages impact customer productivity, and we apologize to customers who were affected during this incident. We are prioritizing the following improvement actions to help reduce the likelihood and impact of similar incidents: - **Enhance deployment safeguards:** Improve rollout practices for high-risk cache and authorization-service changes by using smaller, region-aware rollout increments and validating behavior in higher-traffic regions before broader enablement. - **Implement graceful degradation:** Evaluate graceful-degradation options for authorization checks so dependent experiences can continue working, if a cache dependency is unavailable. - **Strengthen regional telemetry and capacity monitoring:** Strengthen cache capacity and health monitoring, including regional detectors for cache CPU, memory pressure, hit rate, eviction rate, timeout rate, and connection behavior. - **Expand production-scale load testing:** Review cache design and capacity planning for authorization data so changes that increase key cardinality or cache operations are load-tested against representative production-scale traffic before rollout. - **Harden client resilience:** Review cache client timeout and reconnect behavior to reduce the risk of connection storms when cache nodes become slow or saturated. We will continue to validate these actions through the incident review process and track them to completion with the responsible engineering teams. Thanks, Atlassian Customer Support

Get alerted before the next Jira Software outage.

Pulsetic catches degradations minutes before vendors acknowledge them.