Jira Software Status · History · Incident #105960
RESOLVEDDisrupted JIRA availability in US East
Critical · Started Sep 3, 2026 · 3:57 PM
Jira Software Status · History · Incident #105960
RESOLVEDCritical · Started Sep 3, 2026 · 3:57 PM
Duration
1h 4m
Severity
Critical
Detection lead
—
User reports
—
Summary
### Summary On September 3, 2026, between 14:52 and 17:02 UTC, Atlassian customers using Jira and Jira Service Management in the US-East region experienced intermittent errors and degraded access to core Jira experiences, including Rovo, viewing issues, boards, navigation, search-backed experiences, workflow-related actions. The event was triggered by a progressive rollout of a caching change in Atlassian's authorization service. The change increased cache traffic and caused saturation in the cache layer used for authorization checks, which then resulted in timeouts and service errors for some customers in the region. The incident was detected by our monitoring within approximately 12 minutes. We mitigated by deploying the affected cache service and disabling the feature flags associated with the new cache path. The total time to resolution was 2 hours and 10 minutes. ### IMPACT The overall impact occurred on September 3, 2026 between 14:52 and 17:02 UTC. During this period, some customers in the US-East region using Jira and Jira Service Management experienced intermittent failures or degraded performance when accessing experiences that depend on authorization checks. These included issue view, Rovo, board view, navigation, search-backed experiences, API, and workflow-related actions. ### ROOT CAUSE The issue was caused by a progressive rollout of a new cache path in Atlassian’s authorization service. The change was intended to improve scalability for customers with large authorization datasets by changing how some authorization information was cached. During the rollout, both the existing and new cache paths were active for consistency checks, which increased the number of cache entries and cache operations. In the US-East region, where traffic volume was highest, this additional load saturated parts of the shared cache layer. As cache requests timed out, the authorization service made more fallback requests to upstream systems, further increasing load and causing circuit breakers and timeouts. Dependent Jira experiences that require authorization checks then returned intermittent errors or failed to load for affected users. The initial rollout monitoring did not surface the risk early enough because lower-traffic regions did not show the same level of degradation and some alerts were not scoped or routed in a way that highlighted the regional cache saturation before customer impact grew. ### REMEDIAL ACTIONS PLAN & NEXT STEPS We know that outages impact customer productivity, and we apologize to customers who were affected during this incident. We are prioritizing the following improvement actions to help reduce the likelihood and impact of similar incidents: - **Enhance deployment safeguards:** Improve rollout practices for high-risk cache and authorization-service changes by using smaller, region-aware rollout increments and validating behavior in higher-traffic regions before broader enablement. - **Implement graceful degradation:** Evaluate graceful-degradation options for authorization checks so dependent experiences can continue working, if a cache dependency is unavailable. - **Strengthen regional telemetry and capacity monitoring:** Strengthen cache capacity and health monitoring, including regional detectors for cache CPU, memory pressure, hit rate, eviction rate, timeout rate, and connection behavior. - **Expand production-scale load testing:** Review cache design and capacity planning for authorization data so changes that increase key cardinality or cache operations are load-tested against representative production-scale traffic before rollout. - **Harden client resilience:** Review cache client timeout and reconnect behavior to reduce the risk of connection storms when cache nodes become slow or saturated. We will continue to validate these actions through the incident review process and track them to completion with the responsible engineering teams. Thanks, Atlassian Customer Support
Started
Sep 3, 2026 · 3:57 PM
Resolved
Sep 3, 2026 · 5:02 PM
Duration
1h 4m
Severity
Critical
Event timeline
Investigating
Sep 3 · 3:57 PM Jira SoftwareWe are actively investigating reports of a service disruption affecting JIRA customers in US East. We will share updates here as more information is available.
Monitoring
Sep 3 · 4:35 PM Jira SoftwareOur teams have taken mitigation steps to reduce the impact and recovery among affected customers began at 16:15 UTC. We continue to investigate the root cause of the issue, and will share additional updates here as more information is available.
Resolved
Sep 3 · 5:02 PM Jira SoftwareOn September 3, 2026, JIRA experienced a disruption, and services were unavailable to affected users. The issue has now been resolved, and the service is operating normally for all affected customers.
Postmortem
Sep 10 · 6:15 PM Jira Software### Summary On September 3, 2026, between 14:52 and 17:02 UTC, Atlassian customers using Jira and Jira Service Management in the US-East region experienced intermittent errors and degraded access to core Jira experiences, including Rovo, viewing issues, boards, navigation, search-backed experiences, workflow-related actions. The event was triggered by a progressive rollout of a caching change in Atlassian's authorization service. The change increased cache traffic and caused saturation in the cache layer used for authorization checks, which then resulted in timeouts and service errors for some customers in the region. The incident was detected by our monitoring within approximately 12 minutes. We mitigated by deploying the affected cache service and disabling the feature flags associated with the new cache path. The total time to resolution was 2 hours and 10 minutes. ### IMPACT The overall impact occurred on September 3, 2026 between 14:52 and 17:02 UTC. During this period, some customers in the US-East region using Jira and Jira Service Management experienced intermittent failures or degraded performance when accessing experiences that depend on authorization checks. These included issue view, Rovo, board view, navigation, search-backed experiences, API, and workflow-related actions. ### ROOT CAUSE The issue was caused by a progressive rollout of a new cache path in Atlassian’s authorization service. The change was intended to improve scalability for customers with large authorization datasets by changing how some authorization information was cached. During the rollout, both the existing and new cache paths were active for consistency checks, which increased the number of cache entries and cache operations. In the US-East region, where traffic volume was highest, this additional load saturated parts of the shared cache layer. As cache requests timed out, the authorization service made more fallback requests to upstream systems, further increasing load and causing circuit breakers and timeouts. Dependent Jira experiences that require authorization checks then returned intermittent errors or failed to load for affected users. The initial rollout monitoring did not surface the risk early enough because lower-traffic regions did not show the same level of degradation and some alerts were not scoped or routed in a way that highlighted the regional cache saturation before customer impact grew. ### REMEDIAL ACTIONS PLAN & NEXT STEPS We know that outages impact customer productivity, and we apologize to customers who were affected during this incident. We are prioritizing the following improvement actions to help reduce the likelihood and impact of similar incidents: - **Enhance deployment safeguards:** Improve rollout practices for high-risk cache and authorization-service changes by using smaller, region-aware rollout increments and validating behavior in higher-traffic regions before broader enablement. - **Implement graceful degradation:** Evaluate graceful-degradation options for authorization checks so dependent experiences can continue working, if a cache dependency is unavailable. - **Strengthen regional telemetry and capacity monitoring:** Strengthen cache capacity and health monitoring, including regional detectors for cache CPU, memory pressure, hit rate, eviction rate, timeout rate, and connection behavior. - **Expand production-scale load testing:** Review cache design and capacity planning for authorization data so changes that increase key cardinality or cache operations are load-tested against representative production-scale traffic before rollout. - **Harden client resilience:** Review cache client timeout and reconnect behavior to reduce the risk of connection storms when cache nodes become slow or saturated. We will continue to validate these actions through the incident review process and track them to completion with the responsible engineering teams. Thanks, Atlassian Customer Support
Pulsetic catches degradations minutes before vendors acknowledge them.
Stay online, all the time, with Pulsetic's uptime prime.
By Designmodo
Designmodo Inc. 169 Madison Ave, #79627, New York, NY 10016, United States
Copyright © 2010-2026. Pulsetic® is a registered trademark.