Asana Status · History · Incident #6913
RESOLVEDWe are fully down for a subset of customers. We're investigating the issue, and are rolling back a deployment.
Major · Started Sep 2, 2026 · 6:08 PM
Asana Status · History · Incident #6913
RESOLVEDMajor · Started Sep 2, 2026 · 6:08 PM
Duration
41m
Severity
Major
Detection lead
—
User reports
—
Summary
Incident: An internal system responsible for automatically scaling backend server capacity in one of our compute clusters stopped replacing capacity that had been cycled out during routine maintenance. This caused a gradual reduction in available capacity over several hours. A subsequent deployment was activated with insufficient capacity, causing all new requests for that compute cluster to fail. To mitigate the impact, we reverted to a previous release revision, which still had sufficient capacity. Impact: For approximately 30 minutes, customers whose traffic was handled by this compute cluster saw full downtime; other customers were unaffected. No customer data was lost. Moving forward: We have added additional monitoring to detect this type of capacity-scaling failure much earlier, and have added safeguards to prevent deployments from shifting traffic before sufficient healthy capacity is confirmed. _Our metric considers a weighted average of uptime experienced by users at each data center. The number of minutes of downtime shown reflects this weighted average._
Started
Sep 2, 2026 · 6:08 PM
Resolved
Sep 2, 2026 · 6:49 PM
Duration
41m
Severity
Major
Event timeline
Investigating
Sep 2 · 6:08 PM AsanaWe are currently investigating this issue.
Investigating
Sep 2 · 6:13 PM AsanaWe are continuing to investigate this issue.
Investigating
Sep 2 · 6:14 PM AsanaThe rollback seems to have completely resolved errors. We're monitoring closely, but users should be able to access Asana again.
Monitoring
Sep 2 · 6:14 PM AsanaA fix has been implemented and we are monitoring the results.
Resolved
Sep 2 · 6:49 PM AsanaThis incident has been resolved.
Postmortem
Sep 4 · 7:52 PM AsanaIncident: An internal system responsible for automatically scaling backend server capacity in one of our compute clusters stopped replacing capacity that had been cycled out during routine maintenance. This caused a gradual reduction in available capacity over several hours. A subsequent deployment was activated with insufficient capacity, causing all new requests for that compute cluster to fail. To mitigate the impact, we reverted to a previous release revision, which still had sufficient capacity. Impact: For approximately 30 minutes, customers whose traffic was handled by this compute cluster saw full downtime; other customers were unaffected. No customer data was lost. Moving forward: We have added additional monitoring to detect this type of capacity-scaling failure much earlier, and have added safeguards to prevent deployments from shifting traffic before sufficient healthy capacity is confirmed. _Our metric considers a weighted average of uptime experienced by users at each data center. The number of minutes of downtime shown reflects this weighted average._
Pattern
Pulsetic catches degradations minutes before vendors acknowledge them.
Stay online, all the time, with Pulsetic's uptime prime.
By Designmodo
Designmodo Inc. 169 Madison Ave, #79627, New York, NY 10016, United States
Copyright © 2010-2026. Pulsetic® is a registered trademark.