Asana Status · History · Incident #6830

RESOLVED

Asana full outage for some users (2 hours, 25% outage)

Major · Started Aug 31, 2026 · 3:14 PM

  • Duration

    2h 4m

  • Severity

    Major

  • Detection lead

  • User reports

    Peak 3/hr

Summary

Asana full outage for some users (2 hours, 25% outage)

We’ve been working to add a caching layer to our update pipeline, tuning it carefully and rolling it out gradually. On Wednesday, August 26 we enabled it for most use cases, and it initially performed well. On Monday, August 31, a combination of unrelated infrastructure changes and peak traffic pushed the cache past its scaling limits. Once that threshold was crossed, the cache became unusable, and many pods serving read traffic for the Asana application could no longer serve it. We mitigated the incident by reverting the system to use the previous, non-cached code path. The revert was successful, but recovery took longer than we would expect for this class of issue. A fuller analysis is underway. We will follow up with root causes, action items, and improvements, including why recovery took as long as it did. We were fully down for about 25% of our users, for 2 hours, 15 minutes.


  • Started

    Aug 31, 2026 · 3:14 PM

  • Resolved

    Aug 31, 2026 · 5:18 PM

  • Duration

    2h 4m

  • Severity

    Major

Event timeline

How this incident unfolded

  • Investigating

    Aug 31 · 3:14 PM Asana

    We are investigating alerts for slow performance and application errors.

  • Investigating

    Aug 31 · 3:55 PM Asana

    We are continuing to investigate the issue; we have reverted recent changes, and are working to identify the source of the errors.

  • Investigating

    Aug 31 · 4:24 PM Asana

    We have made configuration changes and see partial recovery, but we continue to see some elevated errors.

  • Monitoring

    Aug 31 · 4:47 PM Asana

    We've applied a fix are seeing signs of recovery.

  • Resolved

    Aug 31 · 5:18 PM Asana

    User-facing symptoms have recovered. We'll continue to monitor, and will prioritize a retrospective to understand and prevent similar incidents in the future.

  • Postmortem

    Sep 2 · 9:29 PM Asana

    We’ve been working to add a caching layer to our update pipeline, tuning it carefully and rolling it out gradually. On Wednesday, August 26 we enabled it for most use cases, and it initially performed well. On Monday, August 31, a combination of unrelated infrastructure changes and peak traffic pushed the cache past its scaling limits. Once that threshold was crossed, the cache became unusable, and many pods serving read traffic for the Asana application could no longer serve it. We mitigated the incident by reverting the system to use the previous, non-cached code path. The revert was successful, but recovery took longer than we would expect for this class of issue. A fuller analysis is underway. We will follow up with root causes, action items, and improvements, including why recovery took as long as it did. We were fully down for about 25% of our users, for 2 hours, 15 minutes.

User reports · incident window

Report volume over time

Hourly report count. Red bars mark the incident window.

Peak: 3 reports/hr

← Before incident After →

Get alerted before the next Asana outage.

Pulsetic catches degradations minutes before vendors acknowledge them.