AssemblyAI Status · History · Incident #137922
RESOLVEDOutage on US Async Endpoint
Critical · Started Sep 16, 2026 · 8:22 PM
AssemblyAI Status · History · Incident #137922
RESOLVEDCritical · Started Sep 16, 2026 · 8:22 PM
Duration
1h 14m
Severity
Critical
Detection lead
—
User reports
—
Summary
**Summary** On September 16, 2026, between 20:04 and 21:08 UTC, our US Async API returned elevated server errors. At peak, roughly 50% of async transcription requests on the US endpoint failed, and a portion of the requests that did succeed took longer than usual to complete. Our EU endpoint and real-time/streaming services were not affected. **Root cause** The failure originated in AWS SQS, the managed queuing service our transcription pipeline uses to distribute work. Specifically, the Fair Queues feature began rejecting valid messages with an `InvalidParameterValue` error on a parameter our services had been sending successfully for months. There was no deploy or infrastructure change on our side that triggered this; AWS confirmed it as a service-side issue and rolled back the change on their end later that evening. Because the affected queues sit in the path between our API and our transcription workers, requests that could not be queued failed outright, and a portion of the work that did get through was delayed behind reduced throughput. We mitigated by disabling Fair Queues across our services and reverting to standard SQS queues, which restored normal operation roughly an hour after the first error. **What we're doing about it** * **Faster configuration changes.** Much of our service configuration lives in environment variables, which require task restarts to propagate. We are moving this to a global feature flag system so mitigations like this one take seconds rather than minutes. * **Expedited emergency deploys.** We are adding a reviewed break-glass path so urgent fixes can bypass non-essential CI steps. * **Multi-region failover.** We are prioritizing work to fail the US pipeline over to a second region when a regional provider dependency degrades, rather than relying on feature-level mitigations alone. We're sorry for the disruption. No customer audio or transcript data was lost, and any requests that failed during this window can be safely resubmitted.
Started
Sep 16, 2026 · 8:22 PM
Resolved
Sep 16, 2026 · 9:37 PM
Duration
1h 14m
Severity
Critical
Event timeline
Investigating
Sep 16 · 8:22 PM AssemblyAIWe are currently investigating an issue that is affecting all customers using our Async API on our US endpoint. Users will be receiving a server error. We will update with more information as we learn more. This began at 8pm UTC.
Identified
Sep 16 · 8:24 PM AssemblyAIThe issue has been identified. This is affecting approximately 50% of all async transcriptions on the US endpoint.
Identified
Sep 16 · 8:57 PM AssemblyAIWe are continuing to investigate these elevated errors. We will provide an update as soon as possible.
Identified
Sep 16 · 9:10 PM AssemblyAIWe are seeing a reduction in the number of errors, but we are still working to fully resolve the issue.
Monitoring
Sep 16 · 9:12 PM AssemblyAIWe have mitigated the issue and our Async API service is recovering.
Resolved
Sep 16 · 9:37 PM AssemblyAIThe issue, which was the result of an AWS service issue, has been resolved and we have continue to see good performance since implementing a fix so we are now closing this issue.
Postmortem
Sep 21 · 8:38 PM AssemblyAI**Summary** On September 16, 2026, between 20:04 and 21:08 UTC, our US Async API returned elevated server errors. At peak, roughly 50% of async transcription requests on the US endpoint failed, and a portion of the requests that did succeed took longer than usual to complete. Our EU endpoint and real-time/streaming services were not affected. **Root cause** The failure originated in AWS SQS, the managed queuing service our transcription pipeline uses to distribute work. Specifically, the Fair Queues feature began rejecting valid messages with an `InvalidParameterValue` error on a parameter our services had been sending successfully for months. There was no deploy or infrastructure change on our side that triggered this; AWS confirmed it as a service-side issue and rolled back the change on their end later that evening. Because the affected queues sit in the path between our API and our transcription workers, requests that could not be queued failed outright, and a portion of the work that did get through was delayed behind reduced throughput. We mitigated by disabling Fair Queues across our services and reverting to standard SQS queues, which restored normal operation roughly an hour after the first error. **What we're doing about it** * **Faster configuration changes.** Much of our service configuration lives in environment variables, which require task restarts to propagate. We are moving this to a global feature flag system so mitigations like this one take seconds rather than minutes. * **Expedited emergency deploys.** We are adding a reviewed break-glass path so urgent fixes can bypass non-essential CI steps. * **Multi-region failover.** We are prioritizing work to fail the US pipeline over to a second region when a regional provider dependency degrades, rather than relying on feature-level mitigations alone. We're sorry for the disruption. No customer audio or transcript data was lost, and any requests that failed during this window can be safely resubmitted.
Pattern
Increased Processing Times and Errors on EU Endpoint for Universal-3 Pro Async model
Oct 6, 2026 · 19m
View incident →Increased Processing Times and Errors on US Endpoint for Universal-3.5 Pro Async model
Oct 6, 2026 · 25m
View incident →Rejected Sessions on US Endpoint for Universal-3.6 Pro Realtime Model
Oct 6, 2026 · < 1m
View incident →Add it as a dependency monitor. The Free plan includes one.
Stay online, all the time, with Pulsetic's uptime prime.
By Designmodo
Designmodo Inc. 169 Madison Ave, #79627, New York, NY 10016, United States
Copyright © 2010-2026. Pulsetic® is a registered trademark.