GitHub Status · History · Incident #6658
RESOLVEDActions delays in starting runs
Minor · Started Aug 24, 2026 · 1:56 PM
GitHub Status · History · Incident #6658
RESOLVEDMinor · Started Aug 24, 2026 · 1:56 PM
Duration
37m
Severity
Minor
Detection lead
—
User reports
—
Summary
On August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright. <br /> <br />The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC. <br /><br />To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.
Started
Aug 24, 2026 · 1:56 PM
Resolved
Aug 24, 2026 · 2:34 PM
Duration
37m
Severity
Minor
Event timeline
Investigating
Aug 24 · 1:56 PM GitHubWe are investigating reports of degraded performance for Actions
Investigating
Aug 24 · 2:22 PM GitHubFailures while queuing and running Actions jobs for a subset of customers are now resolving. We are monitoring for full recovery.
Monitoring
Aug 24 · 2:26 PM GitHubThe degradation affecting Actions has been mitigated. We are monitoring to ensure stability.
Resolved
Aug 24 · 2:34 PM GitHubOn August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright. <br /> <br />The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC. <br /><br />To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.
Pulsetic catches degradations minutes before vendors acknowledge them.
Stay online, all the time, with Pulsetic's uptime prime.
By Designmodo
Designmodo Inc. 169 Madison Ave, #79627, New York, NY 10016, United States
Copyright © 2010-2026. Pulsetic® is a registered trademark.