CircleCI Status · History · Incident #146154

RESOLVED

Elevated wait times for machine jobs

Minor · Started Oct 8, 2026 · 1:00 PM

  • Duration

    3h 35m

  • Severity

    Minor

  • Detection lead

    —

  • User reports

    Peak 1/hr

Summary

Elevated wait times for machine jobs

## Summary Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two \(unrelated\) causes, both of which the CircleCI engineering team is actively mitigating: * Available Cloud Computing Capacity * Internal Infrastructure ## Available Cloud Computing Capacity * **Problem**: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs. * **Incidents**: * [Elevated wait times for machine jobs](https://status.circleci.com/incidents/jn756xtsy4xg) * [Increased task wait times for Docker Gen2](https://status.circleci.com/incidents/39rfrjmk5pvd) * [Elevated level of infra fails on customer jobs](https://status.circleci.com/incidents/0z31ldjd8m5v) * [Delay on starting Machine Job Tasks](https://status.circleci.com/incidents/tj6wdnfjwr64) * [Delays starting Gen 2 Docker Jobs](https://status.circleci.com/incidents/22mrv46g6n7g) * **Mitigations and Resolutions**: * We are expanding our set of cloud computing regions to include additional regions with available high-performance instances. * We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions. * We are further expanding the set of instance types we can offer to customers. * We will publish a detailed Incident Report on these incidents on October 9, 2026. ## Internal Infrastructure * **Problem**: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database. * **Incidents:** * [Delays in pipelines, workflows, UI data, and notifications](https://status.circleci.com/incidents/z4hbq8h49j8c) * [Emergency Infrastructure Maintenance](https://status.circleci.com/incidents/g60yc9bvvkmx) * **Mitigations and Resolutions**: * By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them. ## Our Commitment For avoidance of doubt, * These particular incidents are not related to ongoing outages at Github * We are not currently migrating our infrastructure or services * We are not currently under cyberattack Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents. Please reach out to our support team with any questions or concerns.


  • Started

    Oct 8, 2026 · 1:00 PM

  • Resolved

    Oct 8, 2026 · 4:35 PM

  • Duration

    3h 35m

  • Severity

    Minor

Event timeline

How this incident unfolded

  • ◆

    Identified

    Oct 8 · 1:00 PM CircleCI

    Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.

  • ◆

    Identified

    Oct 8 · 1:34 PM CircleCI

    Customers using Linux machine jobs are experiencing elevated wait times, averaging about 11 minutes, with the longest waits exceeding 30 minutes on the medium, arm.medium and arm.large resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:00 UTC.

  • ◆

    Identified

    Oct 8 · 2:04 PM CircleCI

    Customers using Linux machine jobs are experiencing elevated wait times, averaging about 40 minutes, with the longest waits exceeding 50 minutes. Most Linux machine resource classes are affected, including medium, large, xlarge, 2xlarge and their Arm equivalents. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:30 UTC.

  • ◆

    Identified

    Oct 8 · 2:37 PM CircleCI

    Customers using Linux machine jobs and remote Docker are experiencing elevated wait times. Wait times have started to decrease and now average about 20 minutes, with the longest waits exceeding 40 minutes on some resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 15:00 UTC.

  • ◆

    Identified

    Oct 8 · 3:01 PM CircleCI

    A fix has been deployed and wait times are decreasing, but customers using Linux machine jobs and Remote Docker are still experiencing delays. Wait times currently average about 11 minutes, with the longest waits exceeding 35 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 15:30 UTC.

  • ◆

    Identified

    Oct 8 · 3:31 PM CircleCI

    Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 6 minutes. The longest waits, up to about 20 minutes, are on the 2xlarge, arm.2xlarge and gpu.nvidia.small resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:00 UTC.

  • ◆

    Identified

    Oct 8 · 3:59 PM CircleCI

    Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 90 seconds, with the longest waits up to about 6 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:30 UTC.

  • ◉

    Monitoring

    Oct 8 · 4:09 PM CircleCI

    Wait times for customers using Linux machine jobs and remote Docker have returned to normal. We are monitoring to confirm wait times remain stable while we continue to add capacity. We will provide another update by 16:30 UTC.

  • ✓

    Resolved

    Oct 8 · 4:35 PM CircleCI

    Between 12:40 UTC and 16:00 UTC on October 8, customers using Linux machine jobs and remote Docker experienced elevated wait times. The issue has been resolved and wait times have returned to normal. We thank you for your patience while our team worked on implementing a fix.

  • ✓

    Postmortem

    Oct 8 · 5:54 PM CircleCI

    ## Summary Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two \(unrelated\) causes, both of which the CircleCI engineering team is actively mitigating: * Available Cloud Computing Capacity * Internal Infrastructure ## Available Cloud Computing Capacity * **Problem**: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs. * **Incidents**: * [Elevated wait times for machine jobs](https://status.circleci.com/incidents/jn756xtsy4xg) * [Increased task wait times for Docker Gen2](https://status.circleci.com/incidents/39rfrjmk5pvd) * [Elevated level of infra fails on customer jobs](https://status.circleci.com/incidents/0z31ldjd8m5v) * [Delay on starting Machine Job Tasks](https://status.circleci.com/incidents/tj6wdnfjwr64) * [Delays starting Gen 2 Docker Jobs](https://status.circleci.com/incidents/22mrv46g6n7g) * **Mitigations and Resolutions**: * We are expanding our set of cloud computing regions to include additional regions with available high-performance instances. * We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions. * We are further expanding the set of instance types we can offer to customers. * We will publish a detailed Incident Report on these incidents on October 9, 2026. ## Internal Infrastructure * **Problem**: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database. * **Incidents:** * [Delays in pipelines, workflows, UI data, and notifications](https://status.circleci.com/incidents/z4hbq8h49j8c) * [Emergency Infrastructure Maintenance](https://status.circleci.com/incidents/g60yc9bvvkmx) * **Mitigations and Resolutions**: * By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them. ## Our Commitment For avoidance of doubt, * These particular incidents are not related to ongoing outages at Github * We are not currently migrating our infrastructure or services * We are not currently under cyberattack Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents. Please reach out to our support team with any questions or concerns.

User reports · incident window

Report volume over time

Hourly report count. Red bars mark the incident window.

Peak: 1 reports/hr

← Before incident After →

Get an email when CircleCI opens, updates or resolves an incident.

Add it as a dependency monitor. The Free plan includes one.