CloudRepo Status · History · Incident #124724

RESOLVED

Elevated 502 Errors

Minor · Started Jan 13, 2025 · 7:37 AM

  • Duration

    2h 50m

  • Severity

    Minor

  • Detection lead

  • User reports

Summary

Elevated 502 Errors

We have not observed a single 502 across our systems for the past two hours, while close to peak load. We will consider this issue resolved as there is no current customer impact. We will continue to monitor and evaluate internally in order to prevent any future disruption.


  • Started

    Jan 13, 2025 · 7:37 AM

  • Resolved

    Jan 13, 2025 · 10:27 AM

  • Duration

    2h 50m

  • Severity

    Minor

Event timeline

How this incident unfolded

  • Investigating

    Jan 13 · 7:37 AM CloudRepo

    We have received reports of 502 errors which are causing builds to break. Our internal metrics indicate that this is affecting between 1-5% of all requests. We have elevated this to a critical issue and we are actively investigating.

  • Investigating

    Jan 13 · 7:38 AM CloudRepo

    We are continuing to investigate this issue.

  • Investigating

    Jan 13 · 8:13 AM CloudRepo

    While we are identifying the root cause of the issue, we have doubled the size of our clusters (cpu, memory, and network) in order to reduce the frequency of these errors.

  • Monitoring

    Jan 13 · 8:54 AM CloudRepo

    We believe the issue has been caused by an increase in load as well as a potential resource leak of some sort. We have scaled the size of all of our resources by 2x in order to immediately reduce the impact to our partners while we investigate the resource leak. Since we have scaled up at 1415 GMT, we have not seen a single 502 pass through our load balancers. We will continue monitoring closely while we search for root cause.

  • Resolved

    Jan 13 · 10:27 AM CloudRepo

    We have not observed a single 502 across our systems for the past two hours, while close to peak load. We will consider this issue resolved as there is no current customer impact. We will continue to monitor and evaluate internally in order to prevent any future disruption.

Get alerted before the next CloudRepo outage.

Pulsetic catches degradations minutes before vendors acknowledge them.