Blue Canvas Status · History · Incident #124716

RESOLVED

Sync servers crashing

Minor · Started Aug 19, 2026 · 8:05 AM

  • Duration

    1d 8h 36m

  • Severity

    Minor

  • Detection lead

  • User reports

Summary

Sync servers crashing

Systems are operational, no incidents since last update, we don't expect further downtime while we optimize the data center.


  • Started

    Aug 19, 2026 · 8:05 AM

  • Resolved

    Aug 20, 2026 · 4:41 PM

  • Duration

    1d 8h 36m

  • Severity

    Minor

Event timeline

How this incident unfolded

  • Investigating

    Aug 19 · 8:05 AM Blue Canvas

    Syncs and deploys are currently not happening. The underlying servers are crashing for various reasons. We're working on pinpointing the exact issue.

  • Monitoring

    Aug 19 · 11:19 AM Blue Canvas

    We have identified the cause of the intermittent delays affecting branch synchronization and related processing over the past day. A fix has been deployed and the affected processing infrastructure has been replaced. All synchronization services are operating normally and we are monitoring closely to confirm the fix holds under full load. Synchronization jobs that were delayed during this period have resumed automatically — no action is required from customers, and no data was lost.

  • Investigating

    Aug 19 · 1:06 PM Blue Canvas

    Unfortunately the system degraded again after ~1 hour.

  • Monitoring

    Aug 19 · 1:58 PM Blue Canvas

    The "ring" that serves syncs and deploys is stable again. This should get the syncs and deploys running as usual. In the meantime we'll keep an eye out and act immediately if something changes.

  • Monitoring

    Aug 20 · 9:00 AM Blue Canvas

    Earlier today we identified an issue where ~9% of branches experienced extra delay on both syncs and deploys. The fixes from yesterday helped containing the issue so the other 91% of branches could continue as normal. We've resolved this issue by scaling up the units that have to do the work. There are currently no known issues, but we are looking for a sustainable solution, which may involve one or more releases. In technical terms: The scaling and self-repair configuration we've used for the last year is no long viable. We've greatly optimized on CPU, to the point where it is no longer a reliable metric for scaling in or out. We've disabled the scaling, and manually put the system on a configuration that works right now, and continue to fine tune it while we work on a permanent and automated solution.

  • Resolved

    Aug 20 · 4:41 PM Blue Canvas

    Systems are operational, no incidents since last update, we don't expect further downtime while we optimize the data center.

Get alerted before the next Blue Canvas outage.

Pulsetic catches degradations minutes before vendors acknowledge them.