Linode Status · History · Incident #5978

RESOLVED

Service Issue - Host Job Performance Degradation - Several Regions

Minor · Started Jul 28, 2026 · 6:14 PM

  • Duration

    18h 13m

  • Severity

    Minor

  • Detection lead

  • User reports

Summary

Service Issue - Host Job Performance Degradation - Several Regions

On 27 July 2026 at 3:30 UTC, Akamai observed an increase in errors when connecting to the Linode hosting database, primarily affecting Block Storage volume attachments. This resulted in host job failures and limited customer impact, with some users experiencing error messages and interrupted workflows. Elevated timeout rates were noted in logs for certain data center locations, coinciding with the incremental rollout of a new feature flag. Initial investigation revealed intermittent packet drops from the database proxy to client hosts during the TLS handshake. The current theory suggests that a DDoS-protection limit related to path MTU packet too big ICMP messages was reached. When the proxy sent TCP packets with a large MTU, the expected ICMP messages were dropped by Dallas gateway routers due to exceeding the configured allowable rate. This caused database proxy TCP connections to timeout to Compute Hosts. The issue was triggered by the enablement of the new feature flag, which changed the routing path and removed MTU clamping before packets reached the gateways. To mitigate the issue, Akamai rolled back the recent network change across affected Compute sites, starting at 20:50 UTC. As of 22:57 UTC, the rate of service restarts returned to pre-incident levels. Akamai is also planning a change to increase the allowable threshold for packet too big ICMP messages. This summary provides an overview of our current understanding of the incident given the information available. Our investigation is ongoing and any information herein is subject to change.


  • Started

    Jul 28, 2026 · 6:14 PM

  • Resolved

    Jul 29, 2026 · 12:27 PM

  • Duration

    18h 13m

  • Severity

    None

Event timeline

How this incident unfolded

  • Investigating

    Jul 28 · 6:14 PM Linode

    Our team is investigating an issue affecting the Block Storage service in several data center regions. This issue largely impacts attaching and detaching Block Storage volumes. During this time, users may experience volume attach/detach hangs, timeouts and errors with this service. We will share additional updates as we have more information.

  • Identified

    Jul 28 · 7:27 PM Linode

    Our team has identified the issue affecting the Block Storage service in our data centers. We are working quickly to implement a fix, and we will provide an update as soon as the solution is in place.

  • Identified

    Jul 28 · 8:17 PM Linode

    We would like to update that after additional investigation, the impact would manifest in delayed and sometimes failed host jobs, which could include many different actions on Linodes and not only impacting attaching and detaching Block Storage volumes as we mentioned in our initial update, we have updated the title to reflect the updated impact. We are working quickly to implement a fix, and we will provide an update as soon as the solution is in place.

  • Monitoring

    Jul 28 · 10:51 PM Linode

    A fix has been implemented and we are monitoring the results.

  • Resolved

    Jul 29 · 12:27 PM Linode

    We haven’t observed any additional host jobs performance degradation issues, and will now consider this incident resolved. If you continue to experience problems, please <a href="https://cloud.linode.com/support/tickets">open a Support ticket</a> for assistance.

  • Postmortem

    Jul 31 · 6:57 PM Linode

    On 27 July 2026 at 3:30 UTC, Akamai observed an increase in errors when connecting to the Linode hosting database, primarily affecting Block Storage volume attachments. This resulted in host job failures and limited customer impact, with some users experiencing error messages and interrupted workflows. Elevated timeout rates were noted in logs for certain data center locations, coinciding with the incremental rollout of a new feature flag. Initial investigation revealed intermittent packet drops from the database proxy to client hosts during the TLS handshake. The current theory suggests that a DDoS-protection limit related to path MTU packet too big ICMP messages was reached. When the proxy sent TCP packets with a large MTU, the expected ICMP messages were dropped by Dallas gateway routers due to exceeding the configured allowable rate. This caused database proxy TCP connections to timeout to Compute Hosts. The issue was triggered by the enablement of the new feature flag, which changed the routing path and removed MTU clamping before packets reached the gateways. To mitigate the issue, Akamai rolled back the recent network change across affected Compute sites, starting at 20:50 UTC. As of 22:57 UTC, the rate of service restarts returned to pre-incident levels. Akamai is also planning a change to increase the allowable threshold for packet too big ICMP messages. This summary provides an overview of our current understanding of the incident given the information available. Our investigation is ongoing and any information herein is subject to change.

Get alerted before the next Linode outage.

Pulsetic catches degradations minutes before vendors acknowledge them.

Start monitoring free