Skip to content →
  1. Docs
  2. Post mortem on EU sync outage incident on Sept 27th, 2026
Sign upOpen app

Post mortem on EU sync outage incident on Sept 27th, 2026

From September 26 to 27, 2026, live updates were unavailable to users of Linear’s EU service for approximately 10 hours. Service has been restored, and changes made while offline should now sync. We apologize for the disruption.

Overview⁠

The sync source reads a database change stream. Sync servers subscribe to that stream, apply authorization filters for each client connection, and send updates to clients over WebSocket.

Sync source diagram

Timeline (UTC, September 26 - 27)⁠

  • 09:00: September 26th: database maintenance was started by a cron to repack tables in our eu cell
  • 00:00: September 27th: Due to earlier sync server issues, we delayed a rollout of our sync source until late Saturday
  • 00:05: Long running repack transactions prevented the sync source from creating a replication slot in the database
  • 00:52–00:56: the EU sync servers lost their connection to the sync source which replicates database changes via change data capture to the sync server. The servers hit a timing bug when the sync source became available.
  • Near the start of the outage: A monitor fired and continued reporting the failure, but notified an infrastructure alerts channel rather than paging the primary on-call engineer
  • 01:05: A separate “linear-sync has too few available replicas” escalation was acknowledged and then resolved without corrective action. One dashboard showed the sync server pods as running and healthy, while other dashboards unlinked in the monitor would have shown sync to not be sending sync packets to clients.
  • Throughout the outage: Sync traffic telemetry stopped, but there was no reliable page for not processing any traffic. Customers encountered 503 connection failures, for which there was no load-balancer alert. The persistent source-unavailable condition also did not repeatedly page an on-call responder.
  • 10:41: An engineer raised an incident internally and paged the team on call to investigate
  • 10:57: The team restarted affected sync servers, restoring their ability to connect
  • 11:15: The EU service was reported operational and connections restored. The team continued checking service health and processing queued changes.
  • 12:02: The public status-page incident was marked resolved

What happened⁠

Database maintenance delayed the sync source rollout by blocking creation of a replication slot. When the sync source became available, sync servers encountered a race between stream closure and decoding that caused them to drop a reset message. Replacement servers could not become ready after the previous servers stopped serving traffic. Restarting the affected servers restored service.

Detection and response⁠

Our alerts did not reliably page the on-call engineer despite sustained customer impact. The source-unavailable monitor fired, but notified an infrastructure alerts channel instead of paging the engineer on call. A separate low-replica alert was acknowledged and resolved without corrective action because its message was unclear. We also lacked reliable alerts for missing sync traffic and connection failures, and the persistent source-unavailable condition did not trigger repeat pages.

Prevention⁠

We deployed a fix for sync-server reset handling and improved alert routing, wording, missing-data coverage, and repeat notifications. We’re exploring volume-based paging alerts for weekend support volume. We are following up on sync source rollout safety.