← All incidents

Production Incident Atlas

GitHub’s October 2018 incident: network partition and database recovery

A network partition triggered cross-region database failover, leaving clusters with divergent writes and causing more than a day of degraded service while GitHub restored consistency.

GitHub.comSoftware & infrastructureDesign flawMajor severityAvailability

What happened

Incident timeline

6 documented events
  1. During a network partition, Orchestrator began leadership deselection in the primary data center; West Coast and public-cloud nodes reached quorum and began failing clusters over to the West Coast.

  2. Internal monitoring generated alerts for numerous faults. Responders later found unexpected topologies across multiple database clusters.

  3. GitHub paused jobs that wrote metadata, including webhook delivery and Pages builds, to avoid risking data already received.

  4. All database primaries were established in the US East Coast data center; delayed read replicas still caused some users to see inconsistent results.

  5. After replicas synchronized, GitHub failed back to the original topology and began processing the accumulated backlog while keeping incident status red.

  6. All pending webhooks and Pages builds had been processed, system integrity was confirmed, and GitHub updated the site status to green.

Detection

GitHub’s internal monitoring began alerting at 22:54 UTC. By 23:02, responders had found unexpected database-cluster topologies with servers only in the US West Coast data center. The postmortem says an earlier network partition had caused Orchestrator to begin failing clusters over to that site.

Recovery

GitHub restored East Coast primaries by 11:12 UTC on October 22, then waited for replicas to synchronize before returning to the original topology at 16:24 UTC. It kept incident status degraded while processing queued work; all pending webhooks and Pages builds were processed and status returned to green at 23:03 UTC.

Lessons afterward

GitHub said it would configure Orchestrator to prevent cross-region primary promotion, improve status reporting so users could see affected components, and accelerate work toward serving traffic from multiple data centers. Its response prioritized preserving data consistency over shortening the service-degradation window.

Editorial analysis

Analysis

GitHub’s post-incident analysis makes the recovery trade-off explicit: the team accepted a longer period of degraded service to protect the integrity of writes that had reached the two sides of a network partition. Promoting a West Coast primary restored write routing, but the East Coast held a short window of writes that had not replicated. Failing back without resolving that divergence risked consistency.

The incident then became a recovery-throughput problem. Restoring multiple terabytes from remote backups took hours, and read replicas fell behind as traffic returned. GitHub described a nonlinear catch-up curve and added replicas to reduce their utilization so replication could catch up.

The postmortem’s follow-up actions addressed assumptions at several layers: constrain cross-region primary promotion, improve component-level status communication, and continue work toward serving from multiple data centers. The queued webhook and Pages-build backlog also shows why recovery includes background work, not only restoring request handling.

Last reviewed 10/9/2026 by Nexus Lead Architect.

Change notes

  • 10/9/2026: Initial publication from https://github.blog/news-insights/company-news/oct21-post-incident-analysis/

Corrections or additional primary evidence? Contact the editorial team.