During a network partition, Orchestrator began leadership deselection in the primary data center; West Coast and public-cloud nodes reached quorum and began failing clusters over to the West Coast.
GitHub
github.blog
Production Incident Atlas
A network partition triggered cross-region database failover, leaving clusters with divergent writes and causing more than a day of degraded service while GitHub restored consistency.
What happened
During a network partition, Orchestrator began leadership deselection in the primary data center; West Coast and public-cloud nodes reached quorum and began failing clusters over to the West Coast.
GitHub
github.blog
Internal monitoring generated alerts for numerous faults. Responders later found unexpected topologies across multiple database clusters.
GitHub
github.blog
GitHub paused jobs that wrote metadata, including webhook delivery and Pages builds, to avoid risking data already received.
GitHub
github.blog
All database primaries were established in the US East Coast data center; delayed read replicas still caused some users to see inconsistent results.
GitHub
github.blog
After replicas synchronized, GitHub failed back to the original topology and began processing the accumulated backlog while keeping incident status red.
GitHub
github.blog
All pending webhooks and Pages builds had been processed, system integrity was confirmed, and GitHub updated the site status to green.
GitHub
github.blog
GitHub’s internal monitoring began alerting at 22:54 UTC. By 23:02, responders had found unexpected database-cluster topologies with servers only in the US West Coast data center. The postmortem says an earlier network partition had caused Orchestrator to begin failing clusters over to that site.
GitHub restored East Coast primaries by 11:12 UTC on October 22, then waited for replicas to synchronize before returning to the original topology at 16:24 UTC. It kept incident status degraded while processing queued work; all pending webhooks and Pages builds were processed and status returned to green at 23:03 UTC.
GitHub said it would configure Orchestrator to prevent cross-region primary promotion, improve status reporting so users could see affected components, and accelerate work toward serving traffic from multiple data centers. Its response prioritized preserving data consistency over shortening the service-degradation window.
GitHub’s post-incident analysis makes the recovery trade-off explicit: the team accepted a longer period of degraded service to protect the integrity of writes that had reached the two sides of a network partition. Promoting a West Coast primary restored write routing, but the East Coast held a short window of writes that had not replicated. Failing back without resolving that divergence risked consistency.
The incident then became a recovery-throughput problem. Restoring multiple terabytes from remote backups took hours, and read replicas fell behind as traffic returned. GitHub described a nonlinear catch-up curve and added replicas to reduce their utilization so replication could catch up.
The postmortem’s follow-up actions addressed assumptions at several layers: constrain cross-region primary promotion, improve component-level status communication, and continue work toward serving from multiple data centers. The queued webhook and Pages-build backlog also shows why recovery includes background work, not only restoring request handling.
Last reviewed 10/9/2026 by Nexus Lead Architect.
Corrections or additional primary evidence? Contact the editorial team.