← All incidents

Production Incident Atlas

Amazon S3 us-east-1 disruption: capacity removal and a slow restart

An incorrectly entered operational command removed more S3 subsystem capacity than intended, affecting S3 and dependent AWS services in Northern Virginia.

Amazon S3 (AWS US-EAST-1)Software & infrastructureHuman or process failureMajor severityAvailability

What happened

Incident timeline

4 documented events
  1. An authorized S3 team member ran a playbook command intended to remove a small number of servers. An incorrectly entered input caused a larger set to be removed, including capacity used by the index and placement subsystems.

  2. The index subsystem had activated enough capacity to begin servicing S3 GET, LIST, and DELETE requests.

  3. The index subsystem was fully recovered, and its GET, LIST, and DELETE APIs were functioning normally.

  4. The placement subsystem completed recovery; AWS reported that S3 was operating normally.

Detection

The S3 team was investigating a slow billing-system process when a capacity-removal command was executed with an incorrect input. AWS reported that the command removed more servers than intended and that dependent S3 systems became unable to serve requests.

Recovery

The index subsystem began serving GET, LIST, and DELETE requests by 12:26 PM PST and was fully recovered by 1:18 PM. The placement subsystem completed recovery at 1:54 PM, when AWS said S3 was operating normally. Some dependent services then needed more time to drain backlogs.

Lessons afterward

AWS said it changed the capacity-removal tool to remove capacity more slowly and added minimum-capacity safeguards. It also began auditing other operational tools, reprioritized further partitioning of the index subsystem, and changed its Service Health Dashboard administration console to run across multiple regions.

Editorial analysis

Analysis

This incident shows how an operational action can cross subsystem boundaries. The command targeted capacity used by billing, but the removed servers also supported the S3 index and placement subsystems. Because those systems had not undergone a full restart in the larger regions for years, recovery and metadata integrity checks took longer than expected.

Incident communication has dependencies too. AWS reported that its status-dashboard administration console depended on S3, so service-specific status updates were unavailable until 11:37 AM PST. The later change to run that console across multiple regions addressed a separate failure path: restoring a product does not automatically restore the tools used to communicate about it.

The postmortem describes concrete controls: slower capacity removal, minimum-capacity safeguards, reviews of other operational tools, and accelerated subsystem partitioning.

Last reviewed 10/9/2026 by Nexus Lead Architect.

Change notes

  • 10/9/2026: Initial publication from https://aws.amazon.com/message/41926/

Corrections or additional primary evidence? Contact the editorial team.