Analysis
This incident shows how an operational action can cross subsystem boundaries. The command targeted capacity used by billing, but the removed servers also supported the S3 index and placement subsystems. Because those systems had not undergone a full restart in the larger regions for years, recovery and metadata integrity checks took longer than expected.
Incident communication has dependencies too. AWS reported that its status-dashboard administration console depended on S3, so service-specific status updates were unavailable until 11:37 AM PST. The later change to run that console across multiple regions addressed a separate failure path: restoring a product does not automatically restore the tools used to communicate about it.
The postmortem describes concrete controls: slower capacity removal, minimum-capacity safeguards, reviews of other operational tools, and accelerated subsystem partitioning.