My Take on the AWS us-east-1 Incident

An SRE-focused analysis of the November 2025 AWS us-east-1 incident and its cascade from DynamoDB DNS through EC2 and network recovery.

This incident analysis is primarily from an SRE perspective. I reviewed multiple sources and articles to develop the theory below.

Illustration of a Tier 0 service failure cascading through the EC2 control plane into network and load-balancing failures

High-Level Summary: The Cascade Failure

This was not a single event; it was a classic cascade failure.

A highly improbable, latent race condition in a core “Tier 0” service (DynamoDB’s DNS) was the initial spark. The subsequent recovery process from this initial failure then triggered a second, more severe “thundering herd” problem in the EC2 control plane. This, in turn, created a third wave of failures in the network and load-balancing layers.

The key takeaway is that the system’s recovery path was itself a failure mode that had not been tested at this scale.

Detailed Event Timeline

Here is the high-level progression of the cascading failures.

Timeline table for phase one of the AWS incident, from the DynamoDB DNS race condition through DNS recovery

Timeline table for phase two, showing EC2 control-plane lease recovery and congestive collapse

Timeline table for phase three, showing network backlog, NLB flapping, mitigation, and full recovery

The SRE Lens: A Three-Phase Breakdown

Here is what actually happened, explained from an engineering perspective.

Phase 1: The Initial Spark (DynamoDB DNS Failure)

This was a latent bug in a complex, resilient system.

The System: As the diagram shows, the DNS Planner creates “plans” (LB sets + weights). Multiple DNS Enactors (one per AZ) fetch the latest plan and apply it to Route 53.

DynamoDB DNS architecture with a planner creating DNS plans and three availability-zone enactors updating Route 53

The Race Condition:

  1. Enactor-1 (Slow) fetched Plan_A but was heavily delayed.
  2. Planner created Plan_B (a newer plan).
  3. Enactor-2 (Fast) fetched Plan_B, applied it, and—as part of cleanup—marked the “significantly older” Plan_A for deletion.
  4. Enactor-1 (Slow) finally applied its stale Plan_A, overwriting the correct Plan_B. Its initial check that its plan was “newer” was stale due to the long processing delay.
  5. The cleanup process saw that Plan_A (which was now active) was also marked for deletion. It executed the delete, wiping all IP addresses for the regional endpoint.

SRE Lesson: This is a catastrophic blast radius from a single, automated process. The automation, built for resilience, lacked a critical “circuit breaker” (e.g., “NEVER apply a plan that results in 0 records”). The failure of a Tier 0 service (DynamoDB) immediately cascades to all dependent services.

Phase 2: The Cascade (EC2 “Congestive Collapse”)

The recovery from Phase 1 triggered Phase 2.

The Dependency: The EC2 control plane (DWFM) uses DynamoDB to track the state of all physical servers (“droplets”) via a “lease” system.

The Failure:

  1. During the 3-hour DynamoDB outage, all these leases expired.
  2. When DynamoDB came back at 2:25 AM, every DWFM instance tried to renew leases for every server in the region at the same time.
  3. This internal thundering herd completely overwhelmed the DWFM system. It entered congestive collapse: it was so busy processing its massive work queue (and subsequent retries) that it couldn’t make any forward progress.

SRE Lesson: Recovery is a distinct and dangerous failure mode. The system was never designed or tested to recover its entire state from scratch simultaneously. This highlights the need for jitter, exponential backoff, and throttling in recovery workflows, not just in public-facing APIs.

Phase 3: The Aftershock (Network Backlog & NLB Flapping)

The partial recovery from Phase 2 triggered Phase 3.

The Dependency: When an EC2 instance launches, the Network Manager must provision its network configuration.

The Failure:

  1. As EC2 started to recover (post 5:28 AM), the massive backlog of queued network configuration requests flooded the Network Manager, causing it to slow down.
  2. This created a new race condition: EC2 instances were launching but had no network connectivity for several minutes.
  3. This broke the Network Load Balancer (NLB), which was launching new nodes. The NLB’s health checker would see the new node, try to connect, fail (because it had no network yet), and mark the node as “unhealthy.”
  4. Minutes later, the network would connect, the next health check would pass, and the node would be added back. This “flapping” destabilized the load balancers and caused connection errors for many other services (Lambda, Connect, etc.).

SRE Lesson: This shows the danger of tightly coupled dependencies in a startup process. The EC2 instance-ready signal was “true” before the network-ready signal was “true,” violating the assumptions of downstream services like NLB. This is a subtle but critical failure in system integration.

Key SRE Lessons Learned

  1. Test Your Recovery Path: A system’s recovery logic is a first-class feature and a primary source of failure. It must be tested at scale (e.g., via “chaos” drills) to ensure it doesn’t cause a congestive collapse.
  2. Beware “Tier 0” Dependencies: The blast radius of core services like DNS and core databases (DynamoDB) is “everything.” These systems must have robust, human-in-the-loop circuit breakers for destructive actions (like deleting all DNS records).
  3. Automation Needs Guardrails: The DNS Enactor’s cleanup process was the “murder weapon.” Automation that “cleans up” is dangerous. It should never be able to delete an active plan, and it should fail-safe if an operation would reduce capacity to zero.
  4. Thundering Herds are Internal Too: We often design for external thundering herds, but this event shows a massive internal one. Recovery processes must be “desynchronized” using jitter and backoff to prevent self-DDoS.

Originally published on Medium on November 5, 2025.