AWS outage, 7 December 2021
The AWS outage of 7 December 2021: what failed, why, and what AWS changed, from its own incident report.
- Provider
- AWS
- Date
- Duration, hours
- 6.9
- Cause
- Capacity/overload
- Regions
- us-east-1 (N. Virginia)
AWS us-east-1 internal network congestion outage (7 December 2021). The disruption lasted about 6 hours 52 minutes, according to AWS’s own timeline.
What happened. An automated activity to scale capacity of an AWS service triggered unexpected behavior from many clients in AWS’s internal network, causing a surge of connections that overwhelmed networking devices between the internal and main AWS networks. A latent issue stopped clients from backing off, and the congestion also blinded internal monitoring, slowing recovery.
What changes. AWS disabled the triggering scaling activities, committed to fix the client back-off bug, and deployed extra network configuration to protect the affected devices.
Questions
What caused the AWS outage in December 2021?
An automated activity to scale capacity of an AWS service triggered unexpected behavior from many clients in AWS’s internal network, causing a surge of connections that overwhelmed networking devices between the internal and main AWS networks.
Which services were affected?
EC2 APIs, Route 53 APIs, AWS Console, STS, API Gateway, EventBridge, CloudWatch, VPC Endpoints for S3/DynamoDB, RDS/EMR/WorkSpaces provisioning.
Read how shared dependencies spread a failure in why one cloud outage can break unrelated apps, or browse the outage archive.
- Started (UTC)
- 2021-12-07 15:30 UTC
- Resolved (UTC)
- 2021-12-07 22:22 UTC
- Services affected
- EC2 APIs, Route 53 APIs, AWS Console, STS, API Gateway, EventBridge, CloudWatch, VPC Endpoints for S3/DynamoDB, RDS/EMR/WorkSpaces provisioning
- Incident report
- aws.amazon.com
- Sources
Source Link AWS incident report https://aws.amazon.com/message/12721/
