AWS outage, 30 July 2024
The AWS outage of 30 July 2024: what failed, why, and what AWS changed, from its own incident report.
- Provider
- AWS
- Date
- Duration, hours
- 6.9
- Cause
- Software bug
- Regions
- us-east-1 (N. Virginia)
AWS us-east-1 Kinesis Data Streams cell outage (30 July 2024). The disruption lasted about 6 hours 52 minutes, according to AWS’s own timeline.
What happened. A routine deployment on an internal Kinesis Data Streams cell with an unusual workload (very many low-throughput shards) triggered incorrect behavior in the new cell management system. Work was piled onto a few hosts and the cell degraded, which hit AWS services built on that cell.
What changes. AWS is adding data-plane capacity, load shedding and connection-limit changes, and updating the cell management system to handle that workload profile.
Questions
What caused the AWS outage in July 2024?
A routine deployment on an internal Kinesis Data Streams cell with an unusual workload (very many low-throughput shards) triggered incorrect behavior in the new cell management system.
Which services were affected?
CloudWatch Logs, Amazon Data Firehose, S3 event notifications, ECS, Lambda, Redshift, Glue (via internal Kinesis cell).
Read how shared dependencies spread a failure in why one cloud outage can break unrelated apps, or browse the outage archive.
- Started (UTC)
- 2024-07-30 21:45 UTC
- Resolved (UTC)
- 2024-07-31 04:37 UTC
- Services affected
- CloudWatch Logs, Amazon Data Firehose, S3 event notifications, ECS, Lambda, Redshift, Glue (via internal Kinesis cell)
- Incident report
- aws.amazon.com
- Sources
Source Link AWS incident report https://aws.amazon.com/message/073024/
