Web, Payments, Bytes & Ventures

WPBV Radio

Tune into tech and money.

AWS outage, 30 July 2024

The AWS outage of 30 July 2024: what failed, why, and what AWS changed, from its own incident report.

Provider
AWS
Date
Duration, hours
6.9
Cause
Software bug
Regions
us-east-1 (N. Virginia)

AWS us-east-1 Kinesis Data Streams cell outage (30 July 2024). The disruption lasted about 6 hours 52 minutes, according to AWS’s own timeline.

What happened. A routine deployment on an internal Kinesis Data Streams cell with an unusual workload (very many low-throughput shards) triggered incorrect behavior in the new cell management system. Work was piled onto a few hosts and the cell degraded, which hit AWS services built on that cell.

What changes. AWS is adding data-plane capacity, load shedding and connection-limit changes, and updating the cell management system to handle that workload profile.

Questions

What caused the AWS outage in July 2024?

A routine deployment on an internal Kinesis Data Streams cell with an unusual workload (very many low-throughput shards) triggered incorrect behavior in the new cell management system.

Which services were affected?

CloudWatch Logs, Amazon Data Firehose, S3 event notifications, ECS, Lambda, Redshift, Glue (via internal Kinesis cell).

Read how shared dependencies spread a failure in why one cloud outage can break unrelated apps, or browse the outage archive.

Started (UTC)
2024-07-30 21:45 UTC
Resolved (UTC)
2024-07-31 04:37 UTC
Services affected
CloudWatch Logs, Amazon Data Firehose, S3 event notifications, ECS, Lambda, Redshift, Glue (via internal Kinesis cell)
Incident report
aws.amazon.com
Sources
SourceLink
AWS incident reporthttps://aws.amazon.com/message/073024/
WPBV Radio

What are you looking for?

Search by headline, topic or keyword.