OpenAI outage, 11 December 2024
The OpenAI outage of 11 December 2024: what failed, why, and what OpenAI changed, from its own incident report.
- Provider
- OpenAI
- Date
- Duration, hours
- 4.4
- Cause
- Configuration change
- Regions
- Global
OpenAI ChatGPT/API/Sora outage from telemetry deploy (11 December 2024). The disruption lasted about 4 hours 22 minutes, according to OpenAI’s own timeline.
What happened. A new telemetry service’s configuration made every node in each Kubernetes cluster run expensive API operations, overwhelming the Kubernetes API servers in the largest clusters. DNS-based service discovery depends on the control plane, so services failed once DNS caches expired, and engineers were locked out of the control plane needed to roll back.
What changes. OpenAI is implementing more robust phased rollouts with better monitoring, fault-injection testing, emergency control-plane access, and decoupling of the Kubernetes data and control planes.
Questions
What caused the OpenAI outage in December 2024?
A new telemetry service’s configuration made every node in each Kubernetes cluster run expensive API operations, overwhelming the Kubernetes API servers in the largest clusters.
Which services were affected?
ChatGPT, OpenAI API, Sora.
Read how shared dependencies spread a failure in why one cloud outage can break unrelated apps, or browse the outage archive.
- Started (UTC)
- 2024-12-11 23:16 UTC
- Resolved (UTC)
- 2024-12-12 03:38 UTC
- Services affected
- ChatGPT, OpenAI API, Sora
- Incident report
- status.openai.com
- Sources
Source Link OpenAI incident report https://status.openai.com/incidents/ctrsv3lwd797
