Web, Payments, Bytes & Ventures

WPBV Radio

Tune into tech and money.

OpenAI outage, 11 December 2024

The OpenAI outage of 11 December 2024: what failed, why, and what OpenAI changed, from its own incident report.

Provider
OpenAI
Date
Duration, hours
4.4
Cause
Configuration change
Regions
Global

OpenAI ChatGPT/API/Sora outage from telemetry deploy (11 December 2024). The disruption lasted about 4 hours 22 minutes, according to OpenAI’s own timeline.

What happened. A new telemetry service’s configuration made every node in each Kubernetes cluster run expensive API operations, overwhelming the Kubernetes API servers in the largest clusters. DNS-based service discovery depends on the control plane, so services failed once DNS caches expired, and engineers were locked out of the control plane needed to roll back.

What changes. OpenAI is implementing more robust phased rollouts with better monitoring, fault-injection testing, emergency control-plane access, and decoupling of the Kubernetes data and control planes.

Questions

What caused the OpenAI outage in December 2024?

A new telemetry service’s configuration made every node in each Kubernetes cluster run expensive API operations, overwhelming the Kubernetes API servers in the largest clusters.

Which services were affected?

ChatGPT, OpenAI API, Sora.

Read how shared dependencies spread a failure in why one cloud outage can break unrelated apps, or browse the outage archive.

Started (UTC)
2024-12-11 23:16 UTC
Resolved (UTC)
2024-12-12 03:38 UTC
Services affected
ChatGPT, OpenAI API, Sora
Incident report
status.openai.com
Sources
SourceLink
OpenAI incident reporthttps://status.openai.com/incidents/ctrsv3lwd797
WPBV Radio

What are you looking for?

Search by headline, topic or keyword.