Datadog outage, 8 March 2023
The Datadog outage of 8 March 2023: what failed, why, and what Datadog changed, from its own incident report.
- Provider
- Datadog
- Date
- Duration, hours
- 26.9
- Cause
- Software bug
- Regions
- US1, EU1, US3, US4, US5 regions
Datadog multi-region outage from automatic systemd update (8 March 2023). The disruption lasted about 26 hours 55 minutes, according to Datadog’s own timeline.
What happened. A legacy security update channel in Datadog’s base OS image automatically applied a systemd update across many VMs in the same 06:00-07:00 UTC window. When systemd-networkd restarted it deleted routes managed by the Cilium CNI plugin, taking tens of thousands of nodes offline across five regions and three cloud providers.
What changes. Datadog disabled the legacy update channel in all regions and said it will prove the platform can run degraded rather than fully down, plus improve customer guidance and status communication.
Questions
What caused the Datadog outage in March 2023?
A legacy security update channel in Datadog’s base OS image automatically applied a systemd update across many VMs in the same 06:00-07:00 UTC window.
Which services were affected?
All Datadog services (web app, APIs, monitors, data ingestion).
Read how shared dependencies spread a failure in why one cloud outage can break unrelated apps, or browse the outage archive.
- Started (UTC)
- 2023-03-08 06:03 UTC
- Resolved (UTC)
- 2023-03-09 08:58 UTC
- Services affected
- All Datadog services (web app, APIs, monitors, data ingestion)
- Incident report
- datadoghq.com
- Sources
Source Link Datadog incident report https://www.datadoghq.com/blog/2023-03-08-multiregion-infrastructure-connectivity-issue/
