Web, Payments, Bytes & Ventures

WPBV Radio

Tune into tech and money.

Datadog outage, 8 March 2023

The Datadog outage of 8 March 2023: what failed, why, and what Datadog changed, from its own incident report.

Provider
Datadog
Date
Duration, hours
26.9
Cause
Software bug
Regions
US1, EU1, US3, US4, US5 regions

Datadog multi-region outage from automatic systemd update (8 March 2023). The disruption lasted about 26 hours 55 minutes, according to Datadog’s own timeline.

What happened. A legacy security update channel in Datadog’s base OS image automatically applied a systemd update across many VMs in the same 06:00-07:00 UTC window. When systemd-networkd restarted it deleted routes managed by the Cilium CNI plugin, taking tens of thousands of nodes offline across five regions and three cloud providers.

What changes. Datadog disabled the legacy update channel in all regions and said it will prove the platform can run degraded rather than fully down, plus improve customer guidance and status communication.

Questions

What caused the Datadog outage in March 2023?

A legacy security update channel in Datadog’s base OS image automatically applied a systemd update across many VMs in the same 06:00-07:00 UTC window.

Which services were affected?

All Datadog services (web app, APIs, monitors, data ingestion).

Read how shared dependencies spread a failure in why one cloud outage can break unrelated apps, or browse the outage archive.

Started (UTC)
2023-03-08 06:03 UTC
Resolved (UTC)
2023-03-09 08:58 UTC
Services affected
All Datadog services (web app, APIs, monitors, data ingestion)
Incident report
datadoghq.com
Sources
SourceLink
Datadog incident reporthttps://www.datadoghq.com/blog/2023-03-08-multiregion-infrastructure-connectivity-issue/
WPBV Radio

What are you looking for?

Search by headline, topic or keyword.