Slack outage, 4 January 2021
The Slack outage of 4 January 2021: what failed, why, and what Slack changed, from its own incident report.
- Provider
- Slack
- Date
- Cause
- Third-party dependency
- Regions
- Global
Slack outage from AWS Transit Gateway saturation (4 January 2021).
What happened. On the first workday after the holidays, one of the AWS Transit Gateways linking Slack’s VPCs became overloaded and dropped packets because it did not scale fast enough. Packet loss saturated Slack’s web tier, and autoscaling and provisioning problems made recovery harder.
What changes. AWS is reviewing TGW scaling algorithms; Slack will request preemptive TGW upscaling after holidays and fix provisioning, health-check and autoscaling behaviour.
Questions
What caused the Slack outage in January 2021?
On the first workday after the holidays, one of the AWS Transit Gateways linking Slack’s VPCs became overloaded and dropped packets because it did not scale fast enough.
Which services were affected?
Slack (messaging, web tier), Slack internal dashboards/alerting.
Read how shared dependencies spread a failure in why one cloud outage can break unrelated apps, or browse the outage archive.
- Resolved (UTC)
- 2021-01-04 18:40 UTC
- Services affected
- Slack (messaging, web tier), Slack internal dashboards/alerting
- Incident report
- slack.engineering
- Sources
Source Link Slack incident report https://slack.engineering/slacks-outage-on-january-4th-2021/
