Availability zones explained in a provider’s documentation are infrastructure boundaries. An application can run in a cloud region containing several availability zones and still fail when one zone goes down. The region’s infrastructure offers fault-isolation boundaries; the workload must actually use them.
AWS’s fault-isolation guidance describes an availability zone as one or more discrete data centers with separate power, networking and connectivity inside a region. Its resiliency documentation then separates those provider responsibilities from how a customer designs an application.
A region and a zone describe different boundaries
AWS says each region is isolated and contains multiple availability zones. Zones within a region are physically separated while connected through low-latency, redundant networking.
That arrangement supports replication across zones without putting every component in the same facility. AWS describes independent power and cooling infrastructure and deployments separated in time to reduce correlated failures.
The distinction matters when reading an incident report. A problem in one zone is different from a regional service problem or a customer’s deployment affecting every copy of an application.
Multi-zone infrastructure is not automatic application recovery
AWS’s shared-responsibility model for resiliency assigns the provider responsibility for the infrastructure and the customer responsibility for the workload design, configuration and operation.
If every application instance is in one zone, the existence of other zones does not move those instances on its own. If data is available only in the failed zone, a running web server elsewhere may have nothing useful to serve.
A review therefore needs to follow dependencies: traffic routing, application capacity, storage, databases and the services used to recover. The label “multi-AZ” on one component does not describe the whole path.
Follow a request through the failure
Consider a hypothetical checkout with application instances in two zones and a database reachable only in one. Losing the database’s zone can stop checkout even while the second application’s instances remain healthy.
The example is not a report of an AWS incident. It illustrates why redundancy should be assessed at the operation a customer needs to finish rather than counted as a number of servers.
| Component | Recovery question |
|---|---|
| Traffic routing | Can requests reach the healthy instances? |
| Application capacity | Can the remaining deployment handle demand? |
| Data | Is the required state available outside the failed zone? |
| Dependencies | Do authentication and other services remain reachable? |
| Recovery controls | Can the team operate them during the disruption? |
A deployment can avoid a single facility’s failure while remaining vulnerable to an application change rolled out everywhere. Independent infrastructure and controlled software releases address different causes.
Another region adds more separation and more work
A second region changes the fault boundary again. It also requires decisions about data replication, traffic movement, recovery timing and service configuration.
For some workloads, several zones in one region match the required recovery objective. Others need a separate regional recovery plan. The provider’s number of regions does not establish which design a workload uses.
AWS’s multi-location deployment guidance places these decisions in the context of workload requirements. State the desired recovery behavior before treating more locations as a complete plan.
Read outage language carefully
A provider can describe an isolated infrastructure failure while users experience a broader outage through a shared dependency. Both descriptions can be consistent if the application concentrates a critical part of its request path.
For the public communication side, see why a status page can lag an outage. The recovery architecture and the incident message are separate evidence.




