To read a post-incident report, start with the customer impact, reconstruct the timeline, then examine the actions intended to prevent a repeat. Service restoration and permanent remediation can happen on different dates.
A report may be called a postmortem, incident review or root-cause analysis. The name matters less than whether the document gives enough detail to understand the failure. Google’s SRE book describes a postmortem as a record covering impact, recovery, causes and preventive work. Its workbook also examines how vague actions, missing context and unclear ownership weaken that record.
Start with what customers experienced
Look for the affected service, region and operation. A dashboard being available does not establish that payments, downloads or API calls worked. Read the scope of each claim before applying it to your own use.
Pay particular attention to the denominator. “A percentage of requests failed” and “a percentage of users were affected” answer different questions. A user can make several requests, while a failure in one endpoint can coexist with successful traffic elsewhere. Record the measure the provider actually uses rather than converting it into a broader claim.
Separate availability, latency and data integrity. Slow responses, rejected requests and lost data have different consequences for a customer deciding what to reconcile or retry. If the report names a specific operation, start with that operation in your own records.
Rebuild the timeline
Place the start of impact, detection, mitigation and recovery in order. These moments need not coincide.
| Moment | Question to ask |
|---|---|
| Impact began | When did customer operations first fail or degrade? |
| Detection | When did monitoring or a report identify the problem? |
| Mitigation | What reduced the impact while investigation continued? |
| Recovery | What evidence showed the affected service was working again? |
| Permanent work | Which changes were completed after restoration? |
For example, rolling back a release might restore service before the team understands why the release escaped testing. That is an illustrative sequence, not a claim about a particular outage.
Check the timezone before comparing the report with your logs. Preserve the stated time format and any distinction between a provider’s incident window and your own observations. A status update’s publication time is also different from the time of the event it describes.
Our guide to a status page during an outage explains why a public update may arrive after the first customer failures.
Separate the trigger from the conditions
A deployment, maintenance action or traffic spike can trigger a failure. The underlying conditions explain why the system failed to contain it.
An illustrative example is a configuration change accepted by several services at once. The change starts the incident, but the useful questions concern validation, rollout scope, monitoring and rollback. Replacing “configuration change” with “human error” does not answer those questions.
Google’s blameless approach focuses on contributing causes and the information available to the people involved. For a reader, that offers a useful test: does the report explain how the system allowed the failure, or does it mainly identify who touched it?
Avoid treating a confident singular “root cause” label as proof that every contributing factor has been examined. Follow the explanation from trigger to customer effect and check whether each link is supported in the report.
Read the action list as a separate document
Restoring the service proves that recovery work succeeded for that incident. It does not establish that recurrence is impossible.
The SRE workbook stresses concrete actions, ownership, prioritisation and tracking. Apply those criteria to each commitment. “Improve reliability” gives a reader little to evaluate; a named change to a specified control creates a clearer test.
As an illustrative reading exercise, compare these two commitments:
| Commitment | What a reader can assess |
|---|---|
| Improve deployment safety | The intended direction, with little detail |
| Add a validation test for the rejected configuration | The specific failure condition the action addresses |
The second statement still needs evidence of completion. A future commitment and a deployed control should remain separate in your notes. Follow any subsequent update the provider links, rather than assuming every listed item is already finished.
Use the report to guide your own checks
Write a short account-specific note: which operation you used, the relevant interval and the records you need to reconcile. Keep the provider’s documented impact separate from what your own logs show.
A report can explain the infrastructure failure without resolving every customer-side consequence. For a payment or data workflow, a useful next step is to match the affected operation to its request IDs, results and recovery procedure.




