Postmortem template
A blameless incident postmortem template covering impact, root cause, lessons learned and owned action items, with a review checklist.
An outage is over when the service is back. The incident isn't finished until the team understands why it happened and has changed something so it is less likely to happen again. This incident postmortem template gives that review a fixed structure, so the learning ends up in a document with owned action items rather than scattered across an incident channel.
This postmortem template covers impact, the key moments of the timeline, root cause and trigger, what slowed detection and recovery, lessons learned and action items. It ends with a review checklist, so a postmortem only closes once it is complete and shared.
Its structure follows how Google, Amazon Web Services, Microsoft Azure and GitLab write and review postmortems. See the sources for how each of them approaches it.
The template
Postmortem
Write blamelessly: describe what the system allowed, not who slipped. Write "We didn't have monitoring for this condition", not "Person X made a mistake".
Postmortem details
- Title
- Incident ID
- Incident date
- Postmortem date
- Severity
- Owner
- Contributors
- Reviewer
- Status
- Public postmortem
- Incident report
- Root cause analysis
Summary
Write this last, as if it were going straight to your CEO by email: who was affected, for how long, the cause, how it was mitigated and what will prevent it. It should make sense without the rest of the document.
Impact
Use real numbers. If a figure is an estimate, say how it was estimated.
- Users or customers affected
- Regions
- Duration
- Business impact
- SLO or error budget impact
- Team impact
Timeline highlights
Only the key moments, in UTC. Start at the trigger, not at the page. The detailed timeline stays in the incident report.
| Time (UTC) | Event | Phase |
|---|---|---|
Root cause and trigger
The trigger set the incident off; the root cause is the condition that let it become an incident. Link the RCA for the full analysis.
- Root cause
- Trigger
- Root cause type
Detection, diagnosis and mitigation
How could each phase have been faster? Slow detection or recovery deserves its own action item.
- How did we learn of the impact?
- How could we have detected it in half the time?
- What made the cause hard to find?
- How could we have diagnosed it in half the time?
- How did we confirm the service was really back?
- How could we have mitigated it in half the time?
Lessons learned
What went well
What went wrong
Where we got lucky
Action items
One owner per action, a verifiable end state and a tracking link. Include at least one Major or Critical action, unless stakeholders accept the risk of recurrence.
| Action | Type | Priority | Owner | Due date | Tracking link | Status |
|---|---|---|---|---|---|---|
Review checklist
Supporting information
- Dashboards
- Logs
- Incident chat
- Related incidents
What is a postmortem?
A postmortem is a written review of an incident, completed after service is restored. Google's SRE book defines it as a record of the incident, its impact, the actions taken to mitigate it, its root causes and the follow-up actions that prevent it from happening again. A postmortem template keeps that record consistent from one incident to the next.
Postmortems are blameless. They assume everyone acted in good faith with the information they had, and they ask what the system allowed rather than who made a mistake. Teams that expect blame stop reporting problems, and the review loses the details it depends on.
The same document goes by other names: post-incident review at Microsoft and GitLab, and correction of error (COE) at Amazon.
When should you write a postmortem?
Agree on the criteria before the next incident, so nobody has to debate them afterwards. Google's SRE book lists common triggers:
- User-visible downtime or degradation beyond a set threshold
- Data loss of any kind
- On-call intervention, such as a rollback or rerouting traffic
- Resolution time above a set threshold
- A monitoring failure, where the incident was found manually
Any stakeholder can also ask for one. AWS adds that an incident doesn't require an outage: a near miss, or a system behaving unexpectedly while still working, is worth reviewing too. GitLab requires a review for every S1 and S2 incident and lets anyone request one for lower severities.
Write it while the details are fresh. Google's published example went out less than a week after the incident closed, and Azure publishes its final reviews generally within 14 days.
What should a postmortem include?
Postmortem details
Give the postmortem one owner, not a committee. Google lists four owners as a sign of a weak postmortem: contributors help, one person drives it to completion. Name a reviewer as well. At GitLab, a peer "Bar Raiser" approves every review before it can close. Link the incident report and the RCA instead of repeating them.
Summary
Write it last. AWS's guidance is to write the summary as if it were going to your company's main stakeholder by email: who was affected, for how long, how it was mitigated and what will prevent it. It should make sense on its own.
Impact
Use numbers: failed requests, affected customers, minutes of impact, error budget used. Google splits impact into user, revenue and team impact, because the hours responders spent are part of the cost. If a figure is an estimate, say how it was estimated.
Timeline highlights
Record only the moments that matter: trigger, detection, escalation, mitigation and resolution. Start at the trigger, such as a deployment or a traffic spike, not at the page, and use UTC. The minute-by-minute timeline stays in the incident report.
Root cause and trigger
Keep the two apart. The trigger set the incident off; the root cause is the condition that let it become an incident. In Google's example postmortem, the trigger was a sudden traffic increase and the root cause a latent resource leak. The root cause type tells you what kind of fix to look for: GitLab expects an incident caused by a code change to produce actions that stop similar defects from escaping, not just a patch. For the full analysis, use the root cause analysis template.
Detection, diagnosis and mitigation
This section asks how the response could have been faster. For each phase, AWS's correction-of-error process asks how you could cut the time in half. GitLab treats slow detection or recovery as a contributing cause that needs its own corrective action.
Lessons learned
Three lists, from Google's template: what went well, what went wrong and where you got lucky. The last one captures near misses, the things that limited the damage by chance and won't be there next time.
Action items
This is the part that changes something. Each action has one owner, a priority, a due date and a link to the ticket that tracks it. Make the end state verifiable ("alert when connection-pool use exceeds 80%", not "improve monitoring") and cover prevention and detection, not only repair. Google's postmortem checklist asks for at least one high-priority action, or explicit agreement from stakeholders that the risk of recurrence is accepted.
Review checklist
The checklist is based on Google's postmortem checklist and GitLab's completion criteria. A postmortem is done when impact is assessed, every action is owned and tracked, blameful language is gone, the reviewer has approved it and it has been shared.
Supporting information
Link dashboards, logs and the incident chat rather than pasting them in. List related incidents too, including similar ones that aren't exact repeats. A pattern across several incidents is often a bigger finding than any single postmortem.
The following example continues the INC-2026-017 incident used in the incident report and root cause analysis templates.
Title: Database connection pool exhaustion caused elevated API errors Incident: INC-2026-017, SEV-2, September 3, 2026 Owner: Platform Engineering lead Reviewer: Staff engineer, Database Team Status: Final, action items open Public postmortem: Not required
Summary
Between 14:03 and 14:31 UTC, about 18% of public API requests failed. Long-running analytical queries held database connections open until the shared pool was exhausted. Responders terminated the queries and raised connection capacity. We are enforcing a query timeout on the analytics workload, moving it to its own connection pool and adding a saturation alert.
Impact
18% of API requests failed for 28 minutes. Dashboards and the status page were unaffected, and no data was lost. Four engineers spent about an hour each on the response.
Timeline highlights
| Time (UTC) | Event | Phase |
|---|---|---|
| 13:58 | Scheduled analytics job starts long-running queries | Trigger |
| 14:07 | API error-rate alert fires | Detected |
| 14:12 | Incident declared, Database Team paged | Escalated |
| 14:31 | Queries terminated and capacity raised; error rates normal | Mitigated |
| 14:45 | Incident resolved | Resolved |
Root cause and trigger
Root cause: No query timeout was enforced on the connection pool shared by analytics and customer-facing traffic. Trigger: A scheduled analytics job. Type: Capacity
Detection, diagnosis and mitigation
The error-rate alert fired four minutes after impact began. A connection-pool saturation alert would have fired before any request failed. Connecting the errors to the pool took eight minutes because the database runbook had no step for connection exhaustion.
Lessons learned
- Went well: the error-rate alert fired within four minutes, and the incident was declared within five more.
- Went wrong: analytics and API traffic shared one pool, and nothing alerted on its saturation.
- Got lucky: the job ran before peak US traffic.
Action items
| Action | Type | Priority | Owner | Due date |
|---|---|---|---|---|
| Enforce a query timeout on the analytics workload | Prevent | Critical | Data Engineering | September 12 |
| Alert when connection-pool use exceeds 80% | Detect | Major | Platform Engineering | September 10 |
| Move analytical queries to an isolated connection pool | Prevent | Major | Database Team | September 30 |
| Add connection-exhaustion diagnostics to the database runbook | Mitigate | Medium | Site Reliability Engineering | September 15 |
Postmortem vs. incident report vs. root cause analysis
An incident usually produces all three, in this order.
| Document | Answers | When it's written |
|---|---|---|
| Incident report | What happened, what broke, what did we do? | During or right after the incident |
| Root cause analysis | What specifically caused it, and what let it happen? | Once the investigation has progressed |
| Postmortem | What did we learn, and what will we change? | Within a week or two of resolution |
The postmortem builds on the other two rather than reconstructing the incident. It takes the facts from the incident report and the cause from the root cause analysis, and turns them into lessons and owned action items. Low-severity incidents may only need the report.
Writing a public postmortem
Some incidents need a version for customers. GitLab publishes a public root cause analysis for every S1 incident, and Google recommends sharing postmortems as widely as possible, "perhaps even with your customers". A public postmortem is a separate document written from the internal one, not the internal one with names removed.
Azure's post-incident reviews answer the same six questions every time:
- What happened?
- What went wrong and why?
- How did we respond?
- How are we making incidents like this less likely or less impactful?
- How can customers make incidents like this less impactful?
- How can we make our incident communications more useful?
Keep the specifics that rebuild trust: times, scope, the cause in plain language and the changes you are making. Leave out internal names, provisional theories and anything that identifies an individual customer or user. Publish it where customers looked during the incident, such as your status page, and link it from the incident's final update.
Best practices
-
Describe the system, not the person. GitLab's guidance gives examples of the difference:
Write Instead of "The deployment process didn't catch this issue" "Person X made a mistake" "We didn't have monitoring for this condition" "They didn't follow the process" "The runbook didn't have a step for this scenario" "The on-call engineer should have known" -
Publish within a week or two. Google's example of a weak postmortem was published four months after the incident. The outage in that case study did happen again, and a late write-up loses the details that would have helped.
-
Keep the language factual. Dramatic descriptions distract from the findings and make people defensive. Back every claim with data.
-
Involve everyone who took part. A postmortem written by one team alone tends to miss the contributing factors other teams saw.
-
Share it widely. The value of a postmortem grows with the number of people who learn from it. Default to access for the whole engineering organization.
-
Track actions to closure. Every action belongs in your ticket system, and someone checks progress. A postmortem with open action items is not finished.
-
Watch for repeats. If incidents mirror earlier ones, ask whether actions take too long to close or whether the right actions were captured at all.
What is the difference between a postmortem and a post-incident review?
Nothing substantive. Both are a written review of an incident after service is restored, covering impact, cause, lessons and follow-up actions. Microsoft and GitLab call it a post-incident review, Amazon calls it a correction of error, and many teams say postmortem, also written post mortem. A post mortem template and a post-incident review template cover the same ground.
What is a blameless postmortem?
A blameless postmortem looks for the conditions in the system and process that allowed an incident, instead of the person who made a change. It assumes everyone acted in good faith with the information they had. People then report problems openly, which gives the review the details it needs.
Who should write the postmortem?
One owner, usually from the team that owns the affected service, with contributions from everyone who took part in the response. GitLab makes the service-owning team responsible for its incident reviews, and a separate reviewer approves the result.
How soon after an incident should you write a postmortem?
Within a week or two of resolution, while details are fresh. Google's published example went out less than a week after the incident closed, and Azure publishes final post-incident reviews generally within 14 days.
Does every incident need a postmortem?
No. Define criteria in advance, such as user-visible downtime above a threshold, any data loss, on-call intervention or a monitoring failure. Any stakeholder can also request one. Minor incidents may only need an incident report.
Build postmortems into your process
Decide before the next incident which severities require a postmortem, who owns it, who reviews it, where it is stored and how its action items are tracked. When that is settled in advance, writing the postmortem is routine rather than a debate.
Used for every qualifying incident, the same postmortem template turns your postmortems into a searchable record of how your systems fail and what you changed in response. That record is what makes the next incident shorter.
Sources
The structure of this template follows how these organizations write and review postmortems. They differ mostly in when a review is required, how quickly it is published and who approves it.
| Organization | How their postmortem process differs |
|---|---|
| Google: SRE Workbook, Postmortem Culture | Postmortems follow objective triggers listed in the SRE book, and any stakeholder can ask for one. Senior engineers review each draft. Action items have one owner, a priority and a tracking bug, and Google's checklist asks for at least one high-priority action. |
| AWS Well-Architected: Perform post-incident analysis | Amazon's correction of error process covers any significant event, including near misses. It uses five whys and asks how detection, diagnosis and mitigation could each take half the time. Every action needs a priority, an owner and a due date. |
| Microsoft Azure: Post Incident Reviews | Reviews of customer-impacting incidents are published on the Azure status history and kept for five years. Each answers the same six questions, including how customers can reduce their own impact. The final review follows generally within 14 days. |
| GitLab: Incident Review | Every S1 and S2 incident gets a review, owned by the team that owns the service. A peer Bar Raiser must approve it, every corrective action is assigned to a team before it closes, and S1 incidents get a public RCA. The on-call handbook adds the blameless-language examples. |