Incident Postmortem
beta · Family process-docs · Phase iterate · Sizes lean, full · ~3,350 tokens
The document a team writes after an incident to explain why it happened and what will change so it does not happen the same way again, triggered by an agreed-in-advance criterion rather than improvised after the fact. It is judged by whether its Action Items land as owned, tracked tickets in the team’s existing backlog, not by whether the write-up reads well.
The short card. Why the document is shaped this way, and the argument behind every rule here, is in
incident-postmortem_companion.md. A fully worked instance is
incident-postmortem_example.md.
When to use
Section titled “When to use”- A specific criterion your team published in advance actually fired: user-visible downtime past a
threshold, data loss, an on-call engineer having to intervene, a resolution time past a threshold, or a
monitoring gap that meant the incident was found by hand rather than by an alert. The canon’s own list of
triggers is in
incident-postmortem_companion.mdsection 1; write your own team’s version of it before you need this document, not while you are filling one in. - One specific event failed, and you need to explain why, not how a period of work went in general.
- The event is over, or at least stable enough to analyze; a postmortem is written after, not during.
- Follow-up work will need to be tracked somewhere your team already tracks work: the product backlog, a risk register, or a RAID log. If nothing here is going to produce a ticket, this document will not do its job. See “Leaving action items only in this document” under anti-patterns below.
When NOT to use
Section titled “When NOT to use”- The trigger is a date on the calendar, not an event. If what happened is “two weeks passed,” reach
for
sprint-retrospective-notesinstead. That document looks back on a period, on a cadence, at how the team worked; this one is event-triggered, about one specific thing that failed. Running a retrospective on an incident produces a blameless discussion of a thing that needed a causal analysis; running a postmortem on an ordinary sprint pathologizes normal work. Neither substitutes for the other. Seeincident-postmortem_companion.mdsection 8. - You only need a live record of what happened, not an analysis of why. A postmortem is written after the fact, once there is time to look at logs, code changes, and decisions; a live incident record kept during the event is a different artifact. One captures what happened; this one explains why and what changes so it does not happen the same way again.
- You need a facilitator-led discussion, not a document. The “after-action review” label sometimes attached to this document type in other catalogs literally means a live, facilitated conversation running something like 30 minutes to 2 hours about a single training event, not a written artifact. The spirit matches; the format does not.
- You are inside an ITIL or ITSM practice tracking a Problem Record. Expect this document to feed that record, not replace it. A Problem Record documents a problem’s full history from detection to closure; a postmortem is the analysis of one incident that record may reference.
- Nothing changed, and nothing needs to. If the honest content of every section would be “this was normal operation,” you do not have an incident that cleared your team’s own trigger criteria. Writing one anyway trains the team to treat postmortems as routine paperwork rather than as a signal that something worth investigating actually happened.
Pick a variant
Section titled “Pick a variant”Lean (five sections) is Summary, Impact, Timeline, Root Causes, and Action Items. It is the sections attested across the widest span of published practice, and it is what a reader still needs even from a one-off, freeform incident report with no fixed template at all.
Full (nine sections) adds Detection, Trigger, Resolution, and Lessons Learned, inserted in place rather than appended, so lean stays a strict ordered subset of full. Notice what the four additions have in common: none of them is content a reader needs to just know what happened and what is being done about it. Detection and Trigger are about how the team found out and why this counted; Resolution and Lessons Learned separate what already happened from what the team now believes differently. Move to full when at least one of these is true:
- the incident is severe enough, cross-team enough, or novel enough that skipping straight from Impact to Root Causes, or from Root Causes to Action Items, would leave a reader unable to tell how it was found, which of your team’s own criteria made it a postmortem at all, what specifically ended it, or what changed about how the team believes it works;
- more than one team could plausibly need to understand the trigger and the resolution, so naming both explicitly matters more than it would for a single-team incident;
- the incident revealed something the team wants to remember beyond the specific fix, which belongs in Lessons Learned rather than folded into Root Causes.
The two-size packaging is this bundle’s own decision, and it is labeled as such. No published postmortem
template this bundle’s research read ships two sizes; each vendor ships exactly one. What the research does
support is that depth genuinely varies in practice: one vendor’s product ships configurable templates keyed
to incident severity, and a real published postmortem from a named organization used no fixed template
structure at all. This library packages that real variation as two sizes; it is not a discovered industry
standard. See incident-postmortem_companion.md section 4.
Quality rubric (self-grade)
Section titled “Quality rubric (self-grade)”Score each 0, 1 or 2. Full below 13 out of 18 has produced exactly what the process-docs family contract
warns against: a document that records feelings or a timeline and commits nobody to anything. Lean is scored
on five of these rows only (see the scope table below), and below 7 of 10 the same failure applies to the
smaller shape.
| # | Criterion | 0 | 1 | 2 |
|---|---|---|---|---|
| 1 | Summary stands alone | No cause, no duration, no resolution state; a reader who stops here learns nothing | Some of cause, duration, or resolution state is present, but not all three | A reader who stops after this section alone knows what happened, roughly how bad it was, and whether it is over |
| 2 | Impact is measured | A category with no number (“some users,” “a while”) | A number is given for one dimension (duration or magnitude) but not both | Every impact row names an affected system or population, a duration, and a magnitude someone could check against a dashboard or a support queue |
| 3 | Timeline is checkable | Vague times (“that afternoon”) or several events bundled into one line | Timestamps are present but could not be checked against a log, alert, or chat transcript without asking the author | Every entry has a timestamp precise enough to verify against a real record, and gaps in the timeline are left visible rather than smoothed over |
| 4 | Root causes are evidenced | One vague category (“human error”) or a person named as the cause | More than one cause is named, but at least one has no evidence behind it | Every listed cause is a specific, checkable condition with evidence (a log line, a config diff, a metric), and none of them is a person’s name |
| 5 | Detection reveals the gap (full) | “We noticed a problem,” no named signal, no delay stated | A signal is named, but the delay between failure and detection is not | A named signal (alert, page, customer report) and the delay between the failure starting and someone finding out, stated as a number someone could check |
| 6 | Trigger is your own (full) | No named criterion, or a borrowed list (the canon’s own example list, restated as though it were the team’s own) | A criterion is named, but it is not one the team could point to in a document that predates the incident | Names the specific, previously published criterion that fired, and where a reader could go check that it was in fact published before this event |
| 7 | Resolution is distinct (full) | “We fixed it,” no action named, no timestamp | An action is named, but it is not clear whether it already happened or is still planned | Names the specific action that ended the incident, who took it, and when, cross-checkable against the Timeline; work still to come lives in Action Items, not here |
| 8 | Lessons learned change something (full) | “We should be more careful,” indistinguishable from what any postmortem could say | Restates the Root Cause in different words rather than stating a changed belief | States what the team now believes differently about how it works, distinct from the specific fix, in language specific enough to apply to a future, different incident |
| 9 | Actions are tracked | A bulleted intention with no owner and no ticket | An owner or a ticket is present, not both, or the ticket has no visible status | Every action names one owner and a ticket in the team’s actual tracker, with a status a reader could check, not a wish written down and left in this document |
Which rows apply to what.
| Document | Rows | Maximum | Score against |
|---|---|---|---|
| full | all 9 | 18 | 13 |
| lean | 1-4 and 9 | 10 | 7 |
Rows 5, 6, 7 and 8 are scored only against full. Lean ships none of Detection, Trigger, Resolution, or Lessons Learned, so grading it on those rows would penalize the choice of variant rather than the quality of the document.
The test behind every cell above: could someone satisfy it without improving the document? A row that counted the number of root causes, timeline entries, or action items would reward padding. Every cell instead asks whether a specific piece of checkable evidence exists, and whether a reader who was not in the room could verify it.
Named anti-patterns (the usual wrecks)
Section titled “Named anti-patterns (the usual wrecks)”- Writing the Summary as a preview of Root Causes. The Summary is a synopsis for someone who may never read past it, not a teaser for the analysis below. If a reader who stops at Summary cannot say what happened, how bad it was, and whether it is over, the section has not done its job.
- Naming a person as a root cause. A root cause can never be a person. A postmortem that lands on an individual has produced blame, not analysis, and directly contradicts the blameless framing this document type is named for.
- Restating a borrowed list of trigger criteria as though it were your team’s own. Naming a trigger your team never actually published before the incident is inventing a rule retroactively, which defeats the entire point of agreeing criteria in advance rather than arguing them during one.
- Leaving action items only inside this document. The moment a postmortem is published, every action item needs a ticket in the tracker the team already uses. A postmortem that ends in a bulleted list at the bottom of the page rather than tickets in a real tracker has produced a wish list, not follow-up work.
- Describing a fix as already resolving the incident when the underlying condition is still present and only masked. If the real fix has not shipped yet, say so, and put it in Action Items rather than Resolution.
- Restating the Root Cause in different words under Lessons Learned. A lesson is what the team now believes about how it works; a root cause is what happened this time. Repeating one under the other’s heading wastes the section that is supposed to generalize past this specific incident.
- Citing an unmeasured percentage as though it were a finding. Every specific MTTR-reduction or recurrence-reduction figure this bundle’s research chased traced to a headline with no method, a single customer testimonial, or a claim absent from the very document it was attributed to. State the value case honestly as an assertion, not as a number nothing behind it supports.
- Running a retrospective on an incident, or a postmortem on a sprint. These are the two members of the
process-docsfamily, and they exist to be told apart, not blended. A retrospective on an incident produces a blameless discussion of something that needed causal analysis; a postmortem on a sprint pathologizes ordinary work. If the trigger is a date, it is the other document.
Pairing with your process
Section titled “Pairing with your process”This bundle ships in the process-docs family alongside sprint-retrospective-notes, and the distinction
between them is this family’s central teaching point rather than an incidental note: a retrospective is
cadence-triggered and looks back on a period, at how the team worked; a postmortem is event-triggered and
looks back on one specific thing that failed and why. Before you start either document, name which trigger
you actually have. Once you have written this one, make sure every Action Items row lands somewhere your
team already tracks work: the product backlog, a risk register, or a RAID log. An action recorded and never
done is a real failure mode for this document type, distinct from the analysis itself being wrong, and it
is the one the family contract names explicitly: owned actions with a place they are tracked, not a list of
observations.
The artifacts
Section titled “The artifacts”incident-postmortem_template-lean.md · ~3,350 tokens
---title: "{{postmortem_title}}"incident_id: "{{incident_id}}"status: "{{status}}"author: "{{author}}"incident_date: "{{incident_date}}"postmortem_date: "{{date}}"doc_type: incident-postmortemsize: leansource_template: incident-postmortemsource_template_version: 0.1.0---
<!--LEAN INCIDENT POSTMORTEM. The five sections attested across the widest span of the published corpus thisbundle's research read: what happened, who and what it hurt, the order it happened in, why it happened, andwhat changes because of it. It is what a reader would still need even from a one-off, freeform report withno fixed template at all. To carry Detection, Trigger, Resolution, and Lessons Learned (seeincident-postmortem_template-full.md), ADD sections; never rename or reorder the ones below, because thefull variant is a strict superset of this one.
A POSTMORTEM IS TRIGGERED BY A CRITERION YOUR TEAM PUBLISHED IN ADVANCE, NOT BY A BAD DAY. Google's SRE bookstates an explicit list of what should prompt one: "Common postmortem triggers include: User-visibledowntime or degradation beyond a certain threshold; Data loss of any kind; On-call engineer intervention(release rollback, rerouting of traffic, etc.); A resolution time above some threshold; A monitoring failure(which usually implies manual incident discovery)," and holds that a team should agree its own version ofthat list before an incident rather than argue it during one. See incident-postmortem_companion.md section 1.
THERE IS NO SINGLE CANONICAL SECTION SET FOR THIS DOCUMENT. The canon's own worked example and its ownfollow-up volume's worked examples use materially different headings from each other. The five sectionsbelow are the ones attested across the widest span of the published corpus this bundle's research read, not"the" template. See incident-postmortem_companion.md section 4.
WHAT AN INCIDENT POSTMORTEM IS, AND IS NOTIt is the document a team writes after an incident to explain why it happened and what will change so itdoes not happen the same way again. It is NOT a sprint retrospective (that looks back on a period, on acadence, at how the team worked; a postmortem is event-triggered, about one specific thing that failed). Itis NOT the incident report: by one vendor's account, "if the incident report answers the question whathappened, the postmortem answers why it happened and what will prevent this from happening again," thoughthat distinction is not standards-tier, and this template treats it as one useful framing rather thansettled fact. It is NOT an After Action Review in the literal sense of that catalog alias: the US Army's ownguide describes a facilitator-led discussion running 30 minutes to 2 hours about a single training event,not a written document. See incident-postmortem_companion.md section 8.
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into incident-postmortem_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.2. Replace each {{placeholder}} with your content. Fill Summary and Impact first; the rest depends on knowing what actually happened before you analyze why.3. If a section does not apply, write "N/A" and one line of why, rather than deleting it silently.4. Before you share it: self-grade against incident-postmortem_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{postmortem_title}}
## Summary
<!-- WHAT A short synopsis of what happened, written before the analysis that follows it. Read first, even though it is usually the last thing written. WHY Attested by the canon's own worked example and three of the five published templates this bundle's research read in full. A reader who stops here should still know what happened, how bad it was, and whether it is over. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Summary). ASK What happened, in two or three sentences someone outside the incident could follow? What is the current status: resolved, mitigated, or still ongoing? GOOD "On July 14, a config change to the Saved Views API's rate limiter caused it to reject 60 percent of legitimate requests for 42 minutes before the change was rolled back. Fully resolved; no data was lost." WEAK "The Saved Views API had some issues." (no cause, no duration, no resolution state; a reader learns nothing from the section built to be read alone) TRAP Writing the Summary as a preview of Root Causes. It is a synopsis for someone who may never read past it, not a teaser for the analysis below. -->
{{summary}}
## Impact
<!-- WHAT Who and what was affected, and how badly: systems, users, duration, and any data loss. WHY The most broadly attested section besides Timeline: the canon's own example and four of the five published templates read carry it. A postmortem that skips straight from Summary to Root Causes leaves the reader unable to judge whether the analysis that follows was worth the time it took to write. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Impact). ASK What system or user population was affected? For how long? What was the magnitude (percentage of requests failed, records affected, revenue)? Was any data lost? PRIORITY List the most severe or most visible impact first. State a real number, not a category; "some users" is not measurable. ROW HINT A good row names one affected system or population, a duration, and a magnitude a reader could check against a dashboard or a support queue. A weak row is a category with no number. GOOD | Saved Views API | 42 minutes | 60 percent of requests returned 429; no data loss | WEAK | The API | A while | Users were affected | TRAP Reporting only technical impact and skipping user or business impact, or the reverse. "The API was down" hides whether anyone outside engineering noticed. -->
| Affected system or population | Duration | Magnitude ||---|---|---|| {{impact_affected}} | {{impact_duration}} | {{impact_magnitude}} |
## Timeline
<!-- WHAT A chronological, timestamped account of what happened, in order. WHY The single most attested title in this bundle's research, appearing in the canon's own example and four of the five published templates read. It is also this document type's sharpest lesson about received wisdom: the SRE book's own postmortem chapter, the source most people would name if asked where "timeline" comes from, contains zero occurrences of the word in its prose. It exists only as a heading in the separately linked worked example. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Timeline) and section 1. ASK In order, what happened, and when? Include the failure's start, detection, escalation, any mitigation attempts, and resolution, each with a timestamp. PRIORITY List every entry in chronological order, earliest first. Do not skip the boring entries; a gap in the timeline is itself information about what nobody was watching. ROW HINT A good row has a timestamp precise enough to check against a log or a chat transcript, and one clear event. A weak row has a vague time ("morning") or bundles several events into one line. GOOD | 14:02 UTC | Rate limiter config change deployed to production | WEAK | Sometime that afternoon | Things started going wrong | TRAP Writing the timeline from memory days later without checking it against logs, alerts, or chat history. A timeline nobody can verify is a story, not a record. -->
| Time (UTC) | Event ||---|---|| {{timeline_time}} | {{timeline_event}} |
## Root Causes
<!-- WHAT The contributing causes, plural by design, each with the evidence behind it. WHY Attested by the canon's own example, and, under varying names (GitLab's "Root Cause Analysis," Elastic's singular "Root Cause," Atlassian's split "Root cause identification" and "Root cause"), close to universal: present in four of the five published templates read. This is also the one section this bundle ships over a live, unresolved argument. A named, cross-citing community (Cook, Allspaw, Hollnagel, Woods, Dekker, Leveson) holds there is no single root cause in a complex system at all; the clearest named defence found still concedes, "sometimes you won't find a root cause. It happens." This template does not settle that argument; it asks for causes, plural, and evidence for each. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Root Causes) and section 6. ASK What conditions, each necessary but not sufficient alone, combined to produce this incident? What evidence (a log line, a config diff, a metric) supports each one? Have you checked whether any candidate cause is actually a person's name? PRIORITY List every contributing cause you have evidence for, not just the first one found. A single-row Root Causes table is usually a sign the investigation stopped too early. ROW HINT A good row names one specific, checkable condition and the evidence behind it. A weak row names a broad category ("human error," "process failure") or a person. GOOD | The rate limiter's config validation did not reject an out-of-range threshold before deploy | Deploy log shows the config passed CI with no validation step for that field | WEAK | Someone made a mistake | Just experience | TRAP Naming a person as a cause. GitLab's own guidance states the rule plainly: "A root cause can **never be a person**." A postmortem that lands on an individual has produced blame, not analysis, and directly contradicts the blameless framing this document type is named for. -->
| Contributing cause | Evidence ||---|---|| {{root_cause}} | {{root_cause_evidence}} |
## Action Items
<!-- WHAT Concrete, owned, tracked follow-up work. WHY Attested by the canon's own example and two published templates read, but the name itself is a pick: published templates write it three different ways and never converge: "Action Items" (two vendor templates), "Corrective actions" (two others), and "Follow-up actions" (a fifth). This bundle uses the canon's own word and says plainly that it is a pick, not a convergence. The canon's own worked example shows what a filled row looks like: "Plug file descriptor leak in search ranking subsystem | prevent | agoogler | Bug 5554825 DONE," a specific owner and a specific tracked bug, not a bulleted wish. The canon's own row combines the ticket ID and status in one field; this template splits them into separate columns for clarity, which is this template's own formatting choice. Two independent practitioner sources agree on where these rows live once the postmortem is finished: "The postmortem document is where action items are born. It is not where they should live," and a second, independent source says nearly the same thing. Link the ticket; do not leave the work only here. Neither source names a risk register or a RAID log as a destination; this library's own family contract offers those alongside the product backlog, which is this library's own convention rather than received postmortem practice. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Action Items). ASK What specific, owned action follows from this incident? Is there already a ticket for it in the tracker your team actually uses? Who owns it, and is the status current? PRIORITY List every action with an owner and a ticket link, prevention items before general cleanup. An action with no owner is not an action item, it is a wish. ROW HINT A good row names one concrete action, its type (the canon's own quoted row uses "prevent"; mitigate, detect, and process are reasonable values this bundle's research did not see quoted), a single named owner, and a linked ticket with a current status. A weak row is a bulleted intention with no owner and no ticket. GOOD | Add config validation to reject out-of-range rate-limiter thresholds before deploy | prevent | Dana Osei | JIRA-4821 | In Progress | WEAK | Be more careful with config changes | | | | | TRAP Leaving an action item's only record inside this document. The moment this postmortem is published, every row here needs a ticket in the tracker your team already uses; an item that lives only in this file will not get done. -->
| Action | Type | Owner | Ticket | Status ||---|---|---|---|---|| {{action_item}} | {{action_item_type}} | {{action_item_owner}} | {{action_item_ticket}} | {{action_item_status}} |incident-postmortem_template-full.md · ~5,250 tokens
---title: "{{postmortem_title}}"incident_id: "{{incident_id}}"status: "{{status}}"author: "{{author}}"incident_date: "{{incident_date}}"postmortem_date: "{{date}}"doc_type: incident-postmortemsize: fullsource_template: incident-postmortemsource_template_version: 0.1.0---
<!--FULL INCIDENT POSTMORTEM. Everything the lean variant carries, plus Detection, Trigger, Resolution, andLessons Learned, inserted in place rather than appended, so lean stays a strict ordered subset. Reach forthis variant when the incident is severe enough, cross-team enough, or novel enough that skipping straightfrom Impact to Root Causes, or from Root Causes to Action Items, would leave a reader unable to tell how theincident was found, which of your team's own criteria made it a postmortem at all, what specifically endedit, or what the team now believes differently about how it works.
THE TWO-SIZE PACKAGING IS THIS BUNDLE'S OWN DECISION, LABELLED AS SUCH. All five published postmortemtemplates this bundle's research read in full are single-size; none ships a named lean/full pair. What thecorpus does support is that depth genuinely varies in practice: one vendor's product ships configurabletemplates and "Dynamic template selection" by incident type or severity, and a real published postmortemfrom a named organization used no fixed template structure at all. This library packages that real variationas two sizes; reach for full where the incident's severity would, in a template-selecting tool, pick thedeeper option. See incident-postmortem_companion.md section 4.
A POSTMORTEM IS TRIGGERED BY A CRITERION YOUR TEAM PUBLISHED IN ADVANCE, NOT BY A BAD DAY. Google's SRE bookstates an explicit list of what should prompt one: "Common postmortem triggers include: User-visibledowntime or degradation beyond a certain threshold; Data loss of any kind; On-call engineer intervention(release rollback, rerouting of traffic, etc.); A resolution time above some threshold; A monitoring failure(which usually implies manual incident discovery)," and holds that a team should agree its own version ofthat list before an incident rather than argue it during one. See incident-postmortem_companion.md section 1.
THERE IS NO SINGLE CANONICAL SECTION SET FOR THIS DOCUMENT. The canon's own worked example and its ownfollow-up volume's worked examples use materially different headings from each other. The nine sectionsbelow are the ones this bundle could attest to a primary source and a published template, reordered forreading rather than kept in the canon's own appendix order. See incident-postmortem_companion.md section 4.
WHAT AN INCIDENT POSTMORTEM IS, AND IS NOTIt is the document a team writes after an incident to explain why it happened and what will change so itdoes not happen the same way again. It is NOT a sprint retrospective (that looks back on a period, on acadence, at how the team worked; a postmortem is event-triggered, about one specific thing that failed). Itis NOT the incident report: by one vendor's account, "if the incident report answers the question whathappened, the postmortem answers why it happened and what will prevent this from happening again," thoughthat distinction is not standards-tier, and this template treats it as one useful framing rather thansettled fact. It is NOT an After Action Review in the literal sense of that catalog alias: the US Army's ownguide describes a facilitator-led discussion running 30 minutes to 2 hours about a single training event,not a written document. It is NOT the ITIL Problem Record: a team working inside that tradition shouldexpect this document to feed the Problem Record, not replace it. See incident-postmortem_companion.mdsection 8.
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into incident-postmortem_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.2. Replace each {{placeholder}} with your content. Fill Summary and Impact first; the rest depends on knowing what actually happened before you analyze why.3. If a section does not apply, write "N/A" and one line of why, rather than deleting it silently.4. Before you share it: self-grade against incident-postmortem_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{postmortem_title}}
## Summary
<!-- WHAT A short synopsis of what happened, written before the analysis that follows it. Read first, even though it is usually the last thing written. WHY Attested by the canon's own worked example and three of the five published templates this bundle's research read in full. A reader who stops here should still know what happened, how bad it was, and whether it is over. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Summary). ASK What happened, in two or three sentences someone outside the incident could follow? What is the current status: resolved, mitigated, or still ongoing? GOOD "On July 14, a config change to the Saved Views API's rate limiter caused it to reject 60 percent of legitimate requests for 42 minutes before the change was rolled back. Fully resolved; no data was lost." WEAK "The Saved Views API had some issues." (no cause, no duration, no resolution state; a reader learns nothing from the section built to be read alone) TRAP Writing the Summary as a preview of Root Causes. It is a synopsis for someone who may never read past it, not a teaser for the analysis below. -->
{{summary}}
## Impact
<!-- WHAT Who and what was affected, and how badly: systems, users, duration, and any data loss. WHY The most broadly attested section besides Timeline: the canon's own example and four of the five published templates read carry it. A postmortem that skips straight from Summary to Root Causes leaves the reader unable to judge whether the analysis that follows was worth the time it took to write. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Impact). ASK What system or user population was affected? For how long? What was the magnitude (percentage of requests failed, records affected, revenue)? Was any data lost? PRIORITY List the most severe or most visible impact first. State a real number, not a category; "some users" is not measurable. ROW HINT A good row names one affected system or population, a duration, and a magnitude a reader could check against a dashboard or a support queue. A weak row is a category with no number. GOOD | Saved Views API | 42 minutes | 60 percent of requests returned 429; no data loss | WEAK | The API | A while | Users were affected | TRAP Reporting only technical impact and skipping user or business impact, or the reverse. "The API was down" hides whether anyone outside engineering noticed. -->
| Affected system or population | Duration | Magnitude ||---|---|---|| {{impact_affected}} | {{impact_duration}} | {{impact_magnitude}} |
## Detection
<!-- WHAT How the incident was discovered: a monitoring alert, an on-call page, a customer report, and the delay between the underlying failure starting and someone finding out. WHY Attested by the canon's own worked example and two of the five published templates read. It is full-only because two other templates read fold this content into a combined heading or drop it entirely, so the concept is real but not yet universal enough for the smallest version of this document. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Detection). ASK What actually surfaced this: an alert, a page, a customer ticket, or an engineer noticing something by chance? How long after the underlying problem started was it detected? Would an earlier signal have caught it sooner? GOOD "Detected by the SavedViewsAPI-HighErrorRate PagerDuty alert, 6 minutes after the rate limiter config deployed. No customer report arrived before the alert fired." WEAK "We noticed the API was having problems." (no named signal, no delay, no source; nothing here tells the next reader whether monitoring worked or the team got lucky) TRAP Recording only that an alert fired, not how long after the failure began. The gap between failure and detection is often the more useful number to reduce. -->
{{detection_narrative}}
## Timeline
<!-- WHAT A chronological, timestamped account of what happened, in order. WHY The single most attested title in this bundle's research, appearing in the canon's own example and four of the five published templates read. It is also this document type's sharpest lesson about received wisdom: the SRE book's own postmortem chapter, the source most people would name if asked where "timeline" comes from, contains zero occurrences of the word in its prose. It exists only as a heading in the separately linked worked example. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Timeline) and section 1. ASK In order, what happened, and when? Include the failure's start, detection, escalation, any mitigation attempts, and resolution, each with a timestamp. PRIORITY List every entry in chronological order, earliest first. Do not skip the boring entries; a gap in the timeline is itself information about what nobody was watching. ROW HINT A good row has a timestamp precise enough to check against a log or a chat transcript, and one clear event. A weak row has a vague time ("morning") or bundles several events into one line. GOOD | 14:02 UTC | Rate limiter config change deployed to production | WEAK | Sometime that afternoon | Things started going wrong | TRAP Writing the timeline from memory days later without checking it against logs, alerts, or chat history. A timeline nobody can verify is a story, not a record. -->
| Time (UTC) | Event ||---|---|| {{timeline_time}} | {{timeline_event}} |
## Trigger
<!-- WHAT The specific, named criterion, agreed by the team in advance, that made this event a postmortem rather than an ordinary day of operations. WHY The narrowest attestation of any section in this bundle, canon plus one published template, which is why it is full-only. It also does the most conceptual work of any section here: it is where a team writes down which of its own agreed-in-advance criteria fired, rather than leaving readers to infer after the fact why this event, and not some other bad day, got a document. Google's own canon publishes an example list of what such criteria might cover: "Common postmortem triggers include: User-visible downtime or degradation beyond a certain threshold; Data loss of any kind; On-call engineer intervention (release rollback, rerouting of traffic, etc.); A resolution time above some threshold; A monitoring failure (which usually implies manual incident discovery)." Name your own team's published criterion here, not this list restated as though it were yours. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Trigger) and section 1. ASK Which of your team's own published criteria fired? Where is that criteria list published, so a reader could check it? Would this event have qualified without that specific criterion? GOOD "Meets our published criterion 'resolution time above 30 minutes' (see team runbook, section 2). Also meets 'on-call engineer intervention': the rate limiter config was rolled back manually." WEAK "This felt like a big enough deal to write up." (no named criterion, nothing a future reader could check against, no way to tell whether the next similar incident would also qualify) TRAP Restating Google's example list, or any other borrowed list, as though it were your team's own criteria. Naming a trigger your team never actually published is inventing a rule retroactively. -->
{{trigger_narrative}}
## Root Causes
<!-- WHAT The contributing causes, plural by design, each with the evidence behind it. WHY Attested by the canon's own example, and, under varying names (GitLab's "Root Cause Analysis," Elastic's singular "Root Cause," Atlassian's split "Root cause identification" and "Root cause"), close to universal: present in four of the five published templates read. This is also the one section this bundle ships over a live, unresolved argument. A named, cross-citing community (Cook, Allspaw, Hollnagel, Woods, Dekker, Leveson) holds there is no single root cause in a complex system at all; the clearest named defence found still concedes, "sometimes you won't find a root cause. It happens." This template does not settle that argument; it asks for causes, plural, and evidence for each. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Root Causes) and section 6. ASK What conditions, each necessary but not sufficient alone, combined to produce this incident? What evidence (a log line, a config diff, a metric) supports each one? Have you checked whether any candidate cause is actually a person's name? PRIORITY List every contributing cause you have evidence for, not just the first one found. A single-row Root Causes table is usually a sign the investigation stopped too early. ROW HINT A good row names one specific, checkable condition and the evidence behind it. A weak row names a broad category ("human error," "process failure") or a person. GOOD | The rate limiter's config validation did not reject an out-of-range threshold before deploy | Deploy log shows the config passed CI with no validation step for that field | WEAK | Someone made a mistake | Just experience | TRAP Naming a person as a cause. GitLab's own guidance states the rule plainly: "A root cause can **never be a person**." A postmortem that lands on an individual has produced blame, not analysis, and directly contradicts the blameless framing this document type is named for. -->
| Contributing cause | Evidence ||---|---|| {{root_cause}} | {{root_cause_evidence}} |
## Resolution
<!-- WHAT What specifically ended the incident: the technical or operational action taken. Distinct from Action Items, which are future work, not what already happened. WHY Attested by the canon's own example and two published templates read. It is full-only because two other templates read fold this content elsewhere: one calls the equivalent moment "Recovery," and another folds it into timestamp fields rather than naming it separately. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Resolution). ASK What action actually stopped the incident: a rollback, a config change, a manual failover, a fix deployed? Who took it, and when (cross-check against the Timeline above)? GOOD "The rate limiter config change was rolled back at 14:44 UTC by the on-call engineer, restoring the previous threshold. Error rate returned to baseline within 3 minutes of the rollback completing." WEAK "We fixed it." (no action named, no owner, no timestamp to check against the Timeline) TRAP Describing a fix that has not actually shipped yet as though it already resolved the incident. If the underlying condition is still present and only masked, say so here and put the real fix in Action Items. -->
{{resolution_narrative}}
## Lessons Learned
<!-- WHAT What the team concludes it should change about how it works, distinct from the specific technical fix recorded under Resolution and the concrete tickets recorded under Action Items. WHY Attested by the canon's own worked example, which groups this content under named subsections, and one published template read. This template does not reproduce those subsection titles verbatim, since only the section's existence, not its exact wording, was confirmed for this bundle; write your own headings below if grouping helps. It is full-only on the same logic as Detection and Resolution: real, but not yet attested widely enough for the smallest version of this document. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Lessons Learned). ASK What does this incident reveal about how the team works, beyond the specific bug: a gap in testing, a monitoring blind spot, an assumption that turned out false? What would you tell a team about to build something similar? GOOD "We assumed config changes below a certain size never needed a canary rollout. This incident shows that assumption is false for anything touching the rate limiter; size is not a good proxy for blast radius here." WEAK "We should be more careful next time." (not a lesson anyone could act on, indistinguishable from what every postmortem says) TRAP Restating the Root Cause in different words. A lesson is what the team now believes about how it works; a root cause is what happened this time. -->
{{lessons_learned}}
## Action Items
<!-- WHAT Concrete, owned, tracked follow-up work. WHY Attested by the canon's own example and two published templates read, but the name itself is a pick: published templates write it three different ways and never converge: "Action Items" (two vendor templates), "Corrective actions" (two others), and "Follow-up actions" (a fifth). This bundle uses the canon's own word and says plainly that it is a pick, not a convergence. The canon's own worked example shows what a filled row looks like: "Plug file descriptor leak in search ranking subsystem | prevent | agoogler | Bug 5554825 DONE," a specific owner and a specific tracked bug, not a bulleted wish. The canon's own row combines the ticket ID and status in one field; this template splits them into separate columns for clarity, which is this template's own formatting choice. Two independent practitioner sources agree on where these rows live once the postmortem is finished: "The postmortem document is where action items are born. It is not where they should live," and a second, independent source says nearly the same thing. Link the ticket; do not leave the work only here. Neither source names a risk register or a RAID log as a destination; this library's own family contract offers those alongside the product backlog, which is this library's own convention rather than received postmortem practice. Deep dive: incident-postmortem_companion.md section 3 (Anatomy > Action Items). ASK What specific, owned action follows from this incident? Is there already a ticket for it in the tracker your team actually uses? Who owns it, and is the status current? PRIORITY List every action with an owner and a ticket link, prevention items before general cleanup. An action with no owner is not an action item, it is a wish. ROW HINT A good row names one concrete action, its type (the canon's own quoted row uses "prevent"; mitigate, detect, and process are reasonable values this bundle's research did not see quoted), a single named owner, and a linked ticket with a current status. A weak row is a bulleted intention with no owner and no ticket. GOOD | Add config validation to reject out-of-range rate-limiter thresholds before deploy | prevent | Dana Osei | JIRA-4821 | In Progress | WEAK | Be more careful with config changes | | | | | TRAP Leaving an action item's only record inside this document. The moment this postmortem is published, every row here needs a ticket in the tracker your team already uses; an item that lives only in this file will not get done. -->
| Action | Type | Owner | Ticket | Status ||---|---|---|---|---|| {{action_item}} | {{action_item_type}} | {{action_item_owner}} | {{action_item_ticket}} | {{action_item_status}} |incident-postmortem_example.md
---title: "Postmortem: Saved-View Aggregate Disclosed Data Across an Entitlement Boundary (DEF-2291)"incident_id: "DEF-2291"status: "final"author: "Marcus Bell (Staff Engineer, Reporting)"incident_date: "2026-07-13"postmortem_date: "2026-07-16"doc_type: incident-postmortemsize: fullsource_template: incident-postmortemsource_template_version: 0.1.0related: - "../bug-report/bug-report_example.md (DEF-2291, the defect this postmortem examines)" - "../test-plan/test-plan_example.md (Saved Views test plan; risk R-05, the suspension and resumption rules)" - "../test-case/test-case_example.md (TC-047, the case that caught this at step 4)" - "../sdd/sdd_example.md (Saved Views design; the entitlement re-check this incident bypassed)" - "../risk-register/risk-register_example.md (program risk register; R-05 was already the test plan's top-tier risk before this incident, and this postmortem's first action escalated it to the steering group on 2026-07-14)" - "../raid-log/raid-log_example.md (program RAID log; where R-05's steering-group escalation is tracked)" - "../product-backlog/product-backlog_example.md (Saved Views product backlog; two action items land here)"---
> **Worked example.** A filled `incident-postmortem`, full variant, examining> [DEF-2291](../bug-report/bug-report_example.md), the same defect the `bug-report`, `test-plan`, `test-case` and> `sdd` examples already describe from their own vantage points. It is dated 2026-07-16, inside Sprint 24> (2026-07-13 to 2026-07-24) and one day after Phase 2 (sharing) testing resumed on 2026-07-15, which is the> earliest a real postmortem write-up could plausibly follow the fix being verified. It closes before two later> documents this postmortem's own action items seed: the `definition-of-done` example's amendment on 2026-07-24,> and the `runbook` example's entitlement-audit reconciliation job, added in build 2.4.0 "after DEF-2291." Neither> is cited below except here, because the body may only know what had already happened by 2026-07-16.>> The incident itself never reached a customer: it was caught on staging by a planned test, not an alert or a> report, which is deliberately unlike the template's own rate-limiter scenario and is the point this example> makes about triggers in the Trigger section below. Figures marked "illustrative" are made up for the example.
# Postmortem: Saved-View Aggregate Disclosed Data Across an Entitlement Boundary (DEF-2291)
## Summary
On 2026-07-13, the first day of Phase 2 (sharing) testing for Saved Views, an automated regression assertion(TC-047 step 4) caught a shared dashboard view whose aggregate total included revenue from a region therecipient was not entitled to see, even though the same response's row-level data was filtered correctly. Thedefect, tracked as DEF-2291, never reached production: it was confined to staging build 2.3.1, found by aplanned test rather than a customer report or a monitoring alert, and exercised only synthetic QA personas. Itstill met Acme's own postmortem trigger and suspended all Phase 2 sharing testing for two days. Fixed andverified 2026-07-15: the aggregate computation now runs behind the entitlement filter instead of ahead of it,and the full 12-combination permission matrix was re-executed from the start before Phase 2 resumed the sameday.
## Impact
| Affected system or population | Duration | Magnitude ||---|---|---|| Production customers | None; the code path never left staging | Zero exposure. Both accounts in the reproduction, R and O, are synthetic QA personas, and the `saved_views` sharing flag had not opened to any real account || Phase 2 (sharing) test execution, Reporting Squad | 2026-07-13 15:10 UTC to 2026-07-15 (about two days) | All Phase 2 testing suspended under the test plan's suspension rule; on resumption the entire 12-combination permission matrix was re-executed from the start rather than only the failing case, per the plan's resumption rule || Confirmed exposure window, staging only | At most 7 days (illustrative): bounded between the sharing code reaching staging under build 2.3.1 by the test plan's 2026-07-06 entry date, and detection on 2026-07-13 | Limited to the two personas used in TC-047 and TC-048. A manual correlation of the existing `shared_view_served` and `permission_check_passed` staging log streams found no other account requesting a region-restricted shared view in that window (illustrative) |
## Detection
Detected by TC-047 step 4, the automated aggregate assertion added at case version 1.1 after Sam Okafor's2026-07-08 security review. The assertion failed on the first pipeline run against the Phase 2 build, themorning Phase 2 execution opened per the test plan's schedule. The same application log entry that shows theentitlement re-check passing (request ID `7f3c-4a11-a91b`, 14:22 UTC, illustrative) is the request the failingassertion flagged: the row filter ran correctly and the log says so, which is exactly why nothing about thisdefect would have shown up in a re-check-failure alert even if one had existed. Anjali Rao (QA Lead) reproducedthe failure manually against the same build, five of five times via the API and three of three in the browser,to rule out a flaky assertion before filing DEF-2291 the same day.
No customer report arrived, and none could have. The build was staging-only, and this class of defect producesno user-visible symptom at all: the row-level data a user actually sees was correct throughout, which is whatmade the leak invisible to every earlier entitlement case in the suite and is the subject of Root Causes below.
## Timeline
| Time (UTC) | Event ||---|---|| 2026-07-06 | Build 2.3.1 deployed to staging with the `saved_views` flag enabled, per the test plan's entry criteria. The sharing code path, including the aggregate-before-filter pattern, is present from this point at the latest || 2026-07-13, ~09:00 (illustrative) | Phase 2 (sharing) execution opens per the test plan's schedule. The automated entitlement regression suite, including TC-047, runs against the Phase 2 build || 2026-07-13, 14:22 | TC-047 step 4's aggregate assertion fails on request ID `7f3c-4a11-a91b` (illustrative). The same request's row-level check passes and is logged as passing || 2026-07-13, ~14:45 (illustrative) | Anjali Rao manually reproduces the failure, 5 of 5 via the API and 3 of 3 in the browser, and files DEF-2291 || 2026-07-13, 15:10 | Triage (Priya Nair, Marcus Bell, Sam Okafor) confirms Severity S1, raises Priority to P1, and Anjali Rao suspends all Phase 2 sharing testing under the test plan's suspension rule, notifying the same three that afternoon || 2026-07-14, ~10:00 (illustrative) | PR 812 (illustrative) merges, moving the aggregate computation behind the entitlement filter. Build 2.3.2 releases || 2026-07-14, later (illustrative) | Risk R-05 is escalated to the steering group with the specific ask raised at triage: fund a platform-level entitlement-aggregate control, or accept the residual formally || 2026-07-15 | Anjali Rao re-verifies against the original TC-047 steps. The full 12-combination permission matrix is re-executed from the start and all 12 pass. Sam Okafor confirms the security review unblocked. Phase 2 sharing testing resumes the same day |
## Trigger
This event meets Acme's own published postmortem criterion: any defect confirmed at Severity S1 on the squad'sfour-level scale (S1 Critical / S2 Major / S3 Minor / S4 Trivial) triggers a postmortem regardless ofenvironment or customer impact, because data exposure is S1 by definition on that scale. It also meets asecond, independent criterion under the same policy: a confirmed entitlement failure that fires the testplan's suspension rule, stopping planned testing outright rather than being logged as a defect to triage,qualifies on its own.
This is deliberately not Google's published trigger list restated as though it were Acme's own. Google's canonleads with user-visible downtime, data loss, on-call intervention, a resolution-time threshold and a monitoringfailure; none of those apply cleanly here. There was no downtime, because nothing was live. There was noon-call intervention, because this was caught by a scheduled test run, not paged. Resolution time is not ameaningful trigger for a defect fixed inside one calendar day. What actually fired was a criterion built forexactly this shape of event: a severity scale that treats any confirmed data-boundary exposure as critical,independent of who was exposed or how many. Without that criterion, this incident would likely have closed asa well-handled defect and nothing more, and the questions this document asks, why the acceptance criteria neverreached this case and what would have caught it sooner, would not have been asked at all.
## Root Causes
| Contributing cause | Evidence ||---|---|| The dashboard tile's aggregate (`total_revenue`, and equally the row-count badge) was computed in the query layer before ViewsController's entitlement re-check ran, so the re-check filtered the returned `rows` array but never touched the already-computed aggregate | PR 812's diff (illustrative) shows the aggregate assignment executing three calls before the entitlement filter is applied; reproduced 5 of 5 via the API against build 2.3.1 || No acceptance criterion for the Saved Views epic addressed what a partially entitled recipient should see inside an aggregate; the agreed criteria stop at the row level (save, load, default, share) | Confirmed against the full acceptance-criteria set for the sharing story at triage on 2026-07-13; the design document's re-check statement was the only source for the expected behavior, not a signed-off criterion || The two earlier entitlement cases, TC-046 (fully entitled) and TC-048 (not entitled), assert only on returned rows, the pattern every earlier case used. Only the partially entitled partition can expose an aggregate leak, and TC-047 could not have caught it before version 1.1 added the aggregate assertion on 2026-07-08 | TC-047's version history. TC-048 passed against the same build 2.3.1 and did not surface this, confirming the gap was specific to partial entitlement, not to the test suite generally |
## Resolution
PR 812 (illustrative) moved the aggregate computation behind the entitlement filter, so `total_revenue` and therow-count badge are now derived from the same permitted row set the response's `rows` array already used. Build2.3.2 released on 2026-07-14. Anjali Rao re-verified on 2026-07-15 against the original TC-047 steps:`total_revenue` returned the AMER-only total for recipient R, and owner O's own view was unchanged. Per the testplan's resumption rule, the entire 12-combination permission matrix was re-executed from the start, not only thefailing case, because an entitlement defect invalidates the assumption behind every result the matrix hadalready produced; all 12 passed. Sam Okafor confirmed the security review unblocked the same day, and Phase 2(sharing) testing resumed on 2026-07-15.
## Lessons Learned
Row-level correctness is not evidence of entitlement correctness. Every earlier entitlement case in this suiteasserted only on returned rows, which is why none of them, and no acceptance criterion, ever exercised thispath. The squad's working assumption now is that any dashboard element derived from a filtered row set, anaggregate, a count, a chart label, needs its own assertion tying it to the same filter, rather than inheritingcorrectness from the rows being right.
The test plan's risk-ranked approach did what it was built to do; the acceptance criteria did not, and werenever going to. Risk R-05 was already the test plan's highest tier before this incident, which is exactly whyTC-047 existed to look for it. The acceptance criteria, by contrast, describe what the business agreed to, andnobody had agreed to specify what a partially entitled recipient sees inside a total, because the question onlybecomes visible once you look at how sharing is implemented. This incident is a data point for trustingrisk-ranked test design over acceptance criteria on exactly the surfaces acceptance criteria are not built toreach, not a reason to distrust either one.
The squad's Definition of Done does not yet require the permission matrix to be re-executed wheneverentitlement-relevant code changes; nothing in it would have caught this earlier than TC-047 did, and nothing init currently prevents the same class of gap on a future change to this code path. Marcus Bell intends to raisethis specifically at Sprint 24's close, rather than folding it into this document's own action items below,because Acme's Definition of Done is owned and amended at sprint planning, not by a postmortem.
## Action Items
| Action | Type | Owner | Ticket | Status ||---|---|---|---|---|| Escalate risk R-05 with the specific ask agreed at triage: fund a platform-level entitlement-aggregate reconciliation control, or accept the residual formally at board level | prevent | Sam Okafor | Risk register R-05 | Escalated to the steering group since 2026-07-14; open || Add an automated reconciliation job that cross-checks every `shared_view_served` event against a matching `permission_check_passed` event for the same request, so a future leak surfaces the same day rather than only at the next planned risk-tier test | detect | Marcus Bell | Product backlog SV-8 (proposed) | Open, raised 2026-07-16 || Add TC-053, asserting the same aggregate-before-filter defect class against the dashboard's row-count badge, the other metric on this computation path | detect | Anjali Rao | Product backlog SV-10, closed on delivery of TC-053 | Done, 2026-07-15 || Write an acceptance criterion for the sharing story covering what a partially entitled recipient sees inside an aggregate, closing the gap Root Causes names above | process | Priya Nair | Product backlog SV-9 (proposed) | Open, raised 2026-07-16 |Provenance
Section titled “Provenance”The reasoning, the history and every source, in the repository:
- Companion - the long-form argument: why these sections, where the sources disagree, and what the bundle refuses to claim
- History - what changed in this bundle, and when
- Research log - every source consulted, with what each one actually supports
- Catalog metadata - the machine-readable record this page is generated from
Catalog record: 9 sections across 1 format(s), methodology SRE/DevOps, typically owned by SRE / IC.