Runbook
beta · Family standing-standards · Phase undefined · Sizes lean, full · ~2,600 tokens
The document a responder opens when a known alert fires, so that the response does not depend on who is awake. Scoped to one situation rather than one service: the trigger that invokes it, the steps that resolve it, and the check that confirms it worked.
The short card. Why the document is shaped this way, and the argument behind every rule here, is in
runbook_companion.md. A fully worked instance is
runbook_example.md.
When to use
Section titled “When to use”- A known alert or condition can fire again, and the response should not depend on who is on call when it does.
- More than one person might have to run this procedure, so the steps cannot live only in one engineer’s head.
- The situation is scoped to one trigger: one alert, one condition, one page. If you can name it in a sentence, it belongs here.
- You want a named signal that tells a responder the situation is actually resolved, not just quiet.
- The work is regulated or audited, and a record of what was checked has to survive the incident.
When NOT to use
Section titled “When NOT to use”- You do not yet have a known trigger. If nothing has paged yet and you are documenting a service in general, write a wiki page or a design doc, not a runbook. A runbook exists because of a trigger, not a topic.
- The steps never vary. If the on-call engineer runs the same deterministic commands every time this alert fires, that is a candidate for automation, not more documentation. Writing a better runbook for a fully deterministic procedure treats a symptom instead of the cause.
- You need to recover a whole system at an alternate facility after a major failure. That is a disaster recovery plan, one register above what this template covers.
- You are documenting an entire service: its overview, deployment, security posture and disaster recovery in one place. That is a standing service-operations manual, a bigger scope than this template, not a bigger size of it. Growing Prerequisites and Access past a handful of lines is usually the first sign this is happening.
- You only need a memory guard against omission, the Gawande sense of a checklist, not a full procedural walkthrough. A checklist is deliberately smaller than a runbook.
- The content only explains, and never directs. “Check the logs and restart the service if needed” is a knowledge base article. A runbook must go further: a specific action for a specific trigger.
Pick a variant
Section titled “Pick a variant”Lean (four sections) is the default: Purpose and Trigger, Procedure, Validation, Review Trigger. It is close to the smallest shape actually published for this kind of document, and a runbook a responder cannot read in the middle of an incident has already failed at its one job.
Full (seven sections) adds Prerequisites and Access, Remediation and Cleanup, and When This Does Not Apply. Notice what the three additions have in common: none of them is content a responder needs while the alert is actively firing. Prerequisites is read before the incident starts, Remediation and Cleanup after it ends, and When This Does Not Apply is a scoping note rather than a step. Move to full when at least one of these is true:
- the environment is regulated, and Validation needs to double as an audit trail;
- more than one team could plausibly run this procedure, so an explicit boundary and named prerequisites keep the wrong team from running the wrong runbook;
- the service is complex enough that “this does not apply here” needs to be said rather than assumed.
Every lean heading appears in full unchanged, in the same order, so growing from lean to full is additive. You never rewrite what you already agreed.
Quality rubric (self-grade)
Section titled “Quality rubric (self-grade)”Score each 0, 1 or 2. Full below 11 out of 16 ships a document a responder cannot run without guessing, which is the one failure this artifact exists to prevent; lean is scored on five of these rows only (see the scope table below), and below 7 of 10 the same failure applies to the smaller shape.
| # | Criterion | 0 | 1 | 2 |
|---|---|---|---|---|
| 1 | Trigger is specific | A general topic, not a stated alert or condition | Names an alert, but no stated system or outcome | Names the specific alert or condition, the system it belongs to, and a one-line outcome you could check later |
| 2 | Style is declared | No statement of general or step-by-step | States which, but later steps drift into the other style | States which explicitly, and every step reads consistently with that choice |
| 3 | Steps resolve, not describe | Steps like “open the dashboard” or “check the logs” with nothing named | Some steps name a command or dashboard; others stay vague | Every step names the exact action and a result a second person could check without asking the author |
| 4 | Validation resists false negatives | “Confirm it’s fixed,” no named signal | One named signal, and it is only the alert clearing | A named signal independent of the alert clearing, with the value that counts as pass and where to look |
| 5 | Prerequisites earn their place (full) | Nothing listed, or a general “have access” | Items listed but not tied to a specific procedure step | Every item is something that would actually block step 1, named specifically, with how to get it before an incident |
| 6 | Cleanup items are owned (full) | Nothing recorded, or an unowned “fix later” | Items listed, but missing an owner or a deadline | Every temporary change has a named owner and a trigger or deadline for reverting it |
| 7 | Near-misses are named honestly (full) | The section is missing or deleted | Filled with a catch-all vague enough to describe anything | Names specific lookalike symptoms and their actual cause, or states “none identified” as a considered answer |
| 8 | Trigger is event-based | A calendar cadence alone, or nothing | An event is named, but no owner | A specific system event is named, paired with a named owner and what they do about it |
Which rows apply to what.
| Document | Rows | Maximum | Score against |
|---|---|---|---|
| full | all 8 | 16 | 11 |
| lean | 1-4 and 8 | 10 | 7 |
Rows 5, 6 and 7 are scored only against full. Lean ships none of Prerequisites and Access, Remediation and Cleanup, or When This Does Not Apply, so grading it on those rows would penalise the choice of variant rather than the quality of the document.
The test behind every cell above: could someone satisfy it without improving the document? A row that counted prerequisite items, cleanup rows, or near-misses would reward padding. Every cell instead asks whether a specific piece of evidence exists, and whether a second person, not the author, could find it.
Named anti-patterns (the usual wrecks)
Section titled “Named anti-patterns (the usual wrecks)”- Silent decay. A runbook does not announce that it has gone stale; the system it describes keeps changing while the document does not, and the gap is invisible until the moment someone actually needs it. The fix is a real Review Trigger: a named event and a named owner, not a hope.
- Written for yourself, not for the reader. The author knows which dashboard, which credentials, which unstated assumptions hold. The next responder does not. The tell is a step that only makes sense to the person who wrote it; the fix is having someone who did not write the runbook try to run it.
- The wiki-versus-code mismatch. The runbook lives in a wiki. The system it describes lives in code. Nothing forces the two to change together, so they drift apart on different schedules maintained by different people.
- A knowledge-base sentence dressed up as a step. “Check the logs and restart the service if needed” explains; it does not direct. A runbook step names the exact log, the exact command, and the result that tells you it worked.
- Too many steps in one document. A responder under pressure cannot scan a runbook that tries to cover several unrelated situations at once. If you are listing multiple unrelated triggers, you are writing two runbooks under one title; split them.
- Not linked from the thing that triggers it. A correct runbook nobody can find during the incident is functionally the same as no runbook. If the alerting tool supports linking an alert to its runbook, do it.
- Calendar-only review, treated as sufficient. A quarterly cadence or an untouched-duration heuristic catches staleness eventually, but neither catches it at the moment the system actually changed, which is the moment it matters. Pair any calendar cadence with a named event trigger.
- Citing an unmeasured number as though it were a finding. Runbooks are widely asserted to shorten incidents and rarely, if ever, independently measured doing so. State the value case honestly, as an assertion, rather than attaching a specific percentage nothing behind it can support.
Pairing with your process
Section titled “Pairing with your process”This bundle ships in the standing-standards family alongside definition-of-done. A definition of done is
a standard a team is judged against; a runbook is an instrument a team executes. Keep the two apart: a
runbook does not certify that work is finished, and a definition of done does not tell a responder what to
type. Where your incident tooling supports it, link this document from the alert it answers, per the Purpose
and Trigger section; that link is what turns a correct runbook into a findable one.
The artifacts
Section titled “The artifacts”runbook_template-lean.md · ~2,600 tokens
---title: "{{runbook_title}}"service_or_system: "{{service_or_system}}"triggering_alert: "{{triggering_alert}}"owner: "{{owner}}"status: "{{status}}"last_updated: "{{date}}"doc_type: runbooksize: leansource_template: runbooksource_template_version: 0.1.0---
<!--LEAN RUNBOOK. The smallest procedure that is still a real runbook: what invokes it, what to do, how toconfirm it worked, and what would make it wrong. Four sections, because a runbook a responder cannot read inthe middle of an incident has already failed at its one job. To carry prerequisites, cleanup, and an explicitboundary (see runbook_template-full.md), ADD sections; never rename or reorder the ones below, because thefull variant is a strict superset of this one.
A RUNBOOK IS THE PROCEDURE FOR ONE TRIGGER, NOT ONE SERVICE. It is scoped to the alert or condition thatopens it, not to everything a responder might ever need to know about the system. If you find yourselflisting several unrelated triggers here, you are probably writing two runbooks under one title. Seerunbook_companion.md sections 1 and 3.
A NOTE ON THE NAME. The Google literature this practice traces back to calls this artifact a "playbook," nota runbook; the word "runbook" appears in that canon only inside a third-party case study, never in Google'sown analytical prose. This library keeps the catalog's own name, runbook, and says so plainly rather thanimplying the canon uses it. See runbook_companion.md section 2.
WHAT A RUNBOOK IS, AND IS NOTIt is the executable procedure a responder opens when a known situation occurs. It is NOT a knowledge basearticle ("check the logs and restart the service if needed" is not a runbook; it must direct a specificaction), NOT a disaster recovery plan (that recovers one or more systems at an alternate facility after amajor failure, one register above this), NOT a checklist in the Gawande sense (a minimal memory guard againstomission, not a full procedural walkthrough), and NOT the standing service-operations manual some sourcespublish under the same name, which covers an entire service rather than one trigger. Seerunbook_companion.md section 8.
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into runbook_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.2. Replace each {{placeholder}} with your content. Fill Purpose and Trigger and Procedure first; everything else depends on them.3. If a section does not apply, write "N/A" and one line of why, rather than deleting it silently.4. Before you ship it: self-grade against runbook_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{runbook_title}}
## Purpose and Trigger
<!-- WHAT What invokes this runbook, the system or service it is about, and the one-line job it does. Name the specific alert, condition, or page, not a general topic. WHY The Google source most of this literature traces to states that "whenever an alert is created, a corresponding playbook entry is usually created." A runbook with no stated trigger cannot be told apart from a wiki page about the service, which is the exact boundary a knowledge base article fails to cross. Deep dive: runbook_companion.md section 3 (Anatomy > Purpose and Trigger) and section 8. ASK What is the specific alert, condition, or page that opens this document? What system or service does it belong to? What is the one-line outcome this runbook exists to produce? Is it linked from the alerting tool itself? GOOD "Triggered by the SavedViewsAPI-HighErrorRate PagerDuty alert on the Saved Views API. Purpose: restore the API error rate below 1 percent within 15 minutes of the page firing." WEAK "This runbook covers the Saved Views service." (no named alert, no condition, no stated outcome; a responder cannot tell this apart from a wiki page) TRAP Naming several unrelated triggers in one document. That is usually a sign you are writing two runbooks under one title; split them. -->
{{purpose_and_trigger}}
## Procedure
<!-- WHAT The steps that resolve the situation the trigger describes, starting from a known state. Say plainly, in one sentence, whether this procedure is written GENERAL (principles, changes slowly) or STEP-BY-STEP (exact commands, reduces variability), and why. WHY Google names this choice contentious inside its own organisation: some argue for general entries that change slowly, others for step-by-step entries that reduce human variability and drive down MTTR, and its own chapter calls it "a contentious topic." This template does not pick a side its source declines to pick; it asks you to state which you are writing. The same source draws one hard line regardless of style: "if your playbooks are a deterministic list of commands that the on-call engineer runs every time a particular alert fires," automate it instead of documenting it. Deep dive: runbook_companion.md section 3 (Anatomy > Procedure) and section 6. ASK Are you writing general guidance or exact commands, and why, given who runs this and how often? What is the known starting state? What are the ordered steps? Has any step here been run unchanged so many times that it is now a candidate for automation instead of documentation? PRIORITY Number the steps in the order a responder actually performs them. Start from a known state rather than "open the dashboard," which hides exactly the context a responder under pressure does not have. ROW HINT A good row names one concrete action, the exact command or decision it requires, and the result that tells the responder it worked before moving on. A weak row is a vague verb with no way to tell whether it succeeded. GOOD | 1 | Run `kubectl rollout restart deploy/saved-views-api -n prod` | Rollout status shows all pods Ready within 3 minutes | WEAK | 1 | Restart the service | It should come back | TRAP Writing "open the dashboard" or "check the logs" as a step with no named dashboard, no named log, and no signal to look for. That is a knowledge-base sentence, not a runbook step. -->
{{procedure_approach}}
| Step | Action | Expected result ||---|---|---|| {{step_number}} | {{step_action}} | {{step_expected_result}} |
## Validation
<!-- WHAT The specific signal that confirms the situation is actually resolved before you close out: a metric back to baseline, an alert clearing, a health check passing. WHY "Confirm it's fixed" cannot be checked by someone else; a named signal can. In a regulated environment this table is also your audit trail, so it should let a reviewer see what was checked, not just that something was. Deep dive: runbook_companion.md section 3 (Anatomy > Validation). ASK What signal tells you the situation is actually resolved, not just quiet? Where do you look for it? How do you tell a false negative (it looks fine but is not) from a real recovery? PRIORITY List every check a responder must clear before closing the incident, not just the first one that looks reassuring. Order them in the sequence you would actually check them. ROW HINT A good row names one specific, observable signal, the value or state that counts as pass, and exactly where to look. A weak row is "confirm it's fixed" with no location and no threshold. GOOD | Error rate | Below 1 percent for 5 consecutive minutes | Saved Views API dashboard, error-rate panel | WEAK | Everything looks fine | OK | Eyeballing the dashboard | TRAP Treating the alert clearing as the only check. An alert can clear because its threshold rearmed, not because the underlying cause is fixed; pair it with a second, independent signal. -->
| Check | What counts as pass | Where to verify it ||---|---|---|| {{validation_check}} | {{validation_signal}} | {{validation_method}} |
## Review Trigger
<!-- WHAT The event that would make this runbook wrong, and the named person or role expected to notice it. Not a calendar date alone. WHY A runbook fails silently: "the decay is silent until the moment of failure," and the mechanism is a process mismatch: the runbook lives in a wiki while the system lives in code, maintained by different people. No source this bundle's research read names an event-plus-owner trigger; every one reaches for a calendar, a quarterly cadence or a 90-day-untouched heuristic. This section is this bundle's own contribution, built to close that gap, not received industry practice. Deep dive: runbook_companion.md section 3 (Anatomy > Review Trigger) and section 6. ASK What system change would make this runbook wrong: a redeploy, a schema change, a dependency swap, an ownership handover? Who is on the hook to notice when that happens? What do they do about it, re-verify, update, or retire the runbook? PRIORITY Pair every date-based cadence with at least one event-based trigger. An event with no named owner is not a trigger, it is a hope. ROW HINT A good row names a specific system change, the person or role who notices it, and what they do next. A weak row is "review quarterly" with nobody named. GOOD | Saved Views API redeployed on a new runtime | On-call lead for Saved Views | Re-run this procedure against the new environment before the next on-call rotation starts | WEAK | Review this runbook quarterly | Team | Update if needed | TRAP Naming only a calendar date. A quarterly review passes every audit and can still be wrong the day an incident actually needs it; pair the date with the event that breaks the document. -->
| Event that would make this wrong | Owner who notices | What they do about it ||---|---|---|| {{review_event}} | {{review_owner}} | {{review_action}} |runbook_template-full.md · ~4,350 tokens
---title: "{{runbook_title}}"service_or_system: "{{service_or_system}}"triggering_alert: "{{triggering_alert}}"owner: "{{owner}}"status: "{{status}}"last_updated: "{{date}}"doc_type: runbooksize: fullsource_template: runbooksource_template_version: 0.1.0---
<!--FULL RUNBOOK. Everything the lean variant carries, plus Prerequisites and Access, Remediation and Cleanup,and When This Does Not Apply. Use it where a runbook needs to stand up to something beyond the moment it isexecuted: a regulated environment where Validation doubles as an audit trail, cross-team ownership wherePrerequisites and an explicit boundary keep the wrong team from running the wrong procedure, or a servicecomplex enough that "this does not apply here" needs to be said rather than assumed.
THIS VARIANT IS A STRICT SUPERSET OF THE LEAN ONE. The four lean sections, Purpose and Trigger, Procedure,Validation, and Review Trigger, appear here in the same order with the same headings and placeholders; fullonly adds the three sections named above. If you started lean and are growing into this, add the newsections; do not reorder or rename anything you already filled in.
A RUNBOOK IS THE PROCEDURE FOR ONE TRIGGER, NOT ONE SERVICE, AND NOT THE STANDING OPERATIONS MANUAL. Thistemplate ships the incident-scoped procedure some sources publish, one trigger, one investigation-and-remediation arc, not the larger, differently scoped service-operations manual other sources publish underthe same name, covering an entire service's overview, deployment, and disaster recovery in one document.That is a bigger scope, not a bigger size of this template; if Prerequisites and Access is growing past ahandful of lines, that is the warning sign. See runbook_companion.md sections 4 and 8.
A NOTE ON THE NAME. The Google literature this practice traces back to calls this artifact a "playbook," nota runbook; across the chapters this bundle's research read, "runbook" appears only inside a third-party casestudy, never in Google's own analytical prose. This library keeps the catalog's own name, runbook, and saysso plainly rather than implying the canon uses it. See runbook_companion.md section 2.
WHAT A RUNBOOK IS, AND IS NOTIt is the executable procedure a responder opens when a known situation occurs. It is NOT a knowledge basearticle ("check the logs and restart the service if needed" is not a runbook; it must direct a specificaction), NOT a disaster recovery plan (NIST scopes that to recovering one or more systems at an alternatefacility after a major failure, one register above this), NOT a checklist in the Gawande sense (a minimalmemory guard against omission, not a full procedural walkthrough), and NOT the standing service-operationsmanual described above. See runbook_companion.md section 8.
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into runbook_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.2. Replace each {{placeholder}} with your content. Fill Purpose and Trigger and Procedure first; everything else depends on them.3. If a section does not apply, write "N/A" and one line of why, rather than deleting it silently. "None identified" is a legitimate, honest answer for When This Does Not Apply.4. Before you ship it: self-grade against runbook_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{runbook_title}}
## Purpose and Trigger
<!-- WHAT What invokes this runbook, the system or service it is about, and the one-line job it does. Name the specific alert, condition, or page, not a general topic. WHY The Google source most of this literature traces to states that "whenever an alert is created, a corresponding playbook entry is usually created." A runbook with no stated trigger cannot be told apart from a wiki page about the service, which is the exact boundary a knowledge base article fails to cross. Deep dive: runbook_companion.md section 3 (Anatomy > Purpose and Trigger) and section 8. ASK What is the specific alert, condition, or page that opens this document? What system or service does it belong to? What is the one-line outcome this runbook exists to produce? Is it linked from the alerting tool itself? GOOD "Triggered by the SavedViewsAPI-HighErrorRate PagerDuty alert on the Saved Views API. Purpose: restore the API error rate below 1 percent within 15 minutes of the page firing." WEAK "This runbook covers the Saved Views service." (no named alert, no condition, no stated outcome; a responder cannot tell this apart from a wiki page) TRAP Naming several unrelated triggers in one document. That is usually a sign you are writing two runbooks under one title; split them. -->
{{purpose_and_trigger}}
## Prerequisites and Access
<!-- WHAT What a responder needs before they can start: access grants, roles, tooling, and any logging that must already be turned on. WHY Of the templates this bundle's research read, only Microsoft's security-operations format names this as its own fixed component: "Prerequisites: The specific requirements you need to complete before starting the investigation... logging that should be turned on and roles and permissions that are required." Everywhere else this content is folded into the first procedure steps, which is why it stays out of the lean variant: naming it separately earns its own weight only once a runbook is formal enough to also carry the other full-only sections. Deep dive: runbook_companion.md section 3 (Anatomy > Prerequisites and Access). ASK What access, role, or permission must the responder already hold? What tooling or dashboard must already be open or configured? What logging must already be turned on for the Procedure steps to work at all? PRIORITY List only what would actually block or slow down step 1 if missing. A nice-to-have is not a prerequisite. ROW HINT A good row names one concrete requirement, why the Procedure needs it, and exactly how a responder gets it before an incident, not during one. A weak row is a vague "have access" with no specifics. GOOD | prod-readonly role in the cluster IAM group | Step 1 requires kubectl access to the prod namespace | Request via the Access Portal, group saved-views-oncall | WEAK | Access to the systems | Needed | Ask your manager | TRAP Letting this section grow past a handful of lines. That is usually a sign the runbook is quietly turning into the kind of standing operations manual that documents an entire service, a different and larger artifact than this template covers. See runbook_companion.md section 8. -->
| Requirement | Why the Procedure needs it | How to get it before an incident ||---|---|---|| {{prereq_item}} | {{prereq_reason}} | {{prereq_how_to_obtain}} |
## Procedure
<!-- WHAT The steps that resolve the situation the trigger describes, starting from a known state. Say plainly, in one sentence, whether this procedure is written GENERAL (principles, changes slowly) or STEP-BY-STEP (exact commands, reduces variability), and why. WHY Google names this choice contentious inside its own organisation: some argue for general entries that change slowly, others for step-by-step entries that reduce human variability and drive down MTTR, and its own chapter calls it "a contentious topic." This template does not pick a side its source declines to pick; it asks you to state which you are writing. The same source draws one hard line regardless of style: "if your playbooks are a deterministic list of commands that the on-call engineer runs every time a particular alert fires," automate it instead of documenting it. Deep dive: runbook_companion.md section 3 (Anatomy > Procedure) and section 6. ASK Are you writing general guidance or exact commands, and why, given who runs this and how often? What is the known starting state? What are the ordered steps? Has any step here been run unchanged so many times that it is now a candidate for automation instead of documentation? PRIORITY Number the steps in the order a responder actually performs them. Start from a known state rather than "open the dashboard," which hides exactly the context a responder under pressure does not have. ROW HINT A good row names one concrete action, the exact command or decision it requires, and the result that tells the responder it worked before moving on. A weak row is a vague verb with no way to tell whether it succeeded. GOOD | 1 | Run `kubectl rollout restart deploy/saved-views-api -n prod` | Rollout status shows all pods Ready within 3 minutes | WEAK | 1 | Restart the service | It should come back | TRAP Writing "open the dashboard" or "check the logs" as a step with no named dashboard, no named log, and no signal to look for. That is a knowledge-base sentence, not a runbook step. -->
{{procedure_approach}}
| Step | Action | Expected result ||---|---|---|| {{step_number}} | {{step_action}} | {{step_expected_result}} |
## Validation
<!-- WHAT The specific signal that confirms the situation is actually resolved before you close out: a metric back to baseline, an alert clearing, a health check passing. WHY "Confirm it's fixed" cannot be checked by someone else; a named signal can. In a regulated environment this table is also your audit trail, so it should let a reviewer see what was checked, not just that something was. Deep dive: runbook_companion.md section 3 (Anatomy > Validation). ASK What signal tells you the situation is actually resolved, not just quiet? Where do you look for it? How do you tell a false negative (it looks fine but is not) from a real recovery? PRIORITY List every check a responder must clear before closing the incident, not just the first one that looks reassuring. Order them in the sequence you would actually check them. ROW HINT A good row names one specific, observable signal, the value or state that counts as pass, and exactly where to look. A weak row is "confirm it's fixed" with no location and no threshold. GOOD | Error rate | Below 1 percent for 5 consecutive minutes | Saved Views API dashboard, error-rate panel | WEAK | Everything looks fine | OK | Eyeballing the dashboard | TRAP Treating the alert clearing as the only check. An alert can clear because its threshold rearmed, not because the underlying cause is fixed; pair it with a second, independent signal. -->
| Check | What counts as pass | Where to verify it ||---|---|---|| {{validation_check}} | {{validation_signal}} | {{validation_method}} |
## Remediation and Cleanup
<!-- WHAT What to undo once the immediate situation is handled: temporary mitigations, feature flags, scaled-up capacity, anything that should not become the new permanent state. WHY This is the evidenced name for this content, drawn from a practitioner template's own section title, "Remediation and cleanup." An item left here unresolved across incidents is itself a signal: the same temporary workaround recurring is often the strongest evidence the underlying cause was never actually fixed. Deep dive: runbook_companion.md section 3 (Anatomy > Remediation and Cleanup). ASK Did you flip a switch, scale something up, or disable a check to get through the incident? What needs to be reverted, by when, and who owns reverting it? What happens if it is never reverted? PRIORITY List every temporary change made during the Procedure, even ones that feel harmless. An unlisted temporary change is the one that becomes permanent by accident. ROW HINT A good row names the specific temporary change, why it must be undone, a named owner, and the trigger or deadline for undoing it. A weak row names a worry with no owner and no deadline. GOOD | Scaled saved-views-api to 12 replicas (from 4) | Temporary capacity added to absorb the retry storm; not load-tested at this size long-term | Dana Osei | Scale back within 24 hours of the alert clearing | WEAK | Scaled things up | Fix later | Team | Someday | TRAP Leaving an item here with no owner or deadline. An undone cleanup item is the strongest predictor that the same workaround will recur, unexamined, at the next incident. -->
| Item to undo | Why it must be undone | Owner | Trigger or deadline ||---|---|---|---|| {{cleanup_item}} | {{cleanup_reason}} | {{cleanup_owner}} | {{cleanup_trigger}} |
## When This Does Not Apply
<!-- WHAT The situations that look like this trigger but are not, and what to do instead. An honestly filled "none identified" is a legitimate answer. WHY No source this bundle's research read publishes a section with this name or job; it is this template's own device for drawing the boundary explicitly rather than leaving a responder to guess it under pressure. A responder who cannot tell whether this is the right document for what they are looking at is exactly the failure the curse-of-knowledge framing describes, a runbook written by someone who cannot picture what the reader does not already know. Deep dive: runbook_companion.md section 3 (Anatomy > When This Does Not Apply) and section 8. ASK What symptom looks like this trigger but has a different cause? How would a responder tell them apart? If this is the wrong runbook, which one is right, and is it linked? PRIORITY List only genuine near-misses a responder could plausibly confuse with this trigger. A lookalike nobody would ever confuse this with does not belong here. ROW HINT A good row names the lookalike symptom, its actual cause, and points to the correct runbook or action. A weak row is vague enough to cover everything, which covers nothing. GOOD | Elevated latency with error rate still under 1 percent | Usually the shared database's connection pool, not the Saved Views API | See db-connection-pool-exhaustion runbook instead | WEAK | Other issues | Different cause | Investigate separately | TRAP Leaving this section empty by omission rather than by a stated "none identified." An empty section nobody thought about is not the same thing as one that was checked and found empty. -->
| Symptom that looks like this trigger | What it actually is | What to do instead ||---|---|---|| {{lookalike_symptom}} | {{actual_cause}} | {{redirect_action}} |
## Review Trigger
<!-- WHAT The event that would make this runbook wrong, and the named person or role expected to notice it. Not a calendar date alone. WHY A runbook fails silently: "the decay is silent until the moment of failure," and the mechanism is a process mismatch: the runbook lives in a wiki while the system lives in code, maintained by different people. No source this bundle's research read names an event-plus-owner trigger; every one reaches for a calendar, a quarterly cadence or a 90-day-untouched heuristic. This section is this bundle's own contribution, built to close that gap, not received industry practice. Deep dive: runbook_companion.md section 3 (Anatomy > Review Trigger) and section 6. ASK What system change would make this runbook wrong: a redeploy, a schema change, a dependency swap, an ownership handover? Who is on the hook to notice when that happens? What do they do about it, re-verify, update, or retire the runbook? PRIORITY Pair every date-based cadence with at least one event-based trigger. An event with no named owner is not a trigger, it is a hope. ROW HINT A good row names a specific system change, the person or role who notices it, and what they do next. A weak row is "review quarterly" with nobody named. GOOD | Saved Views API redeployed on a new runtime | On-call lead for Saved Views | Re-run this procedure against the new environment before the next on-call rotation starts | WEAK | Review this runbook quarterly | Team | Update if needed | TRAP Naming only a calendar date. A quarterly review passes every audit and can still be wrong the day an incident actually needs it; pair the date with the event that breaks the document. -->
| Event that would make this wrong | Owner who notices | What they do about it ||---|---|---|| {{review_event}} | {{review_owner}} | {{review_action}} |---title: "Saved Views Sharing: Entitlement Audit Mismatch"service_or_system: "dashboard-service - Saved Views sharing path (ViewsController + entitlement audit job)"triggering_alert: "SavedViews-EntitlementAuditMismatch (PagerDuty)"owner: "Marcus Bell (Staff Engineer, Reporting)"status: "active"last_updated: "2026-07-28"doc_type: runbooksize: fullsource_template: runbooksource_template_version: 0.1.0---
> **Worked example.** A filled `runbook`, full variant, for the Saved Views sharing path on Acme Analytics'> dashboard-service, the same feature the [`sdd`](../sdd/sdd_example.md), [`test-plan`](../test-plan/test-plan_example.md)> and [`bug-report`](../bug-report/bug-report_example.md) examples describe. Per the `standing-standards` family> contract, a runbook chains onto the shared Acme Analytics thread only loosely, because it belongs to a team> rather than to a moment in the story. This one is dated after the sharing rollout's exit review on> 2026-07-17, after the entitlement defect it responds to was fixed and reverified on 2026-07-15, and after> the risk register's last review on 2026-07-20, so everything it references had already happened.>> Read it alongside [`runbook_guide.md`](runbook_guide.md), the rubric it was graded against. Notice what the> Procedure section does: it states plainly which of its steps are exact commands and which are judgment> calls, rather than treating the whole thing as one style. All identifiers, thresholds and timestamps below> are illustrative.
# Saved Views Sharing: Entitlement Audit Mismatch
## Purpose and Trigger
Triggered by the **SavedViews-EntitlementAuditMismatch** PagerDuty alert. This alert comes from theentitlement-audit reconciliation job, a safeguard the Reporting and Platform teams added in build 2.4.0(illustrative) after DEF-2291 (a shared view's aggregate total disclosed the size of a restricted region'srevenue to a recipient who could not see the underlying rows, closed 2026-07-15). The job compares two auditstreams every 5 minutes: `shared_view_served` events emitted by ViewsController, and `permission_check_passed`events emitted by the dashboard permissions service. It pages when a `shared_view_served` event has nomatching `permission_check_passed` event for the same `request_id` inside the window.
Purpose: inside a 15-minute diagnostic window from page acknowledgment, determine whether a shared view wasactually served without a verified entitlement check. If confirmed, this is the same class of incidentDEF-2291's triage treated as a suspension event rather than a routine defect, not something to file and movepast.
## Prerequisites and Access
| Requirement | Why the Procedure needs it | How to get it before an incident ||---|---|---|| `audit-reader` role on the observability platform's Saved Views audit index | Steps 1 through 3 all query the two audit event streams | Requested through the internal access-request tool, group `reporting-saved-views-oncall`; auto-approved for anyone on the Reporting on-call schedule || Read access to the permissions service's decision log, via `permsvc-cli` | Step 4 confirms the recipient's actual entitlement scope before Security is paged | Granted with the audit-reader role above; Dana Osei co-signs the initial grant for anyone new to the Reporting rotation || Write access to the `saved_views.sharing` flag console | Step 6 disables sharing for one dashboard only, without touching phase 1 (private views) or any other dashboard | Already part of the standard Reporting on-call access bundle; nothing separate to request || The permission-persona mapping for the affected dashboard | Step 4 needs to know what the recipient was entitled to see, to size the exposure, not just that a check was missing | Query the permissions service directly with `permsvc-cli`; there is no separate document listing this |
## Procedure
Written **STEP-BY-STEP** for steps 1 through 4, the diagnostic path: a misread log line here has securityconsequences, and unlike a restart-and-watch procedure, a responder who is not a permissions-service expertcannot safely improvise the query. Step 5 is stated as a decision rather than a command on purpose: whetherto declare a security incident is a judgment call with a named owner, not a step to automate away. Step 6acts on the view already identified in step 1 and leaves no room for improvisation.
| Step | Action | Expected result ||---|---|---|| 1 | Note the `request_id`, `view_id` and `recipient_id` from the alert payload. Query the audit index for `event_type:shared_view_served AND request_id:<id>` | Exactly one matching event, timestamped within 5 minutes of the page || 2 | Query the same index for `event_type:permission_check_passed AND request_id:<id>` | Zero matching events, which is the anomaly the alert exists to catch || 3 | If step 2 returns zero, search for a `permission_check_passed` event on the same `view_id` and recipient, any `request_id`, timestamped within 2 seconds before the served event | If one is found in that window, this is almost certainly a false positive from audit-pipeline lag, not a real mismatch; skip to Validation. If nothing appears in that window, the mismatch is confirmed || 4 | Run `permsvc-cli decision --view <view_id> --user <recipient_id>` and compare its output against the view's `config` filters in the `saved_view` table | The comparison names exactly which fields and rows the recipient was, and was not, entitled to see || 5 | Page Sam Okafor's Security on-call rotation and declare a P1 security incident, stating the confirmed mismatch and the exposure from step 4 | Security acknowledges inside the standard P1 SLA and takes ownership of the incident record || 6 | Disable sharing for the affected dashboard only, by setting `saved_views.sharing` to false scoped to that dashboard ID | The dashboard's Views control shows shared views as unavailable to recipients within 60 seconds |
## Validation
| Check | What counts as pass | Where to verify it ||---|---|---|| No further EntitlementAuditMismatch page for the same dashboard | Zero recurrences in the 30 minutes following step 6 | PagerDuty incident timeline for the alert || The affected dashboard's sharing flag is off | `saved_views.sharing` reads false for that dashboard ID | Flag console, Saved Views section || Security has bounded the exposure | The incident record states which recipient and which rows were affected, and confirms no read has occurred since step 6 | Incident record in the security tracker || The cause is either fixed or explained | A ticket exists against the entitlement-check code path, or the step 3 false-positive finding is recorded with both timestamps compared | Incident record |
## Remediation and Cleanup
| Item to undo | Why it must be undone | Owner | Trigger or deadline ||---|---|---|---|| `saved_views.sharing` disabled for the affected dashboard | Sharing on this one dashboard stays off until the cause is understood; private views and every other dashboard's sharing are unaffected and stay on | Marcus Bell | Re-enable only once Security signs off and the Validation record is complete, not on a fixed date || Recipient sessions active during the exposure window | A session token issued before containment could still hold the disclosed values client-side | Marcus Bell | Force-expire the affected sessions within 4 hours of step 6 || Incident record and timeline | The DEF-2291 triage record is what let the next reporter understand Acme's severity scale for entitlement issues; this incident should leave the same kind of record | Sam Okafor | Before the incident is closed |
## When This Does Not Apply
| Symptom that looks like this trigger | What it actually is | What to do instead ||---|---|---|| The alert fires for a **private** (non-shared) view | The reconciliation job only audits `shared_view_served` events, so a private view triggering it means the job has misclassified that view's scope, not that an entitlement check was skipped | File a defect against the reconciliation job itself, and notify Marcus Bell; this is not a security incident || The alert fires within about 2 seconds of a permissions-service deploy | The two audit streams briefly fall out of order during deploy; this is what step 3's window exists to catch | Wait for the next 5-minute cycle before escalating. If the mismatch is still present after that cycle, treat it as real and continue at step 4 || ViewsController's overall error rate rises with no EntitlementAuditMismatch page | A general availability problem, unrelated to entitlement | Hand off to the dashboard-service reliability on-call as a standard incident; this runbook has nothing further to add |
## Review Trigger
| Event that would make this wrong | Owner who notices | What they do about it ||---|---|---|| The permissions service changes its audit event schema (for example, `permission_check_passed` is renamed, or its `request_id` field moves) | Dana Osei (Platform, owns the permissions service) | Re-verify the queries in steps 1 through 3 against the new schema and update this runbook before Reporting's on-call baton next changes hands || The entitlement-audit reconciliation job's 5-minute window is changed | Marcus Bell (Reporting, owns the job's configuration) | Re-time the 30-minute steady-state check in Validation to at least six times the new window, and update this runbook |Provenance
Section titled “Provenance”The reasoning, the history and every source, in the repository:
- Companion - the long-form argument: why these sections, where the sources disagree, and what the bundle refuses to claim
- History - what changed in this bundle, and when
- Research log - every source consulted, with what each one actually supports
- Catalog metadata - the machine-readable record this page is generated from
Catalog record: 7 sections across 1 format(s), methodology DevOps/SRE, typically owned by SRE / Ops.