Skip to content

Runbook

beta  ·  Family standing-standards  ·  Phase undefined  ·  Sizes lean, full  ·  ~2,600 tokens

The document a responder opens when a known alert fires, so that the response does not depend on who is awake. Scoped to one situation rather than one service: the trigger that invokes it, the steps that resolve it, and the check that confirms it worked.

The short card. Why the document is shaped this way, and the argument behind every rule here, is in runbook_companion.md. A fully worked instance is runbook_example.md.

  • A known alert or condition can fire again, and the response should not depend on who is on call when it does.
  • More than one person might have to run this procedure, so the steps cannot live only in one engineer’s head.
  • The situation is scoped to one trigger: one alert, one condition, one page. If you can name it in a sentence, it belongs here.
  • You want a named signal that tells a responder the situation is actually resolved, not just quiet.
  • The work is regulated or audited, and a record of what was checked has to survive the incident.
  • You do not yet have a known trigger. If nothing has paged yet and you are documenting a service in general, write a wiki page or a design doc, not a runbook. A runbook exists because of a trigger, not a topic.
  • The steps never vary. If the on-call engineer runs the same deterministic commands every time this alert fires, that is a candidate for automation, not more documentation. Writing a better runbook for a fully deterministic procedure treats a symptom instead of the cause.
  • You need to recover a whole system at an alternate facility after a major failure. That is a disaster recovery plan, one register above what this template covers.
  • You are documenting an entire service: its overview, deployment, security posture and disaster recovery in one place. That is a standing service-operations manual, a bigger scope than this template, not a bigger size of it. Growing Prerequisites and Access past a handful of lines is usually the first sign this is happening.
  • You only need a memory guard against omission, the Gawande sense of a checklist, not a full procedural walkthrough. A checklist is deliberately smaller than a runbook.
  • The content only explains, and never directs. “Check the logs and restart the service if needed” is a knowledge base article. A runbook must go further: a specific action for a specific trigger.

Lean (four sections) is the default: Purpose and Trigger, Procedure, Validation, Review Trigger. It is close to the smallest shape actually published for this kind of document, and a runbook a responder cannot read in the middle of an incident has already failed at its one job.

Full (seven sections) adds Prerequisites and Access, Remediation and Cleanup, and When This Does Not Apply. Notice what the three additions have in common: none of them is content a responder needs while the alert is actively firing. Prerequisites is read before the incident starts, Remediation and Cleanup after it ends, and When This Does Not Apply is a scoping note rather than a step. Move to full when at least one of these is true:

  • the environment is regulated, and Validation needs to double as an audit trail;
  • more than one team could plausibly run this procedure, so an explicit boundary and named prerequisites keep the wrong team from running the wrong runbook;
  • the service is complex enough that “this does not apply here” needs to be said rather than assumed.

Every lean heading appears in full unchanged, in the same order, so growing from lean to full is additive. You never rewrite what you already agreed.

Score each 0, 1 or 2. Full below 11 out of 16 ships a document a responder cannot run without guessing, which is the one failure this artifact exists to prevent; lean is scored on five of these rows only (see the scope table below), and below 7 of 10 the same failure applies to the smaller shape.

# Criterion 0 1 2
1 Trigger is specific A general topic, not a stated alert or condition Names an alert, but no stated system or outcome Names the specific alert or condition, the system it belongs to, and a one-line outcome you could check later
2 Style is declared No statement of general or step-by-step States which, but later steps drift into the other style States which explicitly, and every step reads consistently with that choice
3 Steps resolve, not describe Steps like “open the dashboard” or “check the logs” with nothing named Some steps name a command or dashboard; others stay vague Every step names the exact action and a result a second person could check without asking the author
4 Validation resists false negatives “Confirm it’s fixed,” no named signal One named signal, and it is only the alert clearing A named signal independent of the alert clearing, with the value that counts as pass and where to look
5 Prerequisites earn their place (full) Nothing listed, or a general “have access” Items listed but not tied to a specific procedure step Every item is something that would actually block step 1, named specifically, with how to get it before an incident
6 Cleanup items are owned (full) Nothing recorded, or an unowned “fix later” Items listed, but missing an owner or a deadline Every temporary change has a named owner and a trigger or deadline for reverting it
7 Near-misses are named honestly (full) The section is missing or deleted Filled with a catch-all vague enough to describe anything Names specific lookalike symptoms and their actual cause, or states “none identified” as a considered answer
8 Trigger is event-based A calendar cadence alone, or nothing An event is named, but no owner A specific system event is named, paired with a named owner and what they do about it

Which rows apply to what.

Document Rows Maximum Score against
full all 8 16 11
lean 1-4 and 8 10 7

Rows 5, 6 and 7 are scored only against full. Lean ships none of Prerequisites and Access, Remediation and Cleanup, or When This Does Not Apply, so grading it on those rows would penalise the choice of variant rather than the quality of the document.

The test behind every cell above: could someone satisfy it without improving the document? A row that counted prerequisite items, cleanup rows, or near-misses would reward padding. Every cell instead asks whether a specific piece of evidence exists, and whether a second person, not the author, could find it.

  1. Silent decay. A runbook does not announce that it has gone stale; the system it describes keeps changing while the document does not, and the gap is invisible until the moment someone actually needs it. The fix is a real Review Trigger: a named event and a named owner, not a hope.
  2. Written for yourself, not for the reader. The author knows which dashboard, which credentials, which unstated assumptions hold. The next responder does not. The tell is a step that only makes sense to the person who wrote it; the fix is having someone who did not write the runbook try to run it.
  3. The wiki-versus-code mismatch. The runbook lives in a wiki. The system it describes lives in code. Nothing forces the two to change together, so they drift apart on different schedules maintained by different people.
  4. A knowledge-base sentence dressed up as a step. “Check the logs and restart the service if needed” explains; it does not direct. A runbook step names the exact log, the exact command, and the result that tells you it worked.
  5. Too many steps in one document. A responder under pressure cannot scan a runbook that tries to cover several unrelated situations at once. If you are listing multiple unrelated triggers, you are writing two runbooks under one title; split them.
  6. Not linked from the thing that triggers it. A correct runbook nobody can find during the incident is functionally the same as no runbook. If the alerting tool supports linking an alert to its runbook, do it.
  7. Calendar-only review, treated as sufficient. A quarterly cadence or an untouched-duration heuristic catches staleness eventually, but neither catches it at the moment the system actually changed, which is the moment it matters. Pair any calendar cadence with a named event trigger.
  8. Citing an unmeasured number as though it were a finding. Runbooks are widely asserted to shorten incidents and rarely, if ever, independently measured doing so. State the value case honestly, as an assertion, rather than attaching a specific percentage nothing behind it can support.

This bundle ships in the standing-standards family alongside definition-of-done. A definition of done is a standard a team is judged against; a runbook is an instrument a team executes. Keep the two apart: a runbook does not certify that work is finished, and a definition of done does not tell a responder what to type. Where your incident tooling supports it, link this document from the alert it answers, per the Purpose and Trigger section; that link is what turns a correct runbook into a findable one.

runbook_template-lean.md · ~2,600 tokens

---
title: "{{runbook_title}}"
service_or_system: "{{service_or_system}}"
triggering_alert: "{{triggering_alert}}"
owner: "{{owner}}"
status: "{{status}}"
last_updated: "{{date}}"
doc_type: runbook
size: lean
source_template: runbook
source_template_version: 0.1.0
---
<!--
LEAN RUNBOOK. The smallest procedure that is still a real runbook: what invokes it, what to do, how to
confirm it worked, and what would make it wrong. Four sections, because a runbook a responder cannot read in
the middle of an incident has already failed at its one job. To carry prerequisites, cleanup, and an explicit
boundary (see runbook_template-full.md), ADD sections; never rename or reorder the ones below, because the
full variant is a strict superset of this one.
A RUNBOOK IS THE PROCEDURE FOR ONE TRIGGER, NOT ONE SERVICE. It is scoped to the alert or condition that
opens it, not to everything a responder might ever need to know about the system. If you find yourself
listing several unrelated triggers here, you are probably writing two runbooks under one title. See
runbook_companion.md sections 1 and 3.
A NOTE ON THE NAME. The Google literature this practice traces back to calls this artifact a "playbook," not
a runbook; the word "runbook" appears in that canon only inside a third-party case study, never in Google's
own analytical prose. This library keeps the catalog's own name, runbook, and says so plainly rather than
implying the canon uses it. See runbook_companion.md section 2.
WHAT A RUNBOOK IS, AND IS NOT
It is the executable procedure a responder opens when a known situation occurs. It is NOT a knowledge base
article ("check the logs and restart the service if needed" is not a runbook; it must direct a specific
action), NOT a disaster recovery plan (that recovers one or more systems at an alternate facility after a
major failure, one register above this), NOT a checklist in the Gawande sense (a minimal memory guard against
omission, not a full procedural walkthrough), and NOT the standing service-operations manual some sources
publish under the same name, which covers an entire service rather than one trigger. See
runbook_companion.md section 8.
HOW TO FILL THIS IN
1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into
runbook_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For
tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.
2. Replace each {{placeholder}} with your content. Fill Purpose and Trigger and Procedure first; everything
else depends on them.
3. If a section does not apply, write "N/A" and one line of why, rather than deleting it silently.
4. Before you ship it: self-grade against runbook_guide.md, then DELETE every HTML comment. They are
guidance, not content.
-->
# {{runbook_title}}
## Purpose and Trigger
<!-- WHAT What invokes this runbook, the system or service it is about, and the one-line job it does. Name
the specific alert, condition, or page, not a general topic.
WHY The Google source most of this literature traces to states that "whenever an alert is created, a
corresponding playbook entry is usually created." A runbook with no stated trigger cannot be told
apart from a wiki page about the service, which is the exact boundary a knowledge base article
fails to cross. Deep dive: runbook_companion.md section 3 (Anatomy > Purpose and Trigger) and
section 8.
ASK What is the specific alert, condition, or page that opens this document? What system or service
does it belong to? What is the one-line outcome this runbook exists to produce? Is it linked from
the alerting tool itself?
GOOD "Triggered by the SavedViewsAPI-HighErrorRate PagerDuty alert on the Saved Views API. Purpose:
restore the API error rate below 1 percent within 15 minutes of the page firing."
WEAK "This runbook covers the Saved Views service." (no named alert, no condition, no stated outcome;
a responder cannot tell this apart from a wiki page)
TRAP Naming several unrelated triggers in one document. That is usually a sign you are writing two
runbooks under one title; split them. -->
{{purpose_and_trigger}}
## Procedure
<!-- WHAT The steps that resolve the situation the trigger describes, starting from a known state. Say
plainly, in one sentence, whether this procedure is written GENERAL (principles, changes slowly)
or STEP-BY-STEP (exact commands, reduces variability), and why.
WHY Google names this choice contentious inside its own organisation: some argue for general entries
that change slowly, others for step-by-step entries that reduce human variability and drive down
MTTR, and its own chapter calls it "a contentious topic." This template does not pick a side its
source declines to pick; it asks you to state which you are writing. The same source draws one
hard line regardless of style: "if your playbooks are a deterministic list of commands that the
on-call engineer runs every time a particular alert fires," automate it instead of documenting it.
Deep dive: runbook_companion.md section 3 (Anatomy > Procedure) and section 6.
ASK Are you writing general guidance or exact commands, and why, given who runs this and how often?
What is the known starting state? What are the ordered steps? Has any step here been run
unchanged so many times that it is now a candidate for automation instead of documentation?
PRIORITY Number the steps in the order a responder actually performs them. Start from a known state
rather than "open the dashboard," which hides exactly the context a responder under pressure does
not have.
ROW HINT A good row names one concrete action, the exact command or decision it requires, and the
result that tells the responder it worked before moving on. A weak row is a vague verb with no
way to tell whether it succeeded.
GOOD | 1 | Run `kubectl rollout restart deploy/saved-views-api -n prod` | Rollout status shows all
pods Ready within 3 minutes |
WEAK | 1 | Restart the service | It should come back |
TRAP Writing "open the dashboard" or "check the logs" as a step with no named dashboard, no named
log, and no signal to look for. That is a knowledge-base sentence, not a runbook step. -->
{{procedure_approach}}
| Step | Action | Expected result |
|---|---|---|
| {{step_number}} | {{step_action}} | {{step_expected_result}} |
## Validation
<!-- WHAT The specific signal that confirms the situation is actually resolved before you close out: a
metric back to baseline, an alert clearing, a health check passing.
WHY "Confirm it's fixed" cannot be checked by someone else; a named signal can. In a regulated
environment this table is also your audit trail, so it should let a reviewer see what was
checked, not just that something was. Deep dive: runbook_companion.md section 3 (Anatomy >
Validation).
ASK What signal tells you the situation is actually resolved, not just quiet? Where do you look for
it? How do you tell a false negative (it looks fine but is not) from a real recovery?
PRIORITY List every check a responder must clear before closing the incident, not just the first one
that looks reassuring. Order them in the sequence you would actually check them.
ROW HINT A good row names one specific, observable signal, the value or state that counts as pass,
and exactly where to look. A weak row is "confirm it's fixed" with no location and no threshold.
GOOD | Error rate | Below 1 percent for 5 consecutive minutes | Saved Views API dashboard, error-rate
panel |
WEAK | Everything looks fine | OK | Eyeballing the dashboard |
TRAP Treating the alert clearing as the only check. An alert can clear because its threshold rearmed,
not because the underlying cause is fixed; pair it with a second, independent signal. -->
| Check | What counts as pass | Where to verify it |
|---|---|---|
| {{validation_check}} | {{validation_signal}} | {{validation_method}} |
## Review Trigger
<!-- WHAT The event that would make this runbook wrong, and the named person or role expected to notice
it. Not a calendar date alone.
WHY A runbook fails silently: "the decay is silent until the moment of failure," and the mechanism is
a process mismatch: the runbook lives in a wiki while the system lives in code, maintained by
different people. No source this bundle's research read names an event-plus-owner trigger; every
one reaches for a calendar, a quarterly cadence or a 90-day-untouched heuristic. This section is
this bundle's own contribution, built to close that gap, not received industry practice. Deep
dive: runbook_companion.md section 3 (Anatomy > Review Trigger) and section 6.
ASK What system change would make this runbook wrong: a redeploy, a schema change, a dependency
swap, an ownership handover? Who is on the hook to notice when that happens? What do they do
about it, re-verify, update, or retire the runbook?
PRIORITY Pair every date-based cadence with at least one event-based trigger. An event with no named
owner is not a trigger, it is a hope.
ROW HINT A good row names a specific system change, the person or role who notices it, and what they
do next. A weak row is "review quarterly" with nobody named.
GOOD | Saved Views API redeployed on a new runtime | On-call lead for Saved Views | Re-run this
procedure against the new environment before the next on-call rotation starts |
WEAK | Review this runbook quarterly | Team | Update if needed |
TRAP Naming only a calendar date. A quarterly review passes every audit and can still be wrong the
day an incident actually needs it; pair the date with the event that breaks the document. -->
| Event that would make this wrong | Owner who notices | What they do about it |
|---|---|---|
| {{review_event}} | {{review_owner}} | {{review_action}} |

The reasoning, the history and every source, in the repository:

  • Companion - the long-form argument: why these sections, where the sources disagree, and what the bundle refuses to claim
  • History - what changed in this bundle, and when
  • Research log - every source consulted, with what each one actually supports
  • Catalog metadata - the machine-readable record this page is generated from

Catalog record: 7 sections across 1 format(s), methodology DevOps/SRE, typically owned by SRE / Ops.