Production Readiness Review
beta · Family standing-standards · Phase undefined · Sizes lean, full · ~5,000 tokens
The standing checklist a reviewing team consults, and a service’s own team fills in, before responsibility for that service’s production behavior changes hands, or at a periodic re-check of a service already running. Grouped by domain, each criterion carries the evidence that answers it, an owner, and a status; a completed run against one service is the record of applying the questionnaire, not a second kind of document.
The short card. Why the document is shaped this way, and the argument behind every rule here, is in
production-readiness-review_companion.md. A fully worked
instance is
production-readiness-review_example.md.
When to use
Section titled “When to use”- A service is about to change hands, for example because a development team is asking an SRE or platform function to take over standing production ownership, and you want that handoff to depend on a checklist someone can point at, not on which reviewer happens to be asking questions that day.
- More than one service, or more than one handoff, will consult this same instrument. It is a standing questionnaire, reused across services, not a document invented fresh for one occasion.
- You want the criteria a service must clear, and the evidence that would actually satisfy each one, settled before the handoff rather than argued case by case while responsibility is already changing hands.
- You are not sure yet who should hold the checklist, who decides when a finding blocks, or how often a passed review should be revisited. Reviewer and Authority and Review Cadence exist to answer exactly that.
- You need every not-applicable answer to carry a written reason, so a later reader cannot mistake a considered exclusion for an item nobody checked.
When NOT to use
Section titled “When NOT to use”- You are evaluating one specific launch, not a service’s standing ability to be run. That is a launch coordination checklist. The two share engineering-scoped roots, but this review asks whether a service can be run once responsibility changes hands, on a basis expected to outlast one release; the launch checklist asks whether one specific launch is ready to ship.
- You want a standing quality bar every unit of work is judged against, not a readiness gate for one service’s production handoff. That is a definition of done, a per-increment bar the team judges itself against every sprint. A definition of done is necessary but not sufficient input to this review; it is not a substitute for it, and this review is not a substitute for it.
- You are advising on a requested change to a system already in production, not on a service’s overall readiness to be run. That is what a change advisory board reviews; its unit is one requested change, not a service.
- You need the procedure a responder actually executes once a known situation has already happened. That is a runbook. This review checks that a runbook exists; it does not write one.
- Your organization does hardware manufacturing, and you were pointed at “Production Readiness Review” from an aerospace or defense context. NASA and the US Department of Defense publish a document under the identical name that determines whether a manufacturer is ready to produce hardware at scale, judged by production planning and supplier management. It shares nothing with this document but the name.
- You need a per-release go-live review with a rollback trigger, a business-cost-of-delay estimate, and a
post-implementation configuration check. One federal agency runs a document it also calls a Production
Readiness Review before every release, but that content belongs to the release moment; use
launch-coordination-checklistandrunbookfor it instead.
Pick a variant
Section titled “Pick a variant”Lean (five sections) is Scope and Trigger, Readiness Criteria, Not-Applicable Rule, Outcome and Sign-off, and Review Trigger. It carries the load-bearing table, the two rules that keep it honest, the recorded outcome, and the standing checklist’s own staleness trigger. At its lightest this is still a genuine review, the couple-hours version one practitioner reports having seen run, not an exemption from one.
Full (eight sections) adds Reviewer and Authority, Review Cadence, and When This Does Not Apply. Move to full when at least one of these is true:
- more than one plausible reviewer exists, so naming who conducts the review and who is authorized to say a finding blocks is worth its own section rather than left implied;
- the service is expected to be reviewed more than once, so a cadence commitment, one-time, event-driven, or continuous, is worth stating and justifying rather than left to whoever remembers next time;
- your review program covers enough services that treating all of them identically would itself be a failure mode, so a documented lighter path, with a stated minimum it still requires, is worth naming rather than left as an unspoken exception.
Every lean heading appears in full unchanged, in the same order. Growing from lean to full is additive; you never reorder or rename a section you already filled in.
Quality rubric (self-grade)
Section titled “Quality rubric (self-grade)”Score each 0, 1 or 2. Full below 11 out of 18 ships a review a service can pass while still carrying the exact kind of gap this review exists to catch: a criterion nobody can verify, a finding with no owner, or a checklist with no way to notice it has gone stale. Lean is scored on rows 1 to 6 only (see the scope table below), and below 7 of 12 a lean review can still pass a service with the same gaps.
| # | Criterion | 0 | 1 | 2 |
|---|---|---|---|---|
| 1 | Named scope and trigger | No trigger stated, or the review reads as though it applies to every service the same way with no stated boundary | A trigger is named, but a reader cannot tell what this review deliberately does not cover | The trigger matches how this review actually starts here, and at least one thing the review does not cover is named, with where that gap is reviewed instead |
| 2 | Evidence-backed criteria | Rows carry a status word and nothing else, “done,” “yes,” with no way for a second person to check it | Some rows name evidence a second person could go check; others stop at a bare confirmation | Every row’s evidence names something a second person could go check themselves, and every criterion has a named owner |
| 3 | Honest not-applicable reasons | An item is marked N/A with nothing beside it | Some N/A items carry a reason; others are bare marks | Every item marked Not Applicable carries a written reason a later reader could evaluate, not a bare checkmark |
| 4 | Findings, owned and dated | The findings table is blank, or a finding is listed with no owner and no date | Findings are listed, but some carry no owner, no date, or no stated severity | Every finding not fully met carries a named owner, a due date, and its severity, and an accepted condition is explicitly listed rather than implied |
| 5 | No silent all-clear | The outcome and findings table are simply left empty | The outcome is stated, but an empty findings table is left blank rather than addressed | Where no findings exist, the document says so explicitly rather than leaving the table blank, so a reader cannot mistake silence for an all-clear |
| 6 | Owned review trigger | Only a calendar cadence is named, or nothing at all | An event is named, but with no owner, or the next action is only “update the checklist” | A specific kind of event, an incident, a near miss, or a named worry that turned out real, is paired with a named owner whose first move is to check which existing criterion should have caught it |
| 7 | Named blocking authority (full) | No one is named as authorized to call a finding blocking, or a channel or distribution list stands in for a person | A reviewer or role is named, but no escalation path exists for when the reviewer and the reviewed team disagree | A specific role is named as authorized to call a finding blocking, and a named escalation path exists for a disagreement, scaled to the risk of what is being reviewed |
| 8 | Cadence as a position (full) | Cadence is unstated, or left as “we’ll revisit if needed” | A cadence is named, but with no stated reason tied to how this service actually changes | One of one-time, event-driven, or continuous is named deliberately, with a reason tied to how this service changes, and a stale, unexamined one-time review is named as a gap rather than left silent |
| 9 | Bounded lighter path (full) | A category is granted a blanket exemption with nothing left to check | A lighter category is named, but what it still requires is vague or unstated | A specific service or change type is named for the lighter path, and what it still requires, not only what it skips, is stated plainly |
Which rows apply to what.
| Document | Rows | Maximum | Score against |
|---|---|---|---|
| full | all 9 | 18 | 11 |
| lean | 1-6 | 12 | 7 |
Rows 7, 8, and 9 are scored only against full. Lean ships neither Reviewer and Authority, Review Cadence, nor When This Does Not Apply, so grading it on those three rows would penalize the choice of variant rather than the quality of the document.
The test behind every cell above: could someone satisfy it without improving the document? A row that counted criteria rows, named domains, or listed findings would reward padding. Every cell instead asks whether a specific piece of evidence exists, and whether a second person, not the author, could check it without asking who wrote the document.
Named anti-patterns (the usual wrecks)
Section titled “Named anti-patterns (the usual wrecks)”- Engagement that starts too late to change the design. The model’s own stated limitation is that the service is already launched and serving at scale by the time SRE engagement begins, so a reviewer’s findings can only recommend changes a team must retrofit rather than build in from the start.
- No stated goal, so no priority. A review run with no specific goal drops to the bottom of the development team’s own priority list, and gets treated as paperwork rather than as a real check.
- A defensive reviewed team that hides risk. A team that feels itself to be under review, rather than served by the review, goes on the defensive and actively hides potential risk in the system instead of surfacing it.
- A template too heavy to actually fill in. Keeping the checklist current, relevant, and usable matters more than keeping it complete; a template that has grown more complex than necessary stops getting filled in honestly.
- Shallowness that substitutes for the reviewed team’s own attention. A reviewer cannot do another team’s review for them; the value of the exercise is the time the owning team spends with its own system, not the filled-in form a reviewer produces on their behalf.
- The template treated as an inflexible rule. Building this checklist and then treating it as a hard and fast set of hoops a team must jump through, rather than a starting point tailored to what the service actually needs, misuses the instrument.
- Forgetting the humans. The template is not the point; the conversation the owning team has with its own system, prompted by the checklist, is. A review that optimizes for a completed form over that conversation has already failed at its own purpose.
- Drift after a one-time review, with no trigger to notice it. A service reviewed once and never revisited can drift out of the state the review certified without anyone noticing, which is exactly why the standing-standards family requires every member to carry its own named Review Trigger rather than relying on a calendar reminder nobody owns.
Pairing with your process
Section titled “Pairing with your process”This bundle ships in the standing-standards family alongside definition-of-done, definition-of-ready
and launch-coordination-checklist, as a tool: a standing checklist a reviewing team consults, not a
standard the reviewed team is judged against on a cadence. All four are agreed once and consulted
repeatedly rather than authored per occasion, but they answer different questions: a definition of done is
a standard a team is judged against, a definition of ready is the agreement on when a backlog item can be
pulled into a sprint, a launch coordination checklist asks whether one specific launch is ready to ship, and
this review asks whether a service can be run on a standing basis by whoever will carry its pager next.
Keep the boundaries where they belong: this review does not certify one launch, it does not tell a
responder what to type once something has gone wrong, and it checks that a runbook exists without writing
one. Where your organization also runs an Operational Readiness Review under AWS’s name for the practice, or
a per-release review under a federal agency’s use of this same document’s name, this bundle’s scope is the
service’s standing readiness; route release-moment content, rollback triggers, cost-of-delay estimates,
post-implementation configuration checks, to launch-coordination-checklist and runbook instead. Update
this file when the questionnaire itself needs to change, not every time a service is reviewed.
The artifacts
Section titled “The artifacts”production-readiness-review_template-lean.md · ~5,000 tokens
---title: "{{review_title}}"service_name: "{{service_name}}"owner: "{{owner}}"status: "{{status}}"last_updated: "{{date}}"doc_type: production-readiness-reviewsize: leansource_template: production-readiness-reviewsource_template_version: 0.1.0---
<!--LEAN PRODUCTION READINESS REVIEW. The load-bearing core: what this review covers, the criteria a servicemust clear, the rule that keeps a not-applicable answer honest, the outcome and who signed it, and whatwould make this checklist itself wrong. Five sections, because a small review program needs the readinessdiscipline without a named reviewer roster, a stated cadence, or a documented lighter path. To add Reviewerand Authority, Review Cadence, and When This Does Not Apply, seeproduction-readiness-review_template-full.md; ADD sections, never rename or reorder the ones below, becausethe full variant is a strict superset of this one.
THIS CHECKLIST IS STANDING; A COMPLETED REVIEW OF ONE SERVICE IS NOT. The instrument you are filling inbelongs to a reviewing team and is reused across services and across handoffs: "the Production ReadinessReview can be started at any point of the service lifecycle, but the stages at which SRE engagement isapplied have expanded over time." A single completed run of this document against one service, at onehandoff, is the record of applying the standing questionnaire, not a second kind of document. Update thisfile when the questionnaire itself needs to change, not every time a service is reviewed. Seeproduction-readiness-review_companion.md section 1 and section 4.
THIS IS A SERVICE'S STANDING READINESS TO BE RUN, NOT ONE LAUNCH'S READINESS TO SHIP. A production readinessreview and a launch coordination checklist share engineering-scoped roots, but they ask different questions:"a PRR is considered a prerequisite for an SRE team to accept responsibility for managing the productionaspects of a service," evaluated against a service that already exists, while a launch checklist evaluatesone specific launch before it ships. In practice a service handoff and a launch can land on the samecalendar date, and some adopters run this review at every release; that does not collapse the two documentsinto one, because what each one asks about still differs. See production-readiness-review_companion.mdsection 6 and section 8.
THE NAME COLLIDES WITH AN UNRELATED HARDWARE-MANUFACTURING REVIEW. NASA and the US Department of Defenseboth publish a "Production Readiness Review" that determines whether a manufacturer is ready to producehardware at scale, judged by production planning and supplier management rather than service reliability.It shares only the name with this document and no lineage. See production-readiness-review_companion.mdsection 2 and section 8.
THE READINESS CRITERIA DOMAINS ARE BUILT ACROSS SEVERAL SOURCES, NOT FROM ANY ONE SOURCE'S TAXONOMY. Google'sown account of this review's evaluation domains ("System architecture and interservice dependencies,""Instrumentation, metrics, and monitoring," "Emergency response," "Capacity planning," "Change management,""Performance: availability, latency, and efficiency") is licensed CC BY-NC-ND 4.0 with no derivatives, so thistemplate does not adapt that list as its own. The evidence-and-status mechanics below are adapted, withattribution, from Mercari's MIT-licensed Production Readiness Check instead. Seeproduction-readiness-review_companion.md section 3 (Anatomy > Readiness Criteria) and section 5.
WHAT THIS REVIEW IS, AND IS NOTIt is the standing questionnaire a reviewing team consults before, or during, an ownership handoff or aperiodic re-check of a service already running. It is NOT a launch coordination checklist (that reviews onelaunch, not a service's standing ability to be run). It is NOT a definition of done ("a formal descriptionof the state of the Increment," a per-work-item bar the team judges itself against every sprint, not aservice-level, often externally reviewed instrument). It is NOT a change advisory board ("delivers support toa change-management team by advising on requested changes," a portfolio-wide standing panel reviewing manychanges, not one service's readiness). It checks that a runbook exists ("Link to the troubleshootingrunbooks."; "It has OnCall playbooks.") without producing one. And it is NOT the hardware-manufacturingProduction Readiness Review named above, which shares nothing but the name. Seeproduction-readiness-review_companion.md section 8.
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into production-readiness-review_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.2. Replace each {{placeholder}} with your content. Fill Scope and Trigger and Readiness Criteria first; the rest depends on knowing what this review covers and what it found.3. If a section, or one row inside a table, does not apply, write "N/A" and the reason beside it rather than leaving it blank or deleting the section. See Not-Applicable Rule below; this is the same rule applied to the whole document.4. Before you ship it: self-grade against production-readiness-review_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{review_title}}
## Scope and Trigger
<!-- WHAT What "production" means for this review, which service or services it covers, what event starts a review, and what this review deliberately does not cover, with a pointer to where that is reviewed instead. WHY No source read for this bundle agrees on what starts a review. Google's own account is a request: "When a development team requests that SRE take over production management of a service, SRE gauges both the importance of the service and the availability of SRE teams." Mercari's is a hard gate before traffic: "This documentation describes the steps to do Production Readiness Check (PRC), which is required for all services before receiving real production traffic." AWS's Operational Readiness Review ties the trigger to a proposed change rather than a running service at all: "the ORR Lifecycle for New Service and Iterations is initiated when a new service, new feature, or architecture change is proposed." Because no trigger read here is universal, this section asks you to name your own rather than adopting any one source's. This is also where a reader who has met "Production Readiness Review" as a hardware-manufacturing milestone at NASA or the Department of Defense is told plainly that this is a different document with the same name. This section absorbs the exclusion a security review needs: "Security must have its own, in-depth, review." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Scope and Trigger) and section 6. ASK What does "production" mean for this service, in this context? Which service or services does this review cover? What event starts a review here: an ownership handoff request, a hard gate before traffic, a maturity-level promotion, a proposed architecture change, or something else? If your real trigger does not match any of these, write your own. What does this review deliberately not cover, and where is that reviewed instead? GOOD "Production means serving live customer traffic on checkout-service. This review covers checkout-service only, not its upstream inventory dependency. Trigger: the Payments team has asked the SRE guild to take over standing production ownership. This review does not cover a dedicated security assessment; that is tracked separately by the security team's own review." WEAK "This is our production readiness review for the service." (states no trigger, no scope boundary, and no exclusion; a reader cannot tell what this review actually covers or when it applies) TRAP Assuming this document is the hardware-manufacturing Production Readiness Review used in aerospace or defense contexts, which shares nothing but the name; or treating a per-release trigger as interchangeable with a standing-service trigger without saying which one this review actually uses. -->
{{production_scope_and_trigger}}
**This review does not cover:** {{out_of_scope_note}}
## Readiness Criteria
<!-- WHAT The load-bearing section: criteria grouped by domain, each carrying the evidence that answers it, an owner, and its current status. WHY No single source's domain list is copied wholesale; the starting set is drawn across sources. Monitoring appears across several checklists; capacity and performance across several more; architecture and dependencies, change management, and emergency response are named directly in Google's own account of what a PRR evaluates: "System architecture and interservice dependencies," "Instrumentation, metrics, and monitoring," "Emergency response," "Capacity planning," "Change management," "Performance: availability, latency, and efficiency." AWS groups its own questions under "Architecture," "Release quality," and "Event management." The evidence-and-status mechanics are adapted, with attribution, from Mercari's MIT-licensed checklist: "If it is satisfied, check the item in the list and provide evidence (e.g. links to tickets, screenshots, documents, test results, ...) showing that it is satisfied. If evidence is not provided the review process will take much longer." A full-size review may add whether a criterion blocks outright and the service tier that requires it; this variant keeps the smaller question, whether the evidence actually answers it. Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Readiness Criteria) and section 5. ASK For each criterion: what domain does it belong to? What evidence would actually satisfy it, not a bare confirmation? Who owns producing that evidence? What is its current status, honestly recorded? A worry-elicitation question worth asking directly: "What are you worried about?" PRIORITY Order domains the way a reader would actually need them (dependencies and architecture before monitoring and rollout, for instance). ROW HINT A good row's evidence field names something a second person could go check themselves, not a bare "done." A weak row asserts a status with no evidence anyone else could verify. GOOD | Monitoring and alerting | Does checkout-service have dashboards and alerts covering its golden signals? | Dashboard linked; alert routes to the on-call rotation and was tested against a synthetic failure last week | Payments on-call lead | Satisfied | WEAK | Monitoring | Is it monitored? | Yes | Team | Done | (no evidence anyone else could check, and "done" is not an actual status) TRAP Copying one source's domain list wholesale instead of using the set that matches what has actually gone wrong for your own service, and writing a bare confirmation where the evidence field belongs. -->
| Domain | Criterion | Evidence | Owner | Status ||---|---|---|---|---|| {{criteria_domain}} | {{criterion}} | {{criterion_evidence}} | {{criterion_owner}} | {{criterion_status}} |
## Not-Applicable Rule
<!-- WHAT An item marked not applicable in Readiness Criteria always carries a written reason beside it, never a bare mark. WHY This rule is sourced three times independently, in near-identical words. GitLab: "Leave all non-applicable items intact and add 'N/A' or reasons for why in place of the response." Mercari, MIT-licensed and adaptable with attribution: "If a specific item is not applicable to your service check the item and explain, as evidence, why it's not applicable." Federal Student Aid states the same rule as a prohibition: "No item in the PRR should only be marked "N/A" or "Not Applicable;" instead an explanation should be provided as to why a particular item does not apply." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Not-Applicable Rule). ASK Does every criterion marked "Not Applicable" in the Readiness Criteria table carry a reason beside it in this document, not only in someone's memory? If a reason is missing, is that item actually not applicable, or was it skipped? PRIORITY List the items in the order they appear in Readiness Criteria, so each reason sits beside its row. ROW HINT A good row names the criterion and a reason a later reader could check, such as where the thing it covers is reviewed instead. A weak row repeats "N/A" in the reason column. GOOD "Data recovery: marked Not Applicable in Readiness Criteria. Reason: checkout-service holds no persistent data of its own; all state lives in the shared payments ledger, reviewed under that service's own PRR." WEAK "Data recovery: N/A" (no reason; a later reader cannot tell this apart from an item nobody checked) TRAP Leaving a "Not Applicable" mark with nothing beside it. An unexplained N/A and a skipped item are indistinguishable to a later reader, which is exactly what this rule exists to prevent. -->
{{not_applicable_policy}}
**Criteria marked not applicable in this review, and why:**
| Criterion Marked Not Applicable | Reason ||---|---|| {{na_criterion}} | {{na_reason}} |
## Outcome and Sign-off
<!-- WHAT The outcome of the review, each finding not fully met, and who signed. WHY No outcome vocabulary read for this bundle is an industry standard. One practitioner's own four-state model says so directly: "These states are a recommended governance model, not a Google, AWS, or Kubernetes platform behavior." This template borrows three of that model's states, "Ready with conditions" and "Not ready," alongside an unconditional Ready state, and drops the model's fourth state, which concerns a launch whose scope or date moved rather than a service's own readiness. Each finding not met carries its own record: "For each finding, capture: severity and customer or business consequence; exact affected scope; evidence that produced the finding; remediation and named owner; due date and verification method; whether it blocks launch; exception record, if the risk is accepted temporarily." A condition accepted at sign-off is listed, owned, and dated rather than implied: "When FSA Management signs-off on the PRR, they are approving implementation even with the issues listed"; "When gaps are identified, teams either address them before launch or document an exception with an expiration date and a plan to remediate." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Outcome and Sign-off) and section 6. ASK What is the outcome: Ready, Ready with conditions, or Not ready? For every finding not fully met, what is its severity, its owner, and its due date? Where a condition is accepted rather than blocking, is that recorded with an owner and a date, or only implied? Who signed, and in what role? PRIORITY List findings that block the outcome outright before ones accepted as a time-bounded exception, so a reader sees the hard stops first. ROW HINT A good row's owner and due date name someone who could actually be asked for a status update later. A weak row states severity with no owner and no date, which nobody can be held to. GOOD | Missing load test against peak checkout traffic | High | Payments on-call lead | Two weeks from sign-off | Accepted as a time-bounded exception, tracked separately | WEAK | Some gaps | Medium | | | (no owner, no date, nothing anyone could be held to) | TRAP Reading an empty findings table as an all-clear. If no risks were identified, write "No Risks Identified" in the first row; there are always unknown risks. -->
**Outcome:** {{outcome_state}}
| Finding | Severity | Owner | Due Date | Exception (if any) ||---|---|---|---|---|| {{finding_description}} | {{finding_severity}} | {{finding_owner}} | {{finding_due_date}} | {{finding_exception}} |
| Signatory | Role | Date ||---|---|---|| {{signatory_name}} | {{signatory_role}} | {{signatory_date}} |
## Review Trigger
<!-- WHAT The event that would make this checklist itself wrong, and the named person or role expected to notice it. Not a calendar date alone. WHY Unlike this family's first two members, a condition for revising this checklist is genuinely sourced here rather than supplied by this library alone. AWS names three sources for a new question: "Real incidents that you've had in the past," "Near-misses that you've had in the past," and "The failure modes that haven't occurred, but that you're concerned about," and ties incident review directly to the checklist through a named question: "Would any ORR recommendations have reduced or avoided the impact of this event?" That question is curated by a standing role, the Ops Champion, who "challenges the team on their answers to the checklist, provides context on the adoption and prioritization of best practices, and ends up influencing everything from workload architecture to operational culture in the team." One vendor's own gate model tempers how loosely that trigger should be read: "Do not add a question merely because something once went wrong. State the failure it prevents, the evidence that answers it, and the launch types for which it applies." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Review Trigger) and section 6. ASK What specific kind of event, an incident, a near miss, or a named worry that turned out to be real, would make you revise this checklist? Who owns noticing it? When an incident does prompt a look at this checklist, which criterion should have caught it, and why did it not, before deciding the fix is a new row rather than something else entirely? PRIORITY An event with no named owner is not a trigger, it is a hope; do not list one without the other. ROW HINT A good row names a specific event, a named owner, and what they do about it: investigate which existing criterion should have caught it before deciding whether the fix is a new row. A weak row names only a calendar date. GOOD | checkout-service had an incident that Readiness Criteria's monitoring domain should have caught earlier | SRE guild lead | Investigate which specific criterion should have caught it before adding a new one | WEAK | Review this checklist annually | Team | Update if needed | TRAP Adding a new criterion for every incident without first asking which existing one should have caught it. That growth is exactly what this rule exists to prevent. -->
| Event That Would Make This Wrong | Owner Who Notices | What They Do About It ||---|---|---|| {{review_event}} | {{review_owner}} | {{review_action}} |production-readiness-review_template-full.md · ~7,550 tokens
---title: "{{review_title}}"service_name: "{{service_name}}"owner: "{{owner}}"status: "{{status}}"last_updated: "{{date}}"doc_type: production-readiness-reviewsize: fullsource_template: production-readiness-reviewsource_template_version: 0.1.0---
<!--FULL PRODUCTION READINESS REVIEW. Everything the lean variant carries, plus Reviewer and Authority, ReviewCadence, and When This Does Not Apply. Use it once more than one plausible reviewer exists, once a serviceis expected to be reviewed more than once, or once your review program covers enough services that treatingall of them identically would itself be a failure mode. To go back to the load-bearing core, seeproduction-readiness-review_template-lean.md.
THIS VARIANT IS A STRICT SUPERSET OF THE LEAN ONE. The five lean sections, Scope and Trigger, ReadinessCriteria, Not-Applicable Rule, Outcome and Sign-off, and Review Trigger, appear here in the same order withthe same headings and placeholders; full only adds Reviewer and Authority, Review Cadence, and When This DoesNot Apply. If you started lean and are growing into this, add the new sections; do not reorder or renameanything you already filled in.
THIS CHECKLIST IS STANDING; A COMPLETED REVIEW OF ONE SERVICE IS NOT. The instrument you are filling inbelongs to a reviewing team and is reused across services and across handoffs: "the Production ReadinessReview can be started at any point of the service lifecycle, but the stages at which SRE engagement isapplied have expanded over time." A single completed run of this document against one service, at onehandoff, is the record of applying the standing questionnaire, not a second kind of document. Update thisfile when the questionnaire itself needs to change, not every time a service is reviewed. Seeproduction-readiness-review_companion.md section 1 and section 4.
THIS IS A SERVICE'S STANDING READINESS TO BE RUN, NOT ONE LAUNCH'S READINESS TO SHIP. A production readinessreview and a launch coordination checklist share engineering-scoped roots, but they ask different questions:"a PRR is considered a prerequisite for an SRE team to accept responsibility for managing the productionaspects of a service," evaluated against a service that already exists, while a launch checklist evaluatesone specific launch before it ships. In practice a service handoff and a launch can land on the samecalendar date, and some adopters run this review at every release; that does not collapse the two documentsinto one, because what each one asks about still differs. See production-readiness-review_companion.mdsection 6 and section 8.
THE NAME COLLIDES WITH AN UNRELATED HARDWARE-MANUFACTURING REVIEW. NASA and the US Department of Defenseboth publish a "Production Readiness Review" that determines whether a manufacturer is ready to producehardware at scale, judged by production planning and supplier management rather than service reliability.It shares only the name with this document and no lineage. See production-readiness-review_companion.mdsection 2 and section 8.
THE READINESS CRITERIA DOMAINS ARE BUILT ACROSS SEVERAL SOURCES, NOT FROM ANY ONE SOURCE'S TAXONOMY. Google'sown account of this review's evaluation domains ("System architecture and interservice dependencies,""Instrumentation, metrics, and monitoring," "Emergency response," "Capacity planning," "Change management,""Performance: availability, latency, and efficiency") is licensed CC BY-NC-ND 4.0 with no derivatives, so thistemplate does not adapt that list as its own. The evidence-and-status mechanics below are adapted, withattribution, from Mercari's MIT-licensed Production Readiness Check instead. Seeproduction-readiness-review_companion.md section 3 (Anatomy > Readiness Criteria) and section 5.
WHAT THIS REVIEW IS, AND IS NOTIt is the standing questionnaire a reviewing team consults before, or during, an ownership handoff or aperiodic re-check of a service already running. It is NOT a launch coordination checklist (that reviews onelaunch, not a service's standing ability to be run). It is NOT a definition of done ("a formal descriptionof the state of the Increment," a per-work-item bar the team judges itself against every sprint, not aservice-level, often externally reviewed instrument). It is NOT a change advisory board ("delivers support toa change-management team by advising on requested changes," a portfolio-wide standing panel reviewing manychanges, not one service's readiness). It checks that a runbook exists ("Link to the troubleshootingrunbooks."; "It has OnCall playbooks.") without producing one. And it is NOT the hardware-manufacturingProduction Readiness Review named above, which shares nothing but the name. Seeproduction-readiness-review_companion.md section 8.
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into production-readiness-review_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.2. Replace each {{placeholder}} with your content. Fill Scope and Trigger and Readiness Criteria first; the rest depends on knowing what this review covers and what it found.3. If a section, or one row inside a table, does not apply, write "N/A" and the reason beside it rather than leaving it blank or deleting the section. See Not-Applicable Rule below; this is the same rule applied to the whole document.4. Before you ship it: self-grade against production-readiness-review_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{review_title}}
## Scope and Trigger
<!-- WHAT What "production" means for this review, which service or services it covers, what event starts a review, and what this review deliberately does not cover, with a pointer to where that is reviewed instead. WHY No source read for this bundle agrees on what starts a review. Google's own account is a request: "When a development team requests that SRE take over production management of a service, SRE gauges both the importance of the service and the availability of SRE teams." Mercari's is a hard gate before traffic: "This documentation describes the steps to do Production Readiness Check (PRC), which is required for all services before receiving real production traffic." AWS's Operational Readiness Review ties the trigger to a proposed change rather than a running service at all: "the ORR Lifecycle for New Service and Iterations is initiated when a new service, new feature, or architecture change is proposed." Because no trigger read here is universal, this section asks you to name your own rather than adopting any one source's. This is also where a reader who has met "Production Readiness Review" as a hardware-manufacturing milestone at NASA or the Department of Defense is told plainly that this is a different document with the same name. This section absorbs the exclusion a security review needs: "Security must have its own, in-depth, review." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Scope and Trigger) and section 6. ASK What does "production" mean for this service, in this context? Which service or services does this review cover? What event starts a review here: an ownership handoff request, a hard gate before traffic, a maturity-level promotion, a proposed architecture change, or something else? If your real trigger does not match any of these, write your own. What does this review deliberately not cover, and where is that reviewed instead? GOOD "Production means serving live customer traffic on checkout-service. This review covers checkout-service only, not its upstream inventory dependency. Trigger: the Payments team has asked the SRE guild to take over standing production ownership. This review does not cover a dedicated security assessment; that is tracked separately by the security team's own review." WEAK "This is our production readiness review for the service." (states no trigger, no scope boundary, and no exclusion; a reader cannot tell what this review actually covers or when it applies) TRAP Assuming this document is the hardware-manufacturing Production Readiness Review used in aerospace or defense contexts, which shares nothing but the name; or treating a per-release trigger as interchangeable with a standing-service trigger without saying which one this review actually uses. -->
{{production_scope_and_trigger}}
**This review does not cover:** {{out_of_scope_note}}
## Reviewer and Authority
<!-- WHAT Who conducts the review, who the reviewed team is, and who is authorized to say a finding blocks. WHY Two sourced positions disagree on who should hold the checklist at all. Pedro Alves argues for an outside reviewer: "To maximise the potential to identify risks, and remove biases, reviewers should be external to the team. Some degree of familiarity with the domain can be helpful, though," naming "two reviewers tends to be the sweet spot." Laura Nolan argues close to the opposite: "This is where I've very strongly argued that the team themselves should be the ones who define what their criteria are for production readiness, because they know that software best." Mercari sides with Nolan's position in practice: "Verify if each item is satisfied or not by your own team." Who decides a blocking finding is equally unsettled: AWS escalates by criticality, "Any high-criticality findings are escalated to leadership as input to a go or no-go launch decision"; Azure names a single accountable role instead, "The directly responsible individual (DRI) should make the final decisions with input from key stakeholders and technical decision-makers"; and Federal Student Aid raises the required signatory with the release's own risk, "based on the operational risk factors for the release, the CTO Enterprise Architecture and IT Planning Branch will indicate if additional sign-off by FSA Senior Management is required." Alves leaves the choice to the reviewed team itself: "It is generally up to the development team to decide whether and how to follow up on those risks." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Reviewer and Authority) and section 6. ASK Who conducts this review, named by role: someone outside the reviewed team, the reviewed team itself, or both? Who is the reviewed team? Who is authorized to say a finding blocks, and does that authority rise with the risk of what is being reviewed? Who does a blocking finding escalate to when the reviewer and the reviewed team disagree? PRIORITY List the decision authority first, then the reviewer or reviewers who feed evidence to that decision, so a reader sees immediately who can say no. ROW HINT A good row names a role, not only a person, and states what that role is accountable for. A weak row is a distribution list, a channel, or "the team," with no one actually accountable for the decision. GOOD | Decision authority | SRE guild lead | Can accept a finding as blocking; the only role authorized to sign the outcome as "Not ready" | WEAK | Reviewer | Whoever is free | (no accountable role, and no way to tell who could say no) | TRAP A review with no one authorized to say a finding blocks, so every finding is negotiable by default. Naming a reviewer with no way to escalate a disagreement is the same gap from the other side. -->
**Blocking findings escalate to:** {{escalation_authority}}
| Role | Named Holder | Responsibility ||---|---|---|| {{reviewer_role}} | {{reviewer_holder}} | {{reviewer_responsibility}} |
## Readiness Criteria
<!-- WHAT The load-bearing section: criteria grouped by domain, each carrying the evidence that answers it, an owner, its current status, and, in this variant, whether it blocks outright and the tier that requires it. WHY No single source's domain list is copied wholesale; the starting set is drawn across sources. Monitoring appears across several checklists; capacity and performance across several more; architecture and dependencies, change management, and emergency response are named directly in Google's own account of what a PRR evaluates: "System architecture and interservice dependencies," "Instrumentation, metrics, and monitoring," "Emergency response," "Capacity planning," "Change management," "Performance: availability, latency, and efficiency." AWS groups its own questions under "Architecture," "Release quality," and "Event management." The evidence-and-status mechanics are adapted, with attribution, from Mercari's MIT-licensed checklist: "If it is satisfied, check the item in the list and provide evidence (e.g. links to tickets, screenshots, documents, test results, ...) showing that it is satisfied. If evidence is not provided the review process will take much longer." Whether a criterion blocks outright is genuinely disputed: "not all issues are blockers to SRE takeover (there might be design or architectural changes that SREs recommend for service robustness that could take many months to implement)," and one vendor's own gate model separates severity into "Blocking ... Conditional ... Advisory ..." The tier that requires a criterion follows the same logic Mercari applies to its own service levels: "The main factor of choosing a Service Level is the expected SLO." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Readiness Criteria) and section 5. ASK For each criterion: what domain does it belong to? What evidence would actually satisfy it, not a bare confirmation? Who owns producing that evidence? What is its current status, honestly recorded? Does missing it block the outcome outright, or only start a conversation? What tier of service actually requires it? A worry-elicitation question worth asking directly: "What are you worried about?" PRIORITY Order domains the way a reader would actually need them (dependencies and architecture before monitoring and rollout, for instance), and within a domain, list criteria that would block the outcome outright before ones that would only start a conversation. ROW HINT A good row's evidence field names something a second person could go check themselves, not a bare "done." A weak row asserts a status with no evidence anyone else could verify. GOOD | Monitoring and alerting | Does checkout-service have dashboards and alerts covering its golden signals? | Dashboard linked; alert routes to the on-call rotation and was tested against a synthetic failure last week | Payments on-call lead | Satisfied | Blocking | All tiers | WEAK | Monitoring | Is it monitored? | Yes | Team | Done | | | (no evidence anyone else could check, no tier, and "done" is not an actual status) TRAP Copying one source's domain list wholesale instead of using the set that matches what has actually gone wrong for your own service, and writing a bare confirmation where the evidence field belongs. -->
| Domain | Criterion | Evidence | Owner | Status | Blocks? | Tier ||---|---|---|---|---|---|---|| {{criteria_domain}} | {{criterion}} | {{criterion_evidence}} | {{criterion_owner}} | {{criterion_status}} | {{criterion_blocks}} | {{criterion_tier}} |
## Not-Applicable Rule
<!-- WHAT An item marked not applicable in Readiness Criteria always carries a written reason beside it, never a bare mark. WHY This rule is sourced three times independently, in near-identical words. GitLab: "Leave all non-applicable items intact and add 'N/A' or reasons for why in place of the response." Mercari, MIT-licensed and adaptable with attribution: "If a specific item is not applicable to your service check the item and explain, as evidence, why it's not applicable." Federal Student Aid states the same rule as a prohibition: "No item in the PRR should only be marked "N/A" or "Not Applicable;" instead an explanation should be provided as to why a particular item does not apply." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Not-Applicable Rule). ASK Does every criterion marked "Not Applicable" in the Readiness Criteria table carry a reason beside it in this document, not only in someone's memory? If a reason is missing, is that item actually not applicable, or was it skipped? PRIORITY List the items in the order they appear in Readiness Criteria, so each reason sits beside its row. ROW HINT A good row names the criterion and a reason a later reader could check, such as where the thing it covers is reviewed instead. A weak row repeats "N/A" in the reason column. GOOD "Data recovery: marked Not Applicable in Readiness Criteria. Reason: checkout-service holds no persistent data of its own; all state lives in the shared payments ledger, reviewed under that service's own PRR." WEAK "Data recovery: N/A" (no reason; a later reader cannot tell this apart from an item nobody checked) TRAP Leaving a "Not Applicable" mark with nothing beside it. An unexplained N/A and a skipped item are indistinguishable to a later reader, which is exactly what this rule exists to prevent. -->
{{not_applicable_policy}}
**Criteria marked not applicable in this review, and why:**
| Criterion Marked Not Applicable | Reason ||---|---|| {{na_criterion}} | {{na_reason}} |
## Outcome and Sign-off
<!-- WHAT The outcome of the review, each finding not fully met, and who signed. WHY No outcome vocabulary read for this bundle is an industry standard. One practitioner's own four-state model says so directly: "These states are a recommended governance model, not a Google, AWS, or Kubernetes platform behavior." This template borrows three of that model's states, "Ready with conditions" and "Not ready," alongside an unconditional Ready state, and drops the model's fourth state, which concerns a launch whose scope or date moved rather than a service's own readiness. Each finding not met carries its own record: "For each finding, capture: severity and customer or business consequence; exact affected scope; evidence that produced the finding; remediation and named owner; due date and verification method; whether it blocks launch; exception record, if the risk is accepted temporarily." A condition accepted at sign-off is listed, owned, and dated rather than implied: "When FSA Management signs-off on the PRR, they are approving implementation even with the issues listed"; "When gaps are identified, teams either address them before launch or document an exception with an expiration date and a plan to remediate." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Outcome and Sign-off) and section 6. ASK What is the outcome: Ready, Ready with conditions, or Not ready? For every finding not fully met, what is its severity, its owner, and its due date? Where a condition is accepted rather than blocking, is that recorded with an owner and a date, or only implied? Who signed, and in what role? PRIORITY List findings that block the outcome outright before ones accepted as a time-bounded exception, so a reader sees the hard stops first. ROW HINT A good row's owner and due date name someone who could actually be asked for a status update later. A weak row states severity with no owner and no date, which nobody can be held to. GOOD | Missing load test against peak checkout traffic | High | Payments on-call lead | Two weeks from sign-off | Accepted as a time-bounded exception, tracked separately | WEAK | Some gaps | Medium | | | (no owner, no date, nothing anyone could be held to) | TRAP Reading an empty findings table as an all-clear. If no risks were identified, write "No Risks Identified" in the first row; there are always unknown risks. -->
**Outcome:** {{outcome_state}}
| Finding | Severity | Owner | Due Date | Exception (if any) ||---|---|---|---|---|| {{finding_description}} | {{finding_severity}} | {{finding_owner}} | {{finding_due_date}} | {{finding_exception}} |
| Signatory | Role | Date ||---|---|---|| {{signatory_name}} | {{signatory_role}} | {{signatory_date}} |
## Review Cadence
<!-- WHAT Whether this review repeats for a service that has already passed once, on what schedule, and why. WHY Four camps disagree, and none read for this bundle is universal. One-time: Google's own SRE team "assumes its production responsibilities" once "sufficient improvements are made," and Grafana's early practice reads "Once the issues have been fixed, the product has passed the PRR," though Grafana was reconsidering that design even as it wrote it down: "We're looking into a periodic and/or incremental PRR as part of our continuous product improvements." Event-driven plus a fixed annual floor: AWS requires that "In addition to the ORR performed through the SDLC process, at least annually, teams are expected to perform an ORR on their full service using a checklist tailored to that event. This helps verify that they stay up to date with new or updated best practices, and also that nothing has changed within their systems." Continuous, tied to every change: Azure gates release promotion through staged approvals ("it's important to establish release promotion as a formal change control protocol before going live. This process progresses proposed changes through various stages with quality gates"), and Cortex's own Scorecard model can "block deployment based on Scorecard scores" on every attempt. Staged across the whole lifecycle rather than fixed to one cadence: a personal ORR template states the design goal outright, "The key is making ORR a continuous exercise rather than a one-time checklist." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Review Cadence) and section 6. ASK Which of the four positions does this review take: one-time, event-driven plus an annual floor, continuous at every change, or staged across the lifecycle? Why that one, given how this service actually changes? If your team has never revisited a passed review, is that absence itself worth naming? GOOD "Event-driven plus an annual floor: checkout-service is re-reviewed whenever its payment processor changes, and at minimum once a year even if nothing else has changed, following AWS's practice above. Chosen because the service's compliance obligations shift with its processor." WEAK "We'll revisit this if it seems necessary." (names no camp, no schedule, and no reason; a reader cannot tell whether this service has ever actually been re-reviewed) TRAP Defaulting to silence on cadence rather than taking a position. A one-time review is defensible for a first handoff; an unexamined one-time review years later is a gap worth naming, not a closed question. -->
{{review_cadence_position}}
## When This Does Not Apply
<!-- WHAT Which services or changes get a lighter version of this review, and what that lighter version still requires. Not a blanket exemption. WHY No source read for this bundle exempts a service from review outright; every source that speaks to scaling depth scales the review instead of skipping it. GitLab scales per section by maturity gate. Mercari scales per row by service level, chosen from the service's own target SLO: "The main factor of choosing a Service Level is the expected SLO." AWS scales per occasion and workload type, "AWS uses different checklists for different occasions and workload types," and deliberately keeps lesser risks out of the checklist entirely to stay usable: "Medium or low risks aren't included in the ORR to keep it a lightweight process that doesn't overburden teams and reduce their agility and ability to innovate." Even at its lightest, a review that has actually run is still a review: one practitioner describes having "seen people do PRRs that just consisted of filling in a template document over a couple of hours," which is minimal, not absent. Deep dive: production-readiness-review_companion.md section 3 (Anatomy > When This Does Not Apply) and section 7. ASK Which of your services or change types would use a lighter path? What does that lighter path still require, rather than what it skips? If a service has never gone through even the lightest version of this review, is that a gap to name rather than a category to invent? PRIORITY List the smallest, lightest-touch category first, and for each one, name the minimum it still requires before naming what it is excused from. ROW HINT A good row names a specific, checkable criterion for the lighter category and states what it still requires. A weak row grants an exemption with nothing left to check. GOOD | Internal batch job with no customer-facing traffic and no persistent data of its own | Still requires Readiness Criteria's monitoring and on-call rows filled in; exempt from the load-test criterion | WEAK | Small services | Skip the review | (no criterion for "small," and no minimum kept) | TRAP Treating any of this as license for a blanket exemption. No source read for this bundle supports one; every service gets some version of this review. -->
| Service or Change Type | Lighter Check It Still Requires ||---|---|| {{lighter_check_scope}} | {{lighter_check_requirement}} |
## Review Trigger
<!-- WHAT The event that would make this checklist itself wrong, and the named person or role expected to notice it. Not a calendar date alone. WHY Unlike this family's first two members, a condition for revising this checklist is genuinely sourced here rather than supplied by this library alone. AWS names three sources for a new question: "Real incidents that you've had in the past," "Near-misses that you've had in the past," and "The failure modes that haven't occurred, but that you're concerned about," and ties incident review directly to the checklist through a named question: "Would any ORR recommendations have reduced or avoided the impact of this event?" That question is curated by a standing role, the Ops Champion, who "challenges the team on their answers to the checklist, provides context on the adoption and prioritization of best practices, and ends up influencing everything from workload architecture to operational culture in the team." One vendor's own gate model tempers how loosely that trigger should be read: "Do not add a question merely because something once went wrong. State the failure it prevents, the evidence that answers it, and the launch types for which it applies." Deep dive: production-readiness-review_companion.md section 3 (Anatomy > Review Trigger) and section 6. ASK What specific kind of event, an incident, a near miss, or a named worry that turned out to be real, would make you revise this checklist? Who owns noticing it? When an incident does prompt a look at this checklist, which criterion should have caught it, and why did it not, before deciding the fix is a new row rather than something else entirely? PRIORITY Pair every date-based cadence named in Review Cadence with at least one event-based trigger here. An event with no named owner is not a trigger, it is a hope. ROW HINT A good row names a specific event, a named owner, and what they do about it: investigate which existing criterion should have caught it before deciding whether the fix is a new row. A weak row names only a calendar date. GOOD | checkout-service had an incident that Readiness Criteria's monitoring domain should have caught earlier | SRE guild lead | Investigate which specific criterion should have caught it before adding a new one | WEAK | Review this checklist annually | Team | Update if needed | TRAP Adding a new criterion for every incident without first asking which existing one should have caught it. That growth is exactly what this rule exists to prevent. -->
| Event That Would Make This Wrong | Owner Who Notices | What They Do About It ||---|---|---|| {{review_event}} | {{review_owner}} | {{review_action}} |production-readiness-review_example.md
---title: "dashboard-service Production Readiness Review: SRE Ownership Handoff"service_name: "dashboard-service"owner: "Ines Halvorsen (Engineer, SRE)"status: "active"last_updated: "2026-08-22"doc_type: production-readiness-reviewsize: fullsource_template: production-readiness-reviewsource_template_version: 0.1.0---
> **Worked example.** A filled `production-readiness-review`, full variant, for Acme Analytics'> dashboard-service, the same service the [`sdd`](../sdd/sdd_example.md),> [`test-plan`](../test-plan/test-plan_example.md), [`bug-report`](../bug-report/bug-report_example.md) and> [`runbook`](../runbook/runbook_example.md) examples describe. Per the `standing-standards` family> contract, the chaining stays loose by design: this is the standing instrument belonging to the new SRE> function, shown as it stood at one review, not a record only this handoff could> have produced. It is dated **2026-08-22**, after every other date this thread already occupies:> DEF-2291's discovery on 2026-07-13, the Saved Views Sharing exit review and launch on 2026-07-17, the> Reporting Squad Definition of Done's amendment on 2026-07-24, the entitlement-audit runbook's first live> use on 2026-07-28, and the dashboard-scoped kill switch acceptance's own expiry on 2026-08-15. Everything> cited below had already happened.>> This review is the different scenario the family contract asks for: it is not the Saved Views Sharing> launch itself, which> [`launch-coordination-checklist_example.md`](../launch-coordination-checklist/launch-coordination-checklist_example.md)> already covers. It asks whether dashboard-service, as a whole, can be run on a standing basis by whoever> carries its pager next, not whether one change was safe to ship. Read it alongside> [`production-readiness-review_guide.md`](production-readiness-review_guide.md), the rubric it was graded> against. All names not already established in the library's Acme Analytics thread, and every threshold,> date, and status not otherwise cited, are illustrative.
# dashboard-service Production Readiness Review: SRE Ownership Handoff
## Scope and Trigger
Production, for this review, means dashboard-service serving live Saved Views and dashboard traffic to AcmeAnalytics' own customers, not a staging or internal-only environment. This review covers dashboard-serviceas a whole: the ViewsController, the `saved_view` table, and every alert and on-call surface the Reportingteam currently owns, not only the Saved Views Sharing feature that shipped on 2026-07-17. Trigger: theReporting team wants Acme Analytics' newly formed SRE function to carry dashboard-service's pager fromhere on, and this review decides whether SRE says yes. This is not the hardware-manufacturingProduction Readiness Review that NASA or the US Department of Defense would run; nothing here concernsmanufacturing or supplier readiness.
**This review does not cover:** the dashboard permissions service. The entitlement-audit alert reads the`permission_check_passed` events that service emits, but Platform owns it (Dana Osei, per[`runbook_example.md`](../runbook/runbook_example.md)) and this handoff does not move it, so Platformreviews its readiness separately.
## Reviewer and Authority
**Blocking findings escalate to:** Dana Osei (Staff Engineer, Platform), who already holds decisionauthority across every Tier 1 launch this thread has produced (see[`launch-coordination-checklist_example.md`](../launch-coordination-checklist/launch-coordination-checklist_example.md)).If Ines Halvorsen and Marcus Bell disagree about whether a finding below should block SRE's acceptance, Dana'ssign-off decides it.
| Role | Named Holder | Responsibility ||---|---|---|| Decision authority | Ines Halvorsen (Engineer, the new SRE function) | Says whether SRE accepts standing production ownership of dashboard-service, or declines; the outcome below cannot read Ready or Ready with conditions without Ines Halvorsen's signature || Second reviewer | Dana Osei (Staff Engineer, Platform) | Brings the outside-the-product-team perspective this review deliberately wants a second reviewer for, and holds the escalation authority named above || Reviewed team | Marcus Bell (Staff Engineer, Reporting) | Answers for dashboard-service's current state, and owns every finding below assigned to Reporting |
## Readiness Criteria
| Domain | Criterion | Evidence | Owner | Status | Blocks? | Tier ||---|---|---|---|---|---|---|| On-call and incident response | Can SRE actually be paged for dashboard-service today, not only Reporting? | The SavedViews-EntitlementAuditMismatch alert (`runbook_example.md`) routes only to Reporting's own PagerDuty schedule; SRE has not yet been added to that schedule or to any other dashboard-service alert | Ines Halvorsen | Not satisfied | Blocking | All tiers || Monitoring and alerting | Do dashboard-service's existing alerts cover the failure classes SRE would actually be paged for, once added to the rotation? | The entitlement-audit reconciliation job pages when a served shared view has no matching permission check inside its 5-minute window (`runbook_example.md`, Purpose and Trigger); the latency and error-rate panels for ViewsController already exist and are the same panels the Saved Views Sharing rollback triggers reference | Marcus Bell | Satisfied | Blocking | All tiers || Capacity and performance | Is dashboard-service's capacity sized against a stated ceiling SRE could alert against, not only against today's load? | No stated ceiling exists yet; the design record sets a p95 target for rendering a saved view and for the views list, but not a load ceiling to alert against before either target is missed | Dana Osei | Not satisfied | Advisory | All tiers || Deployment and rollback | Can dashboard-service's most recent class of production failure be rolled back by whoever is on call, without needing Reporting's own tribal knowledge? | The entitlement-audit runbook's own procedure disables sharing for one affected dashboard through a documented flag-console action, scoped and reversible without a Reporting engineer | Marcus Bell | Satisfied | Blocking | All tiers || Data and backup, disaster recovery | Does dashboard-service need its own backup and disaster-recovery plan, separate from the database it stores its rows in? | Not applicable; see the Not-Applicable Rule below | Dana Osei | Not Applicable | N/A | All tiers || Runbook existence | Does a runbook already exist for dashboard-service's known failure modes, one SRE could execute without Reporting on the call? | The entitlement-audit mismatch runbook (`runbook_example.md`, last updated 2026-07-28) already gives a step-by-step diagnostic and rollback procedure | Marcus Bell | Satisfied | Blocking | All tiers |
## Not-Applicable Rule
An item marked Not Applicable above carries the reason beside it here, not only a bare mark.
**Criteria marked not applicable in this review, and why:**
| Criterion Marked Not Applicable | Reason ||---|---|| Data and backup, disaster recovery: a dedicated recovery plan for dashboard-service's own storage | dashboard-service's `saved_view` table, and every other row it owns, lives in the shared main Postgres database. Backup and disaster recovery for that database is Platform's own standing responsibility, reviewed under Platform's own instrument, and is not duplicated here |
## Outcome and Sign-off
**Outcome:** Ready with conditions.
| Finding | Severity | Owner | Due Date | Exception (if any) ||---|---|---|---|---|| SRE has not yet been added to any dashboard-service PagerDuty schedule, so today only Reporting can actually be paged | High | Ines Halvorsen | 2026-09-05 | Accepted, dated: SRE shadows Reporting's on-call rotation for two full cycles before taking primary, tracked separately in the on-call rotation tool || dashboard-service has no stated load ceiling SRE could alert against ahead of missing its rendering or views-list latency target | Medium | Dana Osei | 2026-10-01 | Accepted, dated: Dana Osei owns setting a ceiling from the design record's existing p95 targets before SRE's first full quarter of ownership |
| Signatory | Role | Date ||---|---|---|| Ines Halvorsen | Decision authority, SRE | 2026-08-22 || Dana Osei | Second reviewer, Platform | 2026-08-22 || Marcus Bell | Reviewed team owner, Reporting | 2026-08-22 |
## Review Cadence
Event-driven, with a yearly backstop. dashboard-service's risk profile has already changed once without areview: on 2026-07-17 Saved Views Sharing turned a private-only surface into a shared one, and nothingstanding looked at the service as a whole. So it comes back for another look whenever it gains a newexternally reachable surface or SRE's on-call footprint for it changes, and twelve months with neitherevent brings it back anyway, to catch drift that no single event announces.
## When This Does Not Apply
| Service or Change Type | Lighter Check It Still Requires ||---|---|| A dashboard-service change that stays Tier 3 under the Platform team's own launch-coordination-checklist (reversible in a single deploy, reaches no account outside Reporting, touches no permission or billing boundary) | Readiness Criteria's on-call and monitoring rows stay current for whoever already owns the pager. Nothing else in this review applies until the change stops being Tier 3 |
## Review Trigger
| Event That Would Make This Wrong | Owner Who Notices | What They Do About It ||---|---|---|| An entitlement or permission-boundary incident recurs on dashboard-service after SRE has taken over its pager, and Readiness Criteria's on-call or monitoring domain should have caught it | Ines Halvorsen (SRE) | Check the incident's timeline against the on-call and monitoring rows first: did the SavedViews-EntitlementAuditMismatch page reach SRE, and did it fire inside the job's 5-minute window? Only if both rows held does the fix become a new row, agreed with Marcus Bell, who owns the job's configuration |Provenance
Section titled “Provenance”The reasoning, the history and every source, in the repository:
- Companion - the long-form argument: why these sections, where the sources disagree, and what the bundle refuses to claim
- History - what changed in this bundle, and when
- Research log - every source consulted, with what each one actually supports
- Catalog metadata - the machine-readable record this page is generated from
Catalog record: 8 sections across 1 format(s), methodology SRE, typically owned by SRE.