Test Summary Report
beta · Family qa-docs · Phase develop · Sizes lean, full · ~4,300 tokens
The retrospective document that closes a testing effort: what was tested and on which build, what the testing found, what was deliberately or accidentally not tested, and whether the result clears the bar the test plan set. The closing bookend to the test plan, and the artifact a dashboard cannot produce, because the counts are computed and the judgement is not. Whether the genre should exist at all is contested: a named school of practitioners petitioned to have the international standard that defines it withdrawn, and this bundle teaches that dispute rather than settling it.
The short card. Why the document is shaped this way, and the argument behind every rule here, is in
test-summary-report_companion.md. A fully worked instance is
test-summary-report_example.md.
When to use
Section titled “When to use”Read this first, because it changes what the card is claiming. The international standard that defines this document type is one a named school of practitioners campaigned to have withdrawn. In 2014 a petition organised by the International Society for Software Testing,, asked ISO to suspend publication of Parts 4 and 5 of ISO/IEC/IEEE 29119 and to withdraw Parts 1 to 3 - which includes the part that defines this report. Michael Bolton, a leading voice of the opposition, described that standard as an “overstructured process model, focused on relentless, ponderous, wasteful bureaucracy and paperwork, with negligible content on actual testing”. So nothing on this card says the test report is settled practice. It says how to write one worth reading if you are writing one, and it takes the opposition’s sharpest charge as the thing to design against rather than the thing to ignore. Both camps: companion section 6.
And one thing about where the section list came from. That standard is sold, not published. What this bundle read of it is its Clause 3 definition and its table of contents, which is a list of subclause headings and not the requirements underneath them. The sections you are about to fill come from ISTQB’s freely published foundation-level syllabus, from real filled reports, and from published templates - not from the standard’s unread text. The full retrieval position is companion section 1, and it is worth two minutes before you cite anything from this bundle to anyone.
Write one when:
- A testing effort has ended - a release, a cycle, a level, a milestone - and someone has to say whether the bar the test plan set was cleared.
- The result has to travel past the people who watched it happen: to another team, to a sponsor, to a customer, to the version of your own team that exists in six months.
- A sign-off, an audit, a certification or a regulatory submission is downstream, and this report is part of the evidence rather than a courtesy.
- Something was not tested, and a reader who is not told will assume it passed.
- Residual risk is going to production and somebody needs to be on record accepting it, by name.
- The engagement is closing or the team is dispersing, so this is the durable record of what was actually verified and on which build.
When NOT to use
Section titled “When NOT to use”Write something else if:
| You actually need | Because |
|---|---|
| a test plan | the testing has not happened yet. The plan is prospective: scope, approach, entry and exit criteria, agreed before the work starts. This report is its closing bookend and is retrospective. The dependency runs one way and it is structural: this report’s Evaluation Against Exit Criteria section has content only because a plan wrote criteria. Reaching for this template before the cycle means you wanted that one |
| a bug report | one verification failed and an engineer needs to reproduce it. That is one document per defect, carrying steps, environment, and expected against actual. This report counts and characterizes defects by severity and says which remain open; it links to them and never accumulates their reproductions |
| a test status or progress report | testing is still running. A progress report is produced at regular intervals against the plan’s baseline so that somebody can intervene while intervening is still possible; this document is produced once, at the end, and judges. If you are writing the same document every Friday you are writing progress reports, and the completion report is still owed at the end. This library ships no template for one, and its status-report is not it: that is the project-level periodic update for a different reader |
| nothing | a pipeline already publishes the numbers and nobody needs a judgment on top of them. If the dashboard is live, the readers watched it all cycle, and nobody will be asked later to defend the release decision, then a message carrying the exit-criteria call and the residual-risk list does the entire job. A report that only restates the dashboard is strictly worse than the dashboard, because it is also out of date the moment you send it |
The threshold, stated plainly. A tool vendor’s own framing of when the formal document earns its cost is the one this bundle adopts: it becomes useful when results need to be shared “beyond a single iteration or the immediate team”, or to support an audit, a customer sign-off, a regulatory review or a contract. Below that line, do not write one. Writing it anyway is how the document type earned its reputation.
The ship call has no section of its own
Section titled “The ship call has no section of its own”If you are looking for a Release Recommendation heading, it is deliberately not here, and knowing where the call went is the difference between filling this template correctly and thinking a section is missing.
The ship sentence goes at the end of Evaluation Against Exit Criteria, underneath the criteria it rests on, in this shape: which criteria were met, which were not, and therefore ship, do not ship, or ship with these named conditions carried by this named person. One line on why: a recommendation with nothing above it is an opinion, and the same sentence under a graded criteria table is a conclusion.
This library considered a standalone section, tested the idea against its own research, and dropped it - neither readable structural source carries one, and the real filled reports in the corpus do not either, with one certification-lab exception the companion explains. The evidence and the change of mind are in companion section 3.
What this does not change: the test plan card promises that what you found and whether you are shipping belong in a report written afterwards. They do. The verdict simply does not get a heading of its own.
Pick a variant
Section titled “Pick a variant”Lean (six sections) is the default: Scope and What Was Tested, Execution Summary, Defects, Deviations from Planned Testing, Evaluation Against Exit Criteria, and Residual Risk and What Was Not Tested. That is what closing a release actually takes - what it was, what ran, what broke, what did not happen as planned, whether the bar was cleared, and what is being carried forward.
Full (nine sections) adds Impediments and Blocked Progress, Test Deliverables and Reusable Assets, and Lessons Learned. Reach for it when at least one of these is true:
- the reader is an auditor, a regulator, a customer or an accreditation body, and completeness is part of the evidence rather than a style preference;
- testing was done by a vendor, a lab, or more than one team, so who tested what belongs in the record;
- the assets outlive the release and the next team inherits the suites, harnesses, fixtures and environments;
- the reader was not there, and this report is the durable record rather than a step in a conversation that continues;
- no retrospective is going to happen, so this is the only place the learning survives.
Nesting is strict: every lean heading appears in the full variant with the same name in the same order, and full only adds. A lean report grows into a full one without rewriting a line of what is already there.
Two sections are never the ones you cut, whichever variant you pick: Evaluation Against Exit Criteria and Residual Risk and What Was Not Tested. Those are the two your tooling cannot produce, and they are the reason the document exists at all. If the whole budget is one paragraph, write those two and link the dashboard for the rest.
Quality rubric (self-grade before you send it)
Section titled “Quality rubric (self-grade before you send it)”Score each 0, 1 or 2. Under 13 out of 18 and what you have is a dashboard with a cover page: the reader gets counts they already had, and none of the three judgments only a person can make - what was not tested, what risk is being accepted, and whether the bar was cleared.
| # | Criterion | 0 | 1 | 2 |
|---|---|---|---|---|
| 1 | Build is named | No version, build or commit anywhere in the document | A version, but no environment and no reference to the plan it answers | A reader six months from now can name the build, the environment and configuration, the period tested, and the plan whose criteria this grades |
| 2 | Claim is bounded | Nothing says what the result does not cover | A scope line that restates the feature name and stops | The report says what configuration and version the result applies to, and names at least one thing a reader must not read it as covering |
| 3 | Coverage claim honest | “Coverage” is a pass percentage wearing a different word | Coverage is stated, but only against the cases in your own suite | Coverage is stated against something outside the suite - risks, requirements, areas, code - and where breadth was traded for depth, the report says which |
| 4 | Defects are summarized | A tracker export pasted in, or a single total | Counts by severity | Severity and pattern, what is still open, who owns each open item, and links out to the reports rather than their reproductions retyped here |
| 5 | Deviations are filled | Empty, or “none” on a cycle that visibly slipped | Says the plan changed | Names what was planned and not done, what was done instead, and why - enough that the counts above can be read correctly rather than taken at face value |
| 6 | Criteria graded individually | Prose about how the cycle went | One summary judgment covering all criteria at once | Every exit criterion the plan set appears with met or not met and the evidence for it, including the ones that were not met |
| 7 | Conclusion follows criteria | No conclusion, or one that contradicts the grades above it | A conclusion is stated, with nothing tying it to any criterion | The ship sentence sits under the graded criteria, names its conditions and who is accepting them, and a reader can point at the grade that would flip it |
| 8 | Residual risk owned | Not mentioned, or written as an apology for running out of time | Untested areas listed, with no consequence attached | Each item says what could go wrong, who is accepting it and on what basis, and Undetermined appears wherever that is the honest answer |
| 9 | Adds beyond the dashboard | Everything here was already in the tool | One section carries something the tool could not compute | A reader who had the dashboard open all cycle still learns something, and could point at where |
The test behind every cell above: could someone satisfy it without improving the report? A row that counted defects, or counted exclusions, would reward padding, and a forty-row defect table nobody triaged would score the same as one honest paragraph. Every cell instead asks whether a specific piece of evidence is there and whether a second person, not the author, could find it.
Named anti-patterns (the usual wrecks)
Section titled “Named anti-patterns (the usual wrecks)”- The dashboard with a cover page. Execution Summary filled in, everything else thin or absent. The tell is that a reader could have learned more by opening the tool, and that your document was out of date before it was sent. This is the failure the whole document type stands accused of, so it is the one to check for first. Fix: write only what the tool cannot compute, and if that leaves nothing, do not write the report.
- Coverage claimed, pass rate measured. “96 percent coverage” where what was measured is the proportion of executed tests that passed. These are two different claims about two different things: pass rate is a property of the tests you chose to run, and coverage is a claim about the system, stated against risks, requirements, areas or code. A suite exercising a tenth of the product can post a perfect pass rate, and a reader who reads that as coverage treats untested ground as cleared. The security assessment in this bundle’s research shows the honest form: it discloses that “codebase coverage emphasized breadth instead of depth” and that portions outside the control areas received minimal to no coverage. Fix: name the denominator, and where the denominator is your own suite, say so in the same sentence.
- Silence read as coverage. Areas nobody tested go unmentioned, so a reader assumes they passed. Related to the one above and not the same: that one overclaims, this one omits, and omission is the easier of the two to commit by accident. Fix: an explicit not-tested list with a reason against each item, sitting in the same section as the residual risk it creates.
- Complete and empty. Every heading present, nothing decided, because the template asked for a heading. James Christie’s critique aimed squarely at this document type describes the result as “a collection of metrics that say nothing about the quality of the product”, and he is describing reports that complied with the standard rather than reports that ignored it. Fix: delete any section that decides nothing, and make the exit-criteria evaluation carry an actual “therefore”.
- The verdict with nothing above it. A confident release call resting on numbers that do not support it, or on no criteria at all. This is exactly the failure that the missing Release Recommendation heading is designed to prevent, and moving the sentence does not prevent it by itself. Fix: grade the criteria first; the call is the last line of that section, never the first line of the report.
- Residual risk written as apology. “Unfortunately we ran out of time for the migration path.” That is a schedule confession, not a risk statement: it explains why a gap exists and says nothing about what the gap could cost or who is carrying it. Fix: name the exposure, the person accepting it, and the basis for accepting it. The regulated form of this is a written risk-based rationale for every defect left unfixed, and it is a good shape to borrow even when nobody is making you use it.
Pairing with a skill
Section titled “Pairing with a skill”pairs_with: [], and the empty list is a finding rather than an omission. There is no testing or QA skill
in the pm-skills library: none of its tracked skills produces a test plan, a test case, a bug report or a
test report, recorded as finding EC-4 (pm-skills covers no testing or QA work) in this repository’s
STATE.md. The test plan bundle can honestly claim
deliver-edge-cases, because an edge-case catalog is an input to planning coverage. A report of testing
already finished has no such input to take, so this bundle declines the pairing rather than borrowing its
sibling’s. Everything here is filled by hand, from your test plan, your defect tracker and your tool’s run
data.
The artifacts
Section titled “The artifacts”test-summary-report_template-lean.md · ~4,300 tokens
---title: "{{release_or_feature}} Test Summary Report"release_or_feature: "{{release_or_feature}}"build_or_version: "{{build_or_version}}"test_plan_ref: "{{test_plan_ref}}"report_author: "{{report_author}}"test_period: "{{test_period}}"status: "{{status}}"last_updated: "{{date}}"doc_type: test-summary-reportsize: leansource_template: test-summary-reportsource_template_version: 0.1.0---
<!--LEAN TEST SUMMARY REPORT. The smallest report that is still a real report: what was tested and on whichbuild, what ran, what broke, what did not happen as planned, whether the bar the test plan set was cleared,and what is being carried forward untested or unfixed. Use it to close a release or a test cycle for readerswho were broadly in the room. To grow it into the accountability-grade report (seetest-summary-report_template-full.md), ADD sections; never rename or reorder the ones below, because the fullvariant is a strict superset of this one.
IF THE DASHBOARD ALREADY SAYS IT, DO NOT RETYPE IT. The criticism this document type has never fully answeredis that it can satisfy every heading and still say nothing about the quality of the product. Your test toolcomputes the counts better than this page can. What it cannot do is say what was not tested and why, whatresidual risk somebody is accepting, and whether the result clears the bar. Spend your time there; where asection adds nothing to the tool, keep it to a line and a link. See test-summary-report_companion.mdsections 6 and 7.
SOMETIMES THE HONEST ANSWER IS NOT TO WRITE ONE. If everyone who would read this sat in the same standup allcycle, a short message carrying the exit-criteria call and the residual-risk list does the whole job. Writethe document when the result has to travel: beyond one iteration, beyond the immediate team, or into anaudit, a sign-off or a contract.
WHAT A TEST SUMMARY REPORT IS, AND IS NOTIt is the retrospective document that closes a testing effort: what was tested, what the testing found, whatwas deliberately or accidentally not tested, and whether the result clears the bar the test plan set. It isNOT a test plan (that is prospective, written before), NOT a test status or progress report (that is producedat intervals while testing is still running), NOT your test tool's exported run or your CI dashboard (thosecompute the counts; this document carries the judgment), and NOT a bug tracker (one failing verification isone bug report). See test-summary-report_companion.md section 8.
IF YOU ARE LOOKING FOR A "RELEASE RECOMMENDATION" SECTION, IT IS DELIBERATELY NOT HERE. The ship call belongsat the end of Evaluation Against Exit Criteria, underneath the criteria that justify it. A recommendationwith nothing above it is an opinion; the same sentence under a graded criteria table is a conclusion. Thislibrary considered a standalone section and dropped it for want of evidence; the reasoning is intest-summary-report_companion.md section 3.
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into test-summary-report_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.2. Replace each {{placeholder}} with your content. Fill Scope and What Was Tested first, then Evaluation Against Exit Criteria; everything else exists to make that evaluation readable.3. If a section does not apply, write "N/A" and one line of why, rather than deleting it silently.4. Before you share it: self-grade against test-summary-report_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{release_or_feature}} Test Summary Report
## Scope and What Was Tested
<!-- WHAT What this report covers: the build or version under test, the environment and configuration, the test plan it answers to, the period it covers, and the boundary of the claim. WHY Every number below is meaningless without the version that produced it, and the version is the field the one filled report in this bundle's research omits, which names no build anywhere, so nobody reading it later can tell what it was about. Bounding the claim is the other half: a report that does not say what it does not cover gets read as covering everything. Deep dive: test-summary-report_companion.md section 3 (Anatomy > Scope and What Was Tested). ASK Which build, version or commit was tested, in which environment and configuration? Which test plan, and which of its criteria, does this report answer? What period does it cover, and who tested? What does the result explicitly not apply to? GOOD "Covers Claims Intake R7.2, build 7.2.14, on pre-production with the fnol_v4 flag on, tested 2026-03-02 to 2026-03-13 by Nadia Okonkwo and Fabiola Reyes against the R7.2 test plan. Results apply to that build and configuration only. The broker portal integration runs on its own release train and was not exercised here." WEAK "Testing of the claims release. Ran on staging." (no build, no dates, no plan reference and no boundary, so a reader can tell neither what the report is about nor what it leaves out) TRAP Omitting the build or version because everyone currently knows which one it was. In six months this report is the only record, and nobody will. -->
{{scope_and_what_was_tested}}
## Execution Summary
<!-- WHAT The counts - planned, executed, passed, failed, blocked, not run - per area or suite, with coverage stated against something meaningful, and a line below the table saying when the counts were taken and where the live results live. WHY This is the section your tooling can fill, and the only one. Test management tools and CI plugins compute exactly these numbers and draw them better than a document can, so a report that is mostly this section is a dashboard with a cover page, and out of date the moment it is written. Put the counts here so the judgment below has something to stand on, and put the judgment where the criteria are. Deep dive: test-summary-report_companion.md section 3 (Execution Summary) and section 6. ASK How many cases were planned, and how many actually ran? What failed, what was blocked, what never ran at all? What is coverage a fraction of - risk areas, requirements, the plan's own priorities? When were these counts taken, and where does a reader go for the detail? PRIORITY Order rows by the plan's own risk ranking, highest risk first, so a reader scanning the top of the table is reading about the areas that mattered most. Counts are a snapshot: state the date and time they were taken, because they moved the day after. ROW HINT A good row names an area the plan recognizes, gives every count including the ones that did not run, and says what the coverage figure is a fraction of. A weak row is a suite name and a pass percentage. GOOD | Payout calculation (High risk) | 48 | 46 | 42 | 3 | 1 | 2 | 46 of 48 planned cases; all 9 High-risk scenarios executed | WEAK | Regression | | | | | | | 97 percent | TRAP Leading with a pass rate. "97 percent passed" hides whether the 3 percent was a cosmetic label or the permission check, and a percentage is an input to the exit-criteria evaluation below, never a substitute for it. -->
| Area or suite | Planned | Executed | Passed | Failed | Blocked | Not run | Coverage and notes ||---|---|---|---|---|---|---|---|| {{area}} | {{planned}} | {{executed}} | {{passed}} | {{failed}} | {{blocked}} | {{not_run}} | {{coverage_note}} |
{{execution_summary_notes}}
## Defects
<!-- WHAT What was found, at what severity, what is fixed, and what is still open at the moment of writing - summarized and linked, never transcribed. WHY Open defects are the content here; a list of what you fixed is history. The strongest real reports break the found defects down by something a team can act on, such as root cause or component, and give every open one a severity and a stated reason it is not being fixed. Deep dive: test-summary-report_companion.md section 3 (Defects). ASK How many defects were found, at what severities, and what pattern do they show? How many remain open, and which are in scope for the release decision? For each open defect, what would a user experience, why is it not fixed, and who accepted that? What is fixed but not yet verified? PRIORITY Severity first, highest first, and open before closed within a severity. Closed defects are a summary line above the table; every open defect in scope gets a row of its own. ROW HINT A good row identifies the defect, gives severity and status, says what a user would experience, and names both the reason it stands and the person who accepted it. A weak row is a ticket number and a title. GOOD | CLM-4471 | Sev-2 | Open | Payout total rounds down by one cent on multi-currency claims | Fix lands in 7.2.15; Ellen Wray accepted the cent-level variance for the two-week window and finance was notified 2026-03-11 | WEAK | CLM-4471 | High | Open | Rounding bug | Will fix | TRAP Pasting the tracker in. Forty rows of reproduction steps make a worse defect list than the tracker itself and bury the three that matter. Summarize, link out, and spend the words on the open ones. -->
{{defect_summary}}
| Defect | Severity | Status | Impact if it ships | Disposition: why it is open, who accepted it ||---|---|---|---|---|| {{defect_id}} | {{severity}} | {{defect_status}} | {{defect_impact}} | {{defect_disposition}} |
## Deviations from Planned Testing
<!-- WHAT What the plan said would happen, what actually happened, and why the difference - in scope, schedule, depth, environment or data. WHY Deviations are what make the numbers above interpretable: a 98 percent pass rate over half the planned depth is a different result from the same rate over all of it. The evidence for keeping this section is one-sided - both structural sources this bundle could read carry it, both real templates ask for it, and the one real filled report gives it a named subsection - which is why it is in the lean variant, against this library's own earlier internal spec. Deep dive: test-summary-report_companion.md section 3 (Deviations from Planned Testing) and section 4. ASK What did the plan say, and what actually happened? Which areas were tested less deeply than planned, or not at all? What changed in scope, schedule, environment or data mid-effort, and who agreed it? Who knew at the time, and who is finding out from this document? GOOD "The plan scheduled two full regression passes; one ran, because pre-production was rebuilt in week two. Fraud-scoring cases were executed against synthetic data rather than the anonymized production extract, which never arrived. Both were raised with Ellen Wray on 2026-03-06; neither changed the agreed exit criteria." WEAK "Some testing was descoped due to time." (which testing, how much, whose decision, and what does it do to the counts above) TRAP Writing this section after the release decision is made. A deviation that surfaces for the first time in the report is news, and news arriving this late damages trust in everything else on the page. Raise them as they happen and record them here. -->
{{deviations_from_planned_testing}}
## Evaluation Against Exit Criteria
<!-- WHAT Each exit criterion the test plan set, whether it is met, and on what evidence - then the conclusion the testing reaches about releasing. WHY This is the load-bearing section, and the evaluation against agreed criteria is the act that makes this a report rather than an export. Neither readable structural source gives the release call a heading of its own, and neither does this template: a recommendation with nothing above it is an opinion, while the same sentence underneath a graded criteria table is a conclusion. Write the table, then write the sentence: met, not met, and therefore. Deep dive: test-summary-report_companion.md section 3 (Evaluation Against Exit Criteria), which records why this library considered a standalone Release Recommendation section and dropped it for want of evidence. ASK What exactly were the exit criteria, and who agreed them and when? Is each met, partially met or not met, and on what evidence? Where one is not met, what does releasing anyway cost? What does the testing therefore conclude, and who owns the decision this conclusion feeds? PRIORITY One row per criterion, in the plan's own words and the plan's own order. Never rewrite a criterion to match the result. If the plan set no exit criteria, do not invent them after the fact: say so plainly in the conclusion and state the bar you are applying instead. ROW HINT A good row quotes the criterion, gives a plain verdict (met, partially met, not met), points at the evidence, and where it is not met says what that costs. A weak row is a criterion and a tick. GOOD | All 9 High-risk payout scenarios executed with no Sev-2 or worse left open | Not met | CLM-4471 (Sev-2) open, see Defects | Multi-currency claims round down by one cent until 7.2.15 | WEAK | Quality is acceptable | Met | Testing complete | | TRAP The verdict with nothing above it. A confident release sentence sitting on ungraded criteria, or on no criteria at all, is exactly the failure this section's shape exists to prevent. -->
| Exit criterion (as the plan wrote it) | Verdict | Evidence | Consequence if released as is ||---|---|---|---|| {{exit_criterion}} | {{criterion_verdict}} | {{criterion_evidence}} | {{criterion_consequence}} |
{{conclusion_and_release_call}}
## Residual Risk and What Was Not Tested
<!-- WHAT What remains untested, what remains unfixed, and what that exposes - written as risk somebody is accepting, not as a gap somebody forgot. WHY This is the section a generated report cannot produce, and the one downstream readers thank you for. Silence gets read as coverage: an area nobody tested and nobody mentioned reads, later, exactly like an area that passed, which is why the sharpest assessments in this bundle's research warn their readers not to treat unexamined areas as cleared. Naming the untested area, the unfixed defect and the person carrying the consequence is what turns an omission into a decision. Deep dive: test-summary-report_companion.md section 3 (Residual Risk and What Was Not Tested) and section 7. ASK What was not tested at all, and what was tested more shallowly than its risk deserved? What ships unfixed? For each item, what could go wrong, who is accepting it, and how would you find out in production? What is genuinely undetermined, as opposed to fine? PRIORITY Order by what it would cost if it went wrong, worst first. Every row needs a named person accepting it; a residual risk with no name on it has been accepted by nobody. "Undetermined" is a legitimate entry and a better one than a confident guess. ROW HINT A good row names the untested area or unfixed defect, states the exposure as a consequence to someone, names the person accepting it, and says what would detect it in production. A weak row is a topic with the word "risk" next to it. GOOD | Fraud scoring exercised on synthetic data only | A rule tuned on real distributions could misfire on live claims; false declines reach customers | Tomas Brenner | Decline-rate alert on the R7.2 dashboard, reviewed daily for two weeks | WEAK | Fraud scoring | Some risk | QA | Monitor | TRAP Writing residual risk as an apology. "Unfortunately we ran out of time for the migration path" is a schedule confession. "An unmigrated legacy claim opens read-only and the adjuster cannot progress it; accepted by Ellen Wray for the 40 affected claims" is a risk statement. -->
| Untested area or unfixed defect | Exposure: what could go wrong, and to whom | Accepted by | How it surfaces ||---|---|---|---|| {{residual_item}} | {{residual_exposure}} | {{residual_accepted_by}} | {{residual_detection}} |test-summary-report_template-full.md · ~5,950 tokens
---title: "{{release_or_feature}} Test Summary Report"release_or_feature: "{{release_or_feature}}"build_or_version: "{{build_or_version}}"test_plan_ref: "{{test_plan_ref}}"report_author: "{{report_author}}"test_period: "{{test_period}}"distribution: "{{distribution}}"status: "{{status}}"last_updated: "{{date}}"doc_type: test-summary-reportsize: fullsource_template: test-summary-reportsource_template_version: 0.1.0---
<!--FULL TEST SUMMARY REPORT. The accountability-grade report: everything the lean variant carries, plus whatblocked progress, what testing handed over and what of it survives, and what the effort itself taught. Use itwhen the reader was not in the room - a regulated submission, a certification or audit file, a customer orvendor deliverable, a multi-team release, or a report that is the durable record because the engagement isending.
THIS VARIANT IS A STRICT SUPERSET OF THE LEAN ONE. The six lean sections appear here in the same order, withthe same headings and the same placeholders, and three sections are added. If you started lean, you can growinto this without rewriting anything you already filled in. (The guidance comments in the shared sectionscarry a little extra context for the accountability case; the content you wrote does not change.)
NINE FILLED HEADINGS ARE NOT A REPORT. The sharpest criticism aimed at standardized test documentation isthat a report can satisfy every required heading and still say nothing about the quality of the product: acollection of metrics under a complete outline. Completeness of fields is not completeness of thought. Everysection here should decide something, or tell a reader something they could not get from the test tool. Ifone would change nothing by its absence, write "N/A" and one line of why. Seetest-summary-report_companion.md sections 4 and 6.
WHAT A TEST SUMMARY REPORT IS, AND IS NOTIt is the retrospective document that closes a testing effort: what was tested, what the testing found, whatwas deliberately or accidentally not tested, and whether the result clears the bar the test plan set. It isNOT a test plan (that is prospective, written before), NOT a test status or progress report (that is producedat intervals while testing is still running), NOT your test tool's exported run or your CI dashboard (thosecompute the counts; this document carries the judgment), and NOT a bug tracker (one failing verification isone bug report). See test-summary-report_companion.md section 8.
IF YOU ARE LOOKING FOR A "RELEASE RECOMMENDATION" SECTION, IT IS DELIBERATELY NOT HERE. The ship call belongsat the end of Evaluation Against Exit Criteria, underneath the criteria that justify it. A recommendationwith nothing above it is an opinion; the same sentence under a graded criteria table is a conclusion. Thislibrary considered a standalone section and dropped it for want of evidence; the reasoning is intest-summary-report_companion.md section 3.
THE TWO SECTIONS TO WRITE EVEN IF YOU CUT EVERYTHING ELSE are Evaluation Against Exit Criteria and ResidualRisk and What Was Not Tested. They are the two your tooling cannot fill, and the two a later reader cannotreconstruct.
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into test-summary-report_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.2. Replace each {{placeholder}} with your content. Fill Scope and What Was Tested first, then Evaluation Against Exit Criteria; everything else exists to make that evaluation readable.3. If a section does not apply, write "N/A" and one line of why, rather than deleting it silently.4. Before you circulate it: self-grade against test-summary-report_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{release_or_feature}} Test Summary Report
## Scope and What Was Tested
<!-- WHAT What this report covers: the build or version under test, the environment and configuration, the test plan it answers to, the period it covers, and the boundary of the claim. WHY Every number below is meaningless without the version that produced it, and the version is the field the one filled report in this bundle's research omits, which names no build anywhere, so nobody reading it later can tell what it was about. Bounding the claim is the other half: certification reports say on the cover what configuration the result applies to, because a report that does not bound itself gets read as covering everything. Deep dive: test-summary-report_companion.md section 3 (Anatomy > Scope and What Was Tested). ASK Which build, version or commit was tested, in which environment and configuration? Which test plan, and which of its criteria, does this report answer? What period does it cover, and who tested? What does the result explicitly not apply to? GOOD "Covers Claims Intake R7.2, build 7.2.14, on pre-production with the fnol_v4 flag on, tested 2026-03-02 to 2026-03-13 by Nadia Okonkwo and Fabiola Reyes against the R7.2 test plan. Results apply to that build and configuration only. The broker portal integration runs on its own release train and was not exercised here." WEAK "Testing of the claims release. Ran on staging." (no build, no dates, no plan reference and no boundary, so a reader can tell neither what the report is about nor what it leaves out) TRAP Omitting the build or version because everyone currently knows which one it was. In six months this report is the only record, and nobody will. -->
{{scope_and_what_was_tested}}
## Execution Summary
<!-- WHAT The counts - planned, executed, passed, failed, blocked, not run - per area or suite, with coverage stated against something meaningful, and a line below the table saying when the counts were taken and where the live results live. WHY This is the section your tooling can fill, and the only one. Test management tools and CI plugins compute exactly these numbers and draw them better than a document can, so a report that is mostly this section is a dashboard with a cover page, and out of date the moment it is written. Put the counts here so the judgment below has something to stand on, and put the judgment where the criteria are. Deep dive: test-summary-report_companion.md section 3 (Execution Summary) and section 6. ASK How many cases were planned, and how many actually ran? What failed, what was blocked, what never ran at all? What is coverage a fraction of - risk areas, requirements, the plan's own priorities? When were these counts taken, and where does a reader go for the detail? PRIORITY Order rows by the plan's own risk ranking, highest risk first, so a reader scanning the top of the table is reading about the areas that mattered most. Counts are a snapshot: state the date and time they were taken, because they moved the day after. ROW HINT A good row names an area the plan recognizes, gives every count including the ones that did not run, and says what the coverage figure is a fraction of. A weak row is a suite name and a pass percentage. GOOD | Payout calculation (High risk) | 48 | 46 | 42 | 3 | 1 | 2 | 46 of 48 planned cases; all 9 High-risk scenarios executed | WEAK | Regression | | | | | | | 97 percent | TRAP Leading with a pass rate. "97 percent passed" hides whether the 3 percent was a cosmetic label or the permission check, and a percentage is an input to the exit-criteria evaluation below, never a substitute for it. -->
| Area or suite | Planned | Executed | Passed | Failed | Blocked | Not run | Coverage and notes ||---|---|---|---|---|---|---|---|| {{area}} | {{planned}} | {{executed}} | {{passed}} | {{failed}} | {{blocked}} | {{not_run}} | {{coverage_note}} |
{{execution_summary_notes}}
## Defects
<!-- WHAT What was found, at what severity, what is fixed, and what is still open at the moment of writing - summarized and linked, never transcribed. WHY Open defects are the content here; a list of what you fixed is history. The strongest real reports break the found defects down by something a team can act on, such as root cause or component, and give every open one a severity and a stated reason it is not being fixed. In regulated work that reason is explicitly a risk-based rationale, written down and defensible, which is a good standard to hold yourself to even when nobody is auditing you. Deep dive: test-summary-report_companion.md section 3 (Defects). ASK How many defects were found, at what severities, and what pattern do they show? How many remain open, and which are in scope for the release decision? For each open defect, what would a user experience, why is it not fixed, and who accepted that? What is fixed but not yet verified? PRIORITY Severity first, highest first, and open before closed within a severity. Closed defects are a summary line above the table; every open defect in scope gets a row of its own. In audited or contract work, anything other than a clean pass carries an explanatory note. ROW HINT A good row identifies the defect, gives severity and status, says what a user would experience, and names both the reason it stands and the person who accepted it. A weak row is a ticket number and a title. GOOD | CLM-4471 | Sev-2 | Open | Payout total rounds down by one cent on multi-currency claims | Fix lands in 7.2.15; Ellen Wray accepted the cent-level variance for the two-week window and finance was notified 2026-03-11 | WEAK | CLM-4471 | High | Open | Rounding bug | Will fix | TRAP Pasting the tracker in. Forty rows of reproduction steps make a worse defect list than the tracker itself and bury the three that matter. Summarize, link out, and spend the words on the open ones. -->
{{defect_summary}}
| Defect | Severity | Status | Impact if it ships | Disposition: why it is open, who accepted it ||---|---|---|---|---|| {{defect_id}} | {{severity}} | {{defect_status}} | {{defect_impact}} | {{defect_disposition}} |
## Deviations from Planned Testing
<!-- WHAT What the plan said would happen, what actually happened, and why the difference - in scope, schedule, depth, environment or data. WHY Deviations are what make the numbers above interpretable: a 98 percent pass rate over half the planned depth is a different result from the same rate over all of it. The evidence for keeping this section is one-sided - both structural sources this bundle could read carry it, both real templates ask for it, and the one real filled report gives it a named subsection - which is why it sits in the lean variant too, against this library's own earlier internal spec. Deep dive: test-summary-report_companion.md section 3 (Deviations from Planned Testing) and section 4. ASK What did the plan say, and what actually happened? Which areas were tested less deeply than planned, or not at all? What changed in scope, schedule, environment or data mid-effort, and who agreed it? Who knew at the time, and who is finding out from this document? GOOD "The plan scheduled two full regression passes; one ran, because pre-production was rebuilt in week two. Fraud-scoring cases were executed against synthetic data rather than the anonymized production extract, which never arrived. Both were raised with Ellen Wray on 2026-03-06; neither changed the agreed exit criteria." WEAK "Some testing was descoped due to time." (which testing, how much, whose decision, and what does it do to the counts above) TRAP Writing this section after the release decision is made. A deviation that surfaces for the first time in the report is news, and news arriving this late damages trust in everything else on the page. Raise them as they happen and record them here. -->
{{deviations_from_planned_testing}}
## Impediments and Blocked Progress
<!-- WHAT What got in the way of the testing and what it cost: environment outages, missing test data, unavailable dependencies, access that arrived late, people pulled onto other work. WHY On a small team this section is redundant, because everyone who will read the report lived through it. It earns its place the moment the reader was not in the room - a steering group, a client, an auditor, or the team that inherits this area next quarter - all of whom otherwise read a thin result as thin work. It is also the honest evidence base for the one change most likely to pay off next cycle. Deep dive: test-summary-report_companion.md section 3 (Impediments and Blocked Progress) and section 4. ASK What blocked or slowed the work, and for how long? What did it cost in days, coverage or depth? What is still blocking at the time of writing? Who was asked to unblock it, and what happened? GOOD "Pre-production was rebuilt without notice on 2026-03-05 and was unusable for three working days; Sam Quist restored the seeded personas on 2026-03-09. Cost: one of the two planned regression passes. Still open: the anonymized production extract requested on 2026-02-24 has not been delivered, so fraud-scoring depth remains synthetic." WEAK "There were some environment issues during the cycle." (no dates, no cost, no owner, and no way for a reader to tell whether this explains the coverage gaps above) TRAP Using this section as a complaint. Every entry needs a cost, and anything still live needs an ask addressed to a named person. An impediment with neither is a mood. -->
{{impediments_and_blocked_progress}}
## Evaluation Against Exit Criteria
<!-- WHAT Each exit criterion the test plan set, whether it is met, and on what evidence - then the conclusion the testing reaches about releasing. WHY This is the load-bearing section, and the evaluation against agreed criteria is the act that makes this a report rather than an export. Neither readable structural source gives the release call a heading of its own, and neither does this template: a recommendation with nothing above it is an opinion, while the same sentence underneath a graded criteria table is a conclusion. Write the table, then write the sentence: met, not met, and therefore. Deep dive: test-summary-report_companion.md section 3 (Evaluation Against Exit Criteria), which records why this library considered a standalone Release Recommendation section and dropped it for want of evidence. ASK What exactly were the exit criteria, and who agreed them and when? Is each met, partially met or not met, and on what evidence? Where one is not met, what does releasing anyway cost? What does the testing therefore conclude, and who owns the decision this conclusion feeds? PRIORITY One row per criterion, in the plan's own words and the plan's own order. Never rewrite a criterion to match the result. If the plan set no exit criteria, do not invent them after the fact: say so plainly in the conclusion and state the bar you are applying instead. ROW HINT A good row quotes the criterion, gives a plain verdict (met, partially met, not met), points at the evidence, and where it is not met says what that costs. A weak row is a criterion and a tick. GOOD | All 9 High-risk payout scenarios executed with no Sev-2 or worse left open | Not met | CLM-4471 (Sev-2) open, see Defects | Multi-currency claims round down by one cent until 7.2.15 | WEAK | Quality is acceptable | Met | Testing complete | | TRAP The verdict with nothing above it. A confident release sentence sitting on ungraded criteria, or on no criteria at all, is exactly the failure this section's shape exists to prevent. -->
| Exit criterion (as the plan wrote it) | Verdict | Evidence | Consequence if released as is ||---|---|---|---|| {{exit_criterion}} | {{criterion_verdict}} | {{criterion_evidence}} | {{criterion_consequence}} |
{{conclusion_and_release_call}}
## Residual Risk and What Was Not Tested
<!-- WHAT What remains untested, what remains unfixed, and what that exposes - written as risk somebody is accepting, not as a gap somebody forgot. WHY This is the section a generated report cannot produce, and the one downstream readers thank you for. Silence gets read as coverage: an area nobody tested and nobody mentioned reads, later, exactly like an area that passed, which is why the sharpest assessments in this bundle's research warn their readers not to treat unexamined areas as cleared. Naming the untested area, the unfixed defect and the person carrying the consequence is what turns an omission into a decision. Deep dive: test-summary-report_companion.md section 3 (Residual Risk and What Was Not Tested) and section 7. ASK What was not tested at all, and what was tested more shallowly than its risk deserved? What ships unfixed? For each item, what could go wrong, who is accepting it, and how would you find out in production? What is genuinely undetermined, as opposed to fine? PRIORITY Order by what it would cost if it went wrong, worst first. Every row needs a named person accepting it; a residual risk with no name on it has been accepted by nobody. "Undetermined" is a legitimate entry and a better one than a confident guess. ROW HINT A good row names the untested area or unfixed defect, states the exposure as a consequence to someone, names the person accepting it, and says what would detect it in production. A weak row is a topic with the word "risk" next to it. GOOD | Fraud scoring exercised on synthetic data only | A rule tuned on real distributions could misfire on live claims; false declines reach customers | Tomas Brenner | Decline-rate alert on the R7.2 dashboard, reviewed daily for two weeks | WEAK | Fraud scoring | Some risk | QA | Monitor | TRAP Writing residual risk as an apology. "Unfortunately we ran out of time for the migration path" is a schedule confession. "An unmigrated legacy claim opens read-only and the adjuster cannot progress it; accepted by Ellen Wray for the 40 affected claims" is a risk statement. -->
| Untested area or unfixed defect | Exposure: what could go wrong, and to whom | Accepted by | How it surfaces ||---|---|---|---|| {{residual_item}} | {{residual_exposure}} | {{residual_accepted_by}} | {{residual_detection}} |
## Test Deliverables and Reusable Assets
<!-- WHAT What the testing produced and handed over, and what of it survives this release: executed results, logs and evidence, plus suites, harnesses, fixtures, data sets and environments. WHY Two questions in one section. Deliverables tell a reader where the evidence is, which is what makes the report auditable rather than self-attesting, and they let a sensitive or enormous evidence record sit outside the report while the report still says it exists. Reusable assets tell the next team what they inherit, and this is where teams quietly lose the most: fixtures and harnesses built under deadline are rarely recorded anywhere, and practitioners interviewed about test documentation name the missing historical record as a live problem. Deep dive: test-summary-report_companion.md section 3 (Test Deliverables and Reusable Assets) and section 4. ASK Where does the evidence live - results, logs, screenshots, signed artifacts - and for how long? What was built during this effort that is worth keeping? What is handed to whom, and who maintains it now? What should be thrown away rather than inherited? PRIORITY List anything a later reader would need in order to reconstruct the result, and anything a later team would otherwise rebuild from scratch. Name a location and an owner for each; an asset with no owner is an asset that rots. ROW HINT A good row names the artifact, says where it lives, names who owns it now, and says how long it is kept or how long it stays valid. A weak row is a file name. GOOD | Multi-currency payout fixture set (14 claims, 6 currencies) | claims-qa/fixtures/payouts | Nadia Okonkwo | Reusable; regenerate when the FX provider contract changes | WEAK | Test data | Shared drive | | | TRAP Listing deliverables that do not exist yet, or assets nobody has agreed to maintain. A handover table is a set of commitments; an uncommitted row is a promise the next team will discover was never made. -->
| Deliverable or asset | Where it lives | Owner now | Retention or reuse note ||---|---|---|---|| {{deliverable}} | {{deliverable_location}} | {{deliverable_owner}} | {{deliverable_retention}} |
## Lessons Learned
<!-- WHAT What the testing effort itself taught, and what should change next time - about the testing, not about the product. WHY This section is full-variant only for a relationship reason rather than a doubt: a team that runs retrospectives is already doing this work somewhere it will actually be read, and a report that duplicates the retrospective badly serves nobody. It earns its place when there is no retrospective, or when this report is the durable record because the contract is closing or the team is about to disband. Then it is the only place the learning survives, which is the purpose the testing-process literature names for this document in the first place. Deep dive: test-summary-report_companion.md section 3 (Lessons Learned) and section 8. ASK What would you do differently if this cycle started again? Which of the impediments above was predictable? What worked well enough to keep on purpose? Who owns each change, and where is it recorded so it survives this document? GOOD "The anonymized data extract needs requesting at plan time, not at entry: it was the longest lead-time item in the cycle, for the second release running. Nadia Okonkwo to add it to the R7.3 plan's entry criteria. Keeping on purpose: running the permission matrix before functional depth, which surfaced the two highest-severity defects in week one." WEAK "Communication could be better and we need more time for testing." (true of every cycle ever run, owned by nobody, and it changes nothing next time) TRAP Rewriting the retrospective here. If the team ran one, link it and keep this to the few items a reader outside the team needs, each with an owner. Lessons with no owner are a genre, not a record. -->
{{lessons_learned}}test-summary-report_example.md
---title: "Saved Views for Dashboards Test Summary Report"release_or_feature: "Saved Views for Dashboards (phases 1 and 2)"build_or_version: "Builds 2.3.1 and 2.3.2 on staging; performance measured on the pre-prod replica at 2.3.2"test_plan_ref: "../test-plan/test-plan_example.md (Saved Views for Dashboards Test Plan, approved 2026-07-06)"report_author: "Anjali Rao (QA Lead, Reporting)"test_period: "2026-07-06 to 2026-07-17"distribution: "Priya Nair (PM, Reporting), Marcus Bell (Staff Engineer, Reporting), Sam Okafor (Security), Marta Reyes (Program Manager); attached to the 2.3.2 release checklist"status: "final"last_updated: "2026-07-17"related: - "../test-plan/test-plan_example.md (the plan whose exit criteria this report grades)" - "../test-case/test-case_example.md (TC-047, the case that found DEF-2291)" - "../bug-report/bug-report_example.md (DEF-2291, the defect that suspended phase 2)" - "../acceptance-criteria/acceptance-criteria_example.md (the agreed criteria for the default-view story)" - "../prd/prd_example.md (Saved Views PRD; FR-1 to FR-5 are the scope tested)"doc_type: test-summary-reportsize: fullsource_template: test-summary-reportsource_template_version: 0.1.0---
<!--Worked example for the test-summary-report bundle: a full-variant report closing the qa-docs chain on theAcme Analytics "Saved Views for Dashboards" feature. It reports the effort that the test plan scoped, runningthe cases that plan scheduled, and it accounts for DEF-2291, the defect the bug report records. Every figurein it is illustrative, and the load-bearing ones are marked so inline.
WHY THE FULL VARIANT: three groups tested (Reporting QA, Platform and Design Systems), a formal security gatesat inside the cycle, the migration assets outlive the release, and a criterion is graded not met, so thereasoning has to survive being read by people who were not in the room.
THREE THINGS TO STUDY. First, a pass rate of 96.1 percent sits above a criteria table with one criterion notmet and one partially met: the total and the judgement disagree, and the judgement is the document. Second,criterion 6 is graded Met over a narrower base than planned, because a criterion about the disposition offindings cannot be failed by not looking. Third, criterion 3 is graded Not met and left that way; the exitreview's decision to release anyway is recorded underneath it as a decision, not folded into the grade.-->
# Saved Views for Dashboards Test Summary Report
## Scope and What Was Tested
This report closes the testing of **Saved Views for Dashboards, phases 1 and 2**, against the[Saved Views test plan](../test-plan/test-plan_example.md) approved on 2026-07-06. It covers**2026-07-06 to 2026-07-17** and grades the plan's six exit criteria.
**Builds.** Execution began on **build 2.3.1** on staging with the `saved_views` flag enabled. **Build2.3.2** was cut on 2026-07-14 to carry the fix for DEF-2291 and was the build used from 2026-07-15 onward.Both builds are named throughout, because the two are not interchangeable: 2.3.2 changed where dashboardaggregates are computed, and most of phase 1 was verified before that change existed. The Execution Summarysays which areas' outcomes stand on which build.
**Environments.** Staging, reseeded nightly from the anonymized production snapshot, for functional,end-to-end and accessibility work. The migration sandbox, holding a copy of the legacy key-value store, forthe dry run and rollback rehearsal. The pre-prod replica, frozen for the window, for performance only. Noentitlement assertion was made on the pre-prod replica at any point: its anonymization collapses the regiongrants (illustrative), which would make the restricted persona appear fully entitled and every entitlementcase pass without proving anything.
**What was tested.** Functional requirements FR-1 to FR-5 of the [Saved Views PRD](../prd/prd_example.md)across both rollout phases, exercised through the five `ViewsController` endpoints, the `saved_view` tablemigration and its rollback, the `default_view_id` preference field, and the Views control in the dashboardfrontend. Non-functional coverage: the shared-view entitlement boundary, view-list load performance,degradation when a saved config references a deleted filter field, and WCAG 2.2 AA keyboard andscreen-reader operation of the Views control.
**Who tested.** Anjali Rao (API, end-to-end, the permission matrix), Lee Zhang (migration and rollback),Marcus Bell (performance), Sofia Marino (accessibility), Sam Okafor (security review on 2026-07-10).
**What this report does not cover, stated so that silence is not read as clearance.**
- **The dashboard permissions service.** Platform owns and tests it. What was verified here is this feature's *use* of it, specifically that a shared view re-checks the recipient's access when the recipient reads it.- **FR-6, the stale-view change indicator.** Not built in this release, so there was nothing to test.- **The legacy report flow.** The standing regression suite ran green against both builds but was not re-planned, re-scoped or re-read for this release, and no claim is made about it here.- **Adoption.** Measured after launch on the KPI dashboard. Testing cannot verify it and did not try.- **Any build other than 2.3.1 and 2.3.2, and any environment other than the three named above.** The performance figures in particular are claims about the pre-prod replica at its frozen data profile and about nothing else.
**A note on severity labels.** The test plan writes its severity bar as "Sev-1 or Sev-2". Acme's trackeruses the four-level scale S1 Critical / S2 Major / S3 Minor / S4 Trivial. They are the same scale under twospellings; the criterion below is quoted in the plan's own words and graded against tracker severities.
## Execution Summary
Rows are in the plan's own risk order, highest first, so the top of this table is the part of the releasethat mattered most.
| Area or suite | Planned | Executed | Passed | Failed | Blocked | Not run | Coverage and notes ||---|---|---|---|---|---|---|---|| Shared-view entitlement (High, register R-05) | 12 | 12 | 12 | 0 | 0 | 0 | The full permission matrix: 3 personas against 4 filter scopes, including TC-046, TC-047 and TC-048, which between them exhaust the entitlement partitions. 7 of the 12 had run on 2.3.1 when TC-047 failed and the matrix was suspended; all 12 were executed from the start on 2.3.2. Denominator is the matrix the plan defined, not the sharing surface || Config migration and rollback (High, register R-02) | 18 | 18 | 18 | 0 | 0 | 0 | 12 seeded known-bad legacy configs reconciled, plus 6 rollback and schema-version cases. Run on 2.3.1 in the migration sandbox, 2026-07-08 to 2026-07-09, and not re-executed on 2.3.2. Coverage is of the 12 config shapes that were seeded, not of the shapes production holds || Stale-field degradation (Medium-high) | 9 | 9 | 8 | 1 | 0 | 0 | The PRD reliability requirement and the design's `stale_fields` path. Eight ran on 2.3.1, where TC-052 raised DEF-2277 on 2026-07-08; that case was re-verified on 2.3.2. The open failure is DEF-2298 (S3) || View-list load performance (Medium, register R-06) | 6 | 6 | 5 | 1 | 0 | 0 | Two thresholds at three view counts (50, 100 and 150 views on one dashboard) on the pre-prod replica, 2.3.2 only. The single failure is view-list load at 150 views: DEF-2304 || Default-view resolution (Medium) | 14 | 13 | 13 | 0 | 1 | 0 | 9 cases derive from the [agreed acceptance criteria](../acceptance-criteria/acceptance-criteria_example.md); 5 came from test design and no criterion names them. Run on 2.3.1 only. The blocked case needs a mid-session access revocation that Platform could not schedule || Accessibility of the Views control (Medium) | 14 | 11 | 10 | 1 | 0 | 3 | NVDA on Windows completed in full on 2.3.1; the accessible-name case was re-verified on 2.3.2. The 3 not run are the VoiceOver on macOS pass, which never started. The open failure is DEF-2286 (S3) || Rename and delete (Low) | 8 | 8 | 8 | 0 | 0 | 0 | Owner and non-owner against an existing and a deleted row, which is the equivalence partition set and deliberately no more. Run on 2.3.1 only || **Total** | **81** | **77** | **74** | **3** | **1** | **3** | |
Counts were taken on **2026-07-16 at 17:00 UTC**, at the close of execution. They will move: DEF-2298 andDEF-2286 are open and their cases will flip when the fixes land. The live results are in the test tool'ssprint 14 run record; this table is a reading of it at one moment, not a replacement for it.
**On the 96.1 percent, because someone will quote it.** 74 of 77 executed cases passed. That figure is aproperty of the cases this team chose to write and chose to run, and it says nothing about how much of theproduct they reach. Four of the seven areas above were verified on 2.3.1 and never re-executed against2.3.2, three planned accessibility cases were never attempted, and the coverage claim in each row names itsown denominator for that reason. The plan refused to make a pass rate an exit criterion, and this cycle isthe argument for that refusal: 96.1 percent would have cleared any plausible pass-rate bar on a release thatalso suspended testing for an S1 data-exposure defect and finished with a criterion graded not met.
## Defects
**Eleven defects were raised in the window** (illustrative): one S1, three S2, five S3, two S4. Eight areclosed. Build 2.3.2 carried the DEF-2291 fix and, because it was the only build cut after entry, also pickedup the fixes for DEF-2277, DEF-2284 and the five lower-severity defects already merged and waiting.
**Three are open**, and they are the content of this section. By area, the open set clusters where therelease is thinnest rather than spreading evenly: one in performance, two in the user-facing edges of thefeature (the stale-field notice and keyboard focus). Nothing is open against entitlement, migration ordefault-view resolution.
DEF-2291 is closed and still gets a row below, because it is the reason phase 2 was suspended, the reasonthe permission matrix was executed twice, and the evidence behind criterion 2's grade. Its full record,including the cause and the regression guard, is in the [bug report](../bug-report/bug-report_example.md);nothing from it is retyped here.
| Defect | Severity | Status | Impact if it ships | Disposition: why it is open, who accepted it ||---|---|---|---|---|| [DEF-2291](../bug-report/bug-report_example.md) | S1 Critical | Closed | A recipient of a shared view could read the magnitude of data they are not entitled to, from an aggregate computed before the entitlement filter | Fixed in 2.3.2 (2026-07-14), verified by Anjali Rao 2026-07-15, full permission matrix re-executed from the start. Listed here because criterion 2 rests on that re-execution || DEF-2304 | S2 Major | Open | A dashboard carrying many saved views opens its view list slowly. p95 612ms at 150 views against the plan's 500ms budget (illustrative), so the feature degrades exactly where the heaviest users are | No fix attempted in the window; found 2026-07-15, after the build was cut. Accepted for release at the 2026-07-17 exit review by Priya Nair, Marcus Bell and Sam Okafor, under the conditions recorded in Evaluation Against Exit Criteria. Priya Nair is accountable for the acceptance; Marcus Bell owns the fix, scheduled for 2.4 || DEF-2286 | S3 Minor | Open | Closing the Views menu with Escape returns focus to the dashboard body instead of the Views trigger, so a keyboard user loses their place and has to tab back. A WCAG 2.2 AA finding | Accepted in writing by Priya Nair on 2026-07-16 under exit criterion 6, deferred to 2.4. Sofia Marino raised it on 2026-07-09 || DEF-2298 | S3 Minor | Open | The stale-field notice says how many filters are missing but does not name them, so an analyst cannot tell what to re-save without opening the view's configuration | Found 2026-07-15 while re-verifying DEF-2277; it could not be seen earlier, because the read path threw an error instead of rendering the notice at all. Accepted by Priya Nair on 2026-07-16, deferred to 2.4. **This one misses an agreed acceptance criterion**, which asks for a message naming the missing filter; see Evaluation Against Exit Criteria |
## Deviations from Planned Testing
**Phase 2 was suspended for two days, and the plan's resumption rule was applied in full.** TC-047 failed atstep 4 on 2026-07-13; Anjali Rao suspended all sharing testing at 15:10 UTC that day, per the plan'ssuspension rule, and notified Priya Nair, Marcus Bell and Sam Okafor the same afternoon. Phase 1 continued,because it does not exercise the sharing path. Testing resumed on 2026-07-15 once 2.3.2 was deployed and SamOkafor confirmed the security review unblocked. Cost: two of the four planned sharing days. Recovered bydropping a second exploratory sharing session, not by shortening the matrix.
**The permission matrix was executed twice, which the plan did not budget for.** Seven of the twelvecombinations had run on 2.3.1 when TC-047 failed. The plan's resumption rule requires the entire matrix tobe re-run from the start rather than the failing case alone, because an entitlement defect invalidates theassumption behind every result that passed before it. All twelve were re-executed on 2.3.2 on 2026-07-15.This is the single largest difference between planned and actual effort in the cycle and it was the rightcall; recording it here is what lets a reader see that the twelve passes are twelve passes on one build, notseven on one and five on another.
**Three accessibility cases were never run.** Sofia Marino completed the NVDA on Windows keyboard andannouncement cases between 2026-07-09 and 2026-07-13 and returned to the Billing program on 2026-07-14, asagreed before the cycle started. The VoiceOver on macOS pass was handed to Anjali Rao for 2026-07-14 to2026-07-16, to be run against Sofia Marino's protocol. It never started: the plan's resumption rule put thefull twelve-combination permission matrix back on the 15th and the resumed sharing cases on the 15th and16th, and one person could not do both. Anjali Rao flagged the clash to Marta Reyes on 2026-07-14 and didnot escalate for a second reviewer, judging the remaining window too short to be worth another team'scontext-switch. That judgement is recorded rather than defended: the consequence is a coverage gap carriedas a residual risk below, and a different call was available.
**One default-view case was blocked.** Verifying that a default which is a shared view the user can nolonger read falls back cleanly requires revoking a recipient's dashboard access mid-session, which is aPlatform operation. Requested from Dana Osei on 2026-07-10; not scheduled inside the window.
**This report is not the one-page summary the plan promised.** The plan's deliverables list commits to "aone-page test summary against the exit criteria". A suspension, an S1, a criterion graded not met and aresidual risk being carried by three named people do not compress to one page honestly. The one-page formexists as the 2.3.2 release checklist entry and links here; this is the record it links to.
## Impediments and Blocked Progress
**The hotfix build was the constraint, not the fix.** The DEF-2291 fix was merged on 2026-07-13, the eveningit was triaged. Build 2.3.2 was cut on 2026-07-14 and deployed to staging that afternoon, so testing resumedon the 15th. Cost: roughly one working day of the two-day suspension was the release train rather than theengineering. Nobody is at fault in that sentence, and it is the single clearest candidate for shortening thenext suspension.
**The accessibility pass lost its stand-in and has no specialist owner.** The VoiceOver on macOS cases movedto Anjali Rao when Sofia Marino returned to Billing on schedule, and the resumption rule then claimed thesame three days. Sofia Marino is committed to Billing through 2026-07-31, so the pass now has neither a datenor anyone qualified to run it. **Still open at the time of writing**, and the ask is to Marta Reyes: eithera Design Systems reviewer for two days before the 2.4 window opens, or an explicit decision to ship theViews control without a VoiceOver pass and record that on the accessibility register.
**The mid-session access revocation was never scheduled.** Requested from Dana Osei on 2026-07-10 and againon 2026-07-14. Cost: one blocked case and one unverified fallback path. **Still open**; the ask is ahalf-hour Platform slot in the 2.4 entry window, which is small enough that the honest reading is that itwas never anyone's priority, including this team's.
**What did not go wrong, because the plan said it probably would.** The plan's top risk to the effort wasstaging contention with the Billing migration in week two, landing exactly on the phase 2 window. The2026-06-29 booking for 13 to 16 July was honoured and the contention never materialised, so the plan'sfallback of running the permission matrix on the pre-prod replica was never invoked. That matters more thana clean risk usually does: TC-047 is explicitly not valid on the replica, whose anonymization collapses theregion grants, so the fallback would have produced twelve passes that proved nothing. The contingency in theplan was wrong, and only luck stopped it being used. It should not survive into the 2.4 plan.
## Evaluation Against Exit Criteria
The six criteria are quoted in the plan's own words and in the plan's own order. None has been reworded tomatch a result.
| Exit criterion (as the plan wrote it) | Verdict | Evidence | Consequence if released as is ||---|---|---|---|| 1. "Every High and Medium-high tier area in the Risk-Ranked Approach has its planned cases executed. Not "most cases": these three areas, complete." | Met | Shared-view entitlement 12 of 12, config migration and rollback 18 of 18, stale-field degradation 9 of 9. No case in these three areas is blocked or unrun | None || 2. "The permission matrix is 100 percent executed with zero failures. This one is absolute; a single failure here is a suspension event, not a defect to triage." | Met, on re-execution | 7 of 12 combinations had run on 2.3.1 when TC-047 failed at step 4; the matrix was suspended, DEF-2291 was fixed in 2.3.2, and all 12 were executed with zero failures on 2026-07-15 under the plan's resumption rule | None for release. The criterion is met against the build being shipped. The failure it records against 2.3.1 is DEF-2291, closed || 3. "Zero open Sev-1 or Sev-2 defects against in-scope requirements." | **Not met** | DEF-2304 (S2 Major) is open against the plan's view-list load budget, carried in the register as risk R-06 | View-list load runs 612ms at p95 against a 500ms budget at 150 views (illustrative), roughly 22 percent over. It degrades rather than breaks, and it degrades worst for the analysts with the most saved views, who are the feature's heaviest users || 4. "Migration reconciliation shows a zero-row mismatch on the dry run, and the rollback rehearsal has been completed once end to end." | Met | Zero mismatch across the seeded legacy extract on 2026-07-08; rollback rehearsed once to the read-only legacy store on 2026-07-09, by Lee Zhang | None. Worth noting what the criterion does not assert: reconciliation counts rows and does not prove a converted config is semantically right. DEF-2277 was exactly that case, a row that migrated correctly and could not be read || 5. "p95 view switch under 1s and p95 view-list load under 500ms on the pre-prod replica at 3x expected view count (illustrative thresholds, taken from the PRD and the register)." | Partially met | View switch p95 0.74s, met. View-list load p95 180ms at 50 views, 340ms at 100, and 612ms at 150, so not met at the 3x count the criterion names (all illustrative). DEF-2304 | As criterion 3. The two criteria fail on one defect, not two problems || 6. "Accessibility findings at AA level are either fixed or accepted in writing by Priya Nair." | Met | Two AA findings. DEF-2284 (missing accessible name on the Views trigger) fixed in 2.3.2 and re-verified 2026-07-15. DEF-2286 accepted in writing by Priya Nair on 2026-07-16 | The criterion is met over 11 of 14 planned cases, on NVDA only. A criterion about the disposition of findings cannot be failed by not looking, so meeting it proves less here than its wording suggests. The unrun VoiceOver pass is carried as a residual risk below |
**What the testing concludes.** Four criteria are met, one of them only because the permission matrix wasre-run in full on a second build; one is partially met; one is not met. The three highest-ranked productrisks the plan carried, entitlement, migration and stale-field degradation, are all cleared on evidence, andthe S1 that interrupted the cycle is fixed, verified and guarded by a regression case. **Testing thereforesupports releasing phases 1 and 2 on build 2.3.2 with DEF-2304 open**, on three conditions: that the phasedrollout in the PRD is followed so the 150-view profile is reached gradually rather than on day one, that ap95 alert on view-list load is live before the first cohort, and that Marcus Bell owns DEF-2304 into 2.4.Testing does not support releasing on a build where criteria 1, 2 or 4 are anything other than met.
**The unmet criterion was not rewritten, and that distinction is the point.** Criterion 3 stays graded Notmet in this document. What happened at the exit review on 2026-07-17 is that Priya Nair, Marcus Bell and SamOkafor, the three signatories the plan's change control names for criteria 2 and 3, decided to release withDEF-2304 open and to carry the exposure under the conditions above. That is a decision about an unmetcriterion. It is not a relaxation of the criterion, and if the plan's bar changes for 2.4 it changes in the2.4 plan, in advance, where the next team can see it before they start.
**One thing no criterion catches.** DEF-2298 is an S3, so it sits comfortably inside criterion 3's bar, andit misses an agreed acceptance criterion: the default-view story asks that a view referencing a deletedfilter "shows a clear message naming the missing filter", and the shipped notice gives a count instead. Theplan's exit criteria are severity bars, and a severity bar cannot see the difference between a cosmetic S3and an S3 that breaks something the business agreed to. Priya Nair accepted it knowing that; it is recordedhere so the acceptance is on the record rather than implied by a number.
## Residual Risk and What Was Not Tested
Ordered by what it would cost if it went wrong. Every row names a person carrying it; an entry with no nameon it would be a risk nobody has accepted.
| Untested area or unfixed defect | Exposure: what could go wrong, and to whom | Accepted by | How it surfaces ||---|---|---|---|| Phase 1 outcomes stand on build 2.3.1 and were not re-executed on 2.3.2 | 2.3.2 moved aggregate computation behind the entitlement filter. Migration, default-view resolution and rename/delete were verified before that change existed, and their evidence is therefore one build old. A regression in any of them would reach analysts as silently as DEF-2291 did | Priya Nair | CI unit and component gate plus the smoke suite, both green on 2.3.2. Neither exercises the migration path or the default-view precedence rules, so in practice this surfaces as a user report || View-list load p95 exceeds the budget at 150 views (DEF-2304, open) | Analysts who accumulate views, who are the feature's most engaged users, wait longest. The program is selling speed, so this erodes the thing being sold rather than breaking it | Priya Nair, accountable, at the 2026-07-17 exit review with Marcus Bell and Sam Okafor co-signing | p95 view-list alert on the reporting latency panel, threshold 550ms (illustrative), live before the first rollout cohort || VoiceOver on macOS was never exercised (3 cases not run) | Undetermined. NVDA passed on 10 of 11 cases, and the two engines diverge most on exactly the custom-menu pattern the Views control uses, so NVDA passing is weak evidence for VoiceOver. A macOS screen-reader user could find the control unusable and the release would not know | Priya Nair | Nothing automated detects this. It surfaces through a support ticket or an accessibility complaint, which is the slowest and most expensive path available || A recipient losing dashboard access while a shared default view is open (1 case blocked) | The fallback is unverified. Worst case the dashboard fails to load rather than falling back to the generic state, which locks a user out of a dashboard they still have rights to | Marcus Bell | Error-rate panel on dashboard open. The fallback path is logged, so the signal exists; nobody is watching it today || Migration coverage is bounded by the 12 seeded known-bad configs | Undetermined, and deliberately so. Nobody enumerated the config shapes production actually holds, so the reconciliation proves the converter handles twelve shapes rather than all of them. An unseen shape fails at cutover, when the legacy store is already read-only | Lee Zhang | Reconciliation counter runs at cutover and halts the migration on any mismatch, which converts a silent loss into a visible stop. That is the mitigation; it is not detection in advance || The rendered-UI entitlement assertion runs once per release, by hand | TC-047's automated form asserts on the API response only. Nothing in the pipeline checks that an unentitled value never appears in a rendered tooltip, chart label or tile. A future change that leaks through the render layer alone passes every automated gate | Sam Okafor | Only the manual step 5, run once per release. Until it is automated, the guard on the class of defect DEF-2291 belongs to is thinner than the regression set implies |
## Test Deliverables and Reusable Assets
| Deliverable or asset | Where it lives | Owner now | Retention or reuse note ||---|---|---|---|| Executed case results, all 77 | Test tool, sprint 14 run record for Saved Views | Anjali Rao | Retained for the life of the 2.3 branch per the release checklist. This report is a reading of it, not a copy || Open-defect list with severities | Tracker, filter `feature = saved-views AND status = open` | Anjali Rao | Live. The three rows in Defects above are a snapshot at 2026-07-16 17:00 UTC || Migration rollback rehearsal record | Migration sandbox run log, 2026-07-09 | Lee Zhang | Required separately by the release checklist; keep until the legacy key-value store is decommissioned || DEF-2291 evidence bundle (screenshot, API response, staging log with the request ID) | Attached to [DEF-2291](../bug-report/bug-report_example.md) | Marcus Bell | Keep as long as the defect record. It is the only artifact showing the pre-fix behaviour || Permission persona set: owner, permitted viewer, restricted viewer | Platform persona fixtures, staging seed | Dana Osei | Reusable and fragile. The personas are manufactured rather than sampled, because the anonymized snapshot carries no restricted filter field; regenerate whenever the entitlement model changes || Entitlement spec directory, including `tests/entitlement/shared_view_restricted_viewer_spec.rb` | Release branch CI | Marcus Bell | In the release regression set, runs on every pipeline execution. TC-053 was added to it after DEF-2291 to cover the row-count badge, which shares the pre-filter computation path || 150-view fixture generator for the pre-prod replica | `reporting-qa/fixtures/view-volume` (illustrative path) | Marcus Bell | Reusable for any view-count performance work. It assumes the frozen replica profile; re-check it after the next replica refresh or the numbers are not comparable || Legacy-store extract with 12 known-bad configs | Migration sandbox | Lee Zhang | Keep until the migration is retired. Its weakness is written into the residual-risk table above: twelve shapes, chosen by hand || Hand-built fallback fixtures from the plan's contingency | Not built | Nobody | **Do not inherit.** The contingency was never triggered, so no fixtures exist. Listed so the next team does not go looking for them |
## Lessons Learned
**Review High-tier cases before they run, and make it an entry criterion.** TC-047 version 1.0 asserted onreturned rows only and would have passed against a build carrying DEF-2291. Step 4, the aggregate assertion,was added at version 1.1 on 2026-07-08 after Sam Okafor reviewed the case, five days before it fired. Thecase review, not the case, is what found the S1. **Anjali Rao** to add security review of every High-tiercase to the 2.4 plan's entry criteria.
**A row-level assertion cannot see an aggregate leak, and that generalises past this feature.** Everyentitlement case written before TC-047 asserted on rows, and the rows were always correct. Anywhere theproduct computes a number over a filtered set, the number needs its own assertion. **Marcus Bell** hascarried this into the entitlement spec directory via TC-053; the wider sweep of other aggregate surfaces hasno owner yet and should get one in 2.4 planning.
**Schedule accessibility against the first stable build, not the last.** Accessibility was ordered sixth ofseven, which put its last three cases after the specialist's agreed departure date and inside the days theresumption rule reclaimed. Nothing about the Views control's keyboard or screen-reader behaviour required alate build, so ordering it late bought nothing and cost three cases. **Anjali Rao** to reorder it in the 2.4plan.
**Retire the plan's staging-contention contingency instead of reusing it.** Running the permission matrix onthe pre-prod replica was written into the plan as the fallback and would have produced twelve meaninglesspasses, because the replica's anonymization collapses exactly the grants the matrix tests. The contingencywas never needed, which is the only reason it caused no harm. **Anjali Rao and Dana Osei** to replace it in2.4 with a second staging slice, or to record that there is no fallback and the booked window is themitigation.
**Keeping on purpose: inheriting product risks from the register rather than re-deriving them.** TheExecution Summary's row order is the plan's risk order, which is the register's risk order. A readerscanning the top of that table is reading the areas the program already agreed mattered most, and nobody hadto argue about ordering at any point in the cycle.
No retrospective was held for this cycle; the team runs one per release train rather than per feature, andthe next falls after 2.4 opens. These five items are therefore recorded here because this document is wherethey survive, and each has a name against it for the same reason.Provenance
Section titled “Provenance”The reasoning, the history and every source, in the repository:
- Companion - the long-form argument: why these sections, where the sources disagree, and what the bundle refuses to claim
- History - what changed in this bundle, and when
- Research log - every source consulted, with what each one actually supports
- Catalog metadata - the machine-readable record this page is generated from
Catalog record: 9 sections across 1 format(s).