Test Plan
beta · Family qa-docs · Phase develop · Sizes lean, full · ~2,650 tokens
The prospective document that scopes and prioritizes a testing effort: what is being tested and what is explicitly not, which areas carry the most product risk and therefore earn the deepest coverage, the measurable conditions that start and end the work, and who runs it where. Written to be read, not to be complete: the plan is the thinking, not the page count.
The short card. Why the document is shaped this way, and the argument behind every rule here, is in
test-plan_companion.md. A fully worked instance is
test-plan_example.md.
When to use
Section titled “When to use”- A release or feature is about to be verified and more than one person needs to know what is covered.
- Testing spans more than one team, vendor or specialism, so coverage assumptions need to be written down rather than assumed.
- There is a gate inside the cycle (a security review, a compliance sign-off, a staged rollout) and someone has to decide in advance what passing it means.
- The context is regulated or audited and the plan is part of the evidence.
- Coverage is going to be argued about afterwards, and you would rather have the argument now, cheaply.
When NOT to use
Section titled “When NOT to use”- The plan would decide nothing. One person, one afternoon, one obvious feature. Write the two entry and exit conditions in the ticket and move on.
- You need a test case. One executable verification with steps, data and an expected result is a different artifact. The plan schedules cases; it does not contain them.
- You need a test report. The plan is written before and is prospective. Results, defect counts and the verdict belong in a report written after.
- You need a definition of done. A DoD is a standing team invariant applied to every increment. A test plan is scoped to one release.
- Your tool already has a “test plan”. It probably means an execution container: a grouping of suites and runs with a name and a date range. That is a complement to this document, not a substitute (see below).
- Continuous delivery with no discrete release. Do not write one plan per deploy. Write one standing plan per product area and revise it in place.
Plan, strategy, case, or report? (the question people actually have)
Section titled “Plan, strategy, case, or report? (the question people actually have)”| You want to say | Write |
|---|---|
| How this organization tests, in general, across projects | A test strategy (standing, rarely changes) |
| What we are testing this time, how deeply, and when we are done | A test plan (this template) |
| The exact steps and expected result for one behavior | A test case |
| What we found, and whether we are shipping | A test report |
Two honest notes. The plan/strategy line is genuinely unsettled: the certification bodies put strategy at the organization or programme level, plenty of teams carry it as a section inside the plan, and practitioners openly disagree about which is written first. If your organization has a standing strategy, say in one line that this plan inherits from it. If it does not, your Risk-Ranked Approach section is your strategy for this release, and that is a legitimate answer rather than a gap.
And the tool trap is real. When Azure Test Plans, TestRail, Xray, Zephyr or qTest says “test plan”, it means a container for suites, runs and configurations. Those tools do not require an approach, a risk ranking or an exit criterion. Teams that believe the tool replaced the document end up with an execution container and no recorded thinking. Keep both: the narrative here, the execution there.
Pick a variant
Section titled “Pick a variant”Lean (five sections) is the default. Scope and Non-Scope, Risk-Ranked Approach, Entry and Exit Criteria, Environment/Data/Ownership, Schedule and Deliverables. One team, one feature or release, a page or two, revised in place.
Full (nine sections) adds Test Levels and Types, Suspension and Resumption, Risks to the Test Effort, and Approvals and Change Control. Use it when at least one of these is true:
- the work is regulated or audited, and completeness is a compliance requirement;
- more than one team or vendor is testing the same release;
- there is a formal gate mid-cycle that can stop the work;
- your process requires an approved, version-controlled artifact.
The test for adding a section is whether its absence would change anything. If it would not, leave it out. Every lean heading appears in full unchanged, so growing is additive and you never rewrite what was agreed.
Quality rubric (self-grade)
Section titled “Quality rubric (self-grade)”Score each 0, 1 or 2. Under 12 out of 18 and the plan will not survive contact with the release.
| # | Criterion | 0 | 1 | 2 |
|---|---|---|---|---|
| 1 | Non-scope is explicit | Nothing excluded | Exclusions listed | Exclusions listed with a reason each, and agreed by a named person |
| 2 | Risk ranking is real | No ranking, or everything is High | Tiers assigned | Tiers assigned with a stated reason, and coverage depth differs by tier |
| 3 | Order follows risk | No order given | Order given | Highest-risk areas are scheduled first, visibly |
| 4 | Criteria are checkable | Adjectives (“ready”, “complete”) | Some measurable | Every criterion is a number, a state, or a named artifact |
| 5 | Exit criteria resist gaming | Pass rate alone | Pass rate plus something | Coverage-of-risk criterion plus a severity rule plus a named agreement with a date |
| 6 | Owners are people | No owners | Teams named | One named individual per area, and per risk |
| 7 | Risks are separated | One merged list | Both present, mixed | Product risk shapes coverage; risk to the effort is separate with contingencies |
| 8 | Test data is named | Not mentioned | Mentioned generally | Specific data for the negative and boundary cases, with who creates it and when |
| 9 | It would be read | Over ten pages of prose | Long but skimmable | Short enough that the team actually reads it, and every section decides something |
Criterion 9 is the one people argue with and the one that matters most. A plan nobody reads is not a safety net.
Named anti-patterns (the usual wrecks)
Section titled “Named anti-patterns (the usual wrecks)”- The plan nobody reads. Written for the kickoff, filed, never reopened. The tell: no decision in the last month traces back to it. Fix by cutting until it is worth reading.
- Completeness theater. Every heading filled, nothing decided, because the template demanded a heading. A document can satisfy every required field and still say nothing. Fix by deleting sections that decide nothing.
- Adjective criteria. “Environment is ready”, “quality is acceptable”. Unfalsifiable, therefore not criteria. Fix with a number, a state or a named artifact.
- Pass-rate exit. “95 percent pass” as the only gate, satisfiable by writing more trivial tests, and able to hide one catastrophic failure inside a comfortable percentage. Fix by pairing it with risk coverage and a severity rule.
- The merged risk list. Product risk and project risk in one table, so neither drives anything. Fix by splitting them: one shapes coverage, the other shapes the schedule.
- The scope with no non-scope. Nothing excluded, so nothing can be found missing, and the argument gets held after the release instead of before it. Fix by writing the exclusions and getting them agreed.
- The plan that became a tracker. Results and status columns creep in and the prospective document turns retrospective. Fix by keeping execution state in the tool and the verdict in the report.
Pairing with a skill
Section titled “Pairing with a skill”pairs_with: [deliver-edge-cases]. There is no testing or QA skill in the pm-skills library: none of
its 68 skills produces a test plan, a test case or a bug report (recorded as finding EC-4 in the repository’s
STATE.md). The one honest pairing is deliver-edge-cases: run it first and feed its edge-case catalog into
the Risk-Ranked Approach section. See companion section 8 for why the two fit together. Everything else in
this template is filled by hand.
The artifacts
Section titled “The artifacts”test-plan_template-lean.md · ~2,650 tokens
---title: "{{release_or_feature}} Test Plan"release_or_feature: "{{release_or_feature}}"plan_owner: "{{plan_owner}}"status: "{{status}}"last_updated: "{{date}}"doc_type: test-plansize: leansource_template: test-plansource_template_version: 0.1.0---
<!--LEAN TEST PLAN. The smallest plan that is still a real plan: what is in and out, how deeply each area will betested and why, the conditions that start and stop the work, who runs it where, and when. Use it for afeature or a release owned by one team. To grow it into a formal, approvable plan (seetest-plan_template-full.md), ADD sections; never rename or reorder the ones below, because the full variantis a strict superset of this one.
THE PLAN IS NOT THE DOCUMENT. The value is in the thinking this forces, not in the page count. A plan nobodyreads is not a safety net. Keep it short enough that the team actually reads it, and specific enough thatsomeone else could check it. Delete any section that would change nothing if it were missing, and say why.See test-plan_companion.md sections 1 and 6.
WHAT A TEST PLAN IS, AND IS NOTIt is the prospective document that scopes and prioritizes a testing effort. It is NOT a test case (that isone executable verification), NOT a test report (that is retrospective, written after), NOT a definition ofdone (that is a standing team invariant), and NOT the thing your test management tool calls a "test plan"(that is an execution container whose only required field is usually a name). Seetest-plan_companion.md section 8.
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into test-plan_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.2. Replace each {{placeholder}} with your content. Fill Scope and Risk-Ranked Approach first; everything else follows from them.3. If a section does not apply, write "N/A" and one line of why, rather than deleting it silently.4. Before you share it: self-grade against test-plan_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{release_or_feature}} Test Plan
## Scope and Non-Scope
<!-- WHAT What is being tested (the concrete test items: builds, services, flag states, requirement IDs) and what is explicitly NOT being tested this cycle, with a reason for each exclusion. WHY The exclusion half is the highest-value content in the document and the most commonly deleted. If nothing is out of scope, nothing can be found missing, and the scope argument happens after the release instead of before it. Writing it down converts an assumption into an agreement. Deep dive: test-plan_companion.md section 3 (Anatomy > Scope and Non-Scope). ASK What exactly is under test, by build or requirement ID? What is deliberately excluded, and why (deferred, covered elsewhere, accepted risk)? Who agreed to the exclusions? GOOD "In scope: FR-1 to FR-5 of the Saved Views PRD, on build 2.3.x behind the saved_views flag. Out of scope: FR-6 (stale-view indicator), Could priority, deferred to the next release; and the permissions service itself, owned and tested by Platform." WEAK "Testing the Saved Views feature." (no test items, no exclusions, nothing anyone can disagree with or be held to) TRAP Leaving out the exclusions because they feel negative. An unstated exclusion is not a smaller scope, it is an unmanaged expectation. -->
{{scope_and_non_scope}}
## Risk-Ranked Approach
<!-- WHAT The areas under test, ranked by product risk, with the reason for each ranking and the depth of testing it earns. This is the section that makes the plan a plan. WHY Risk-based testing exists to steer effort rather than spread it evenly: high-risk areas earn heavier techniques and are tested FIRST, so that if time runs out what is missing is the least important coverage rather than a random sample. A section that only names test types ("functional, regression") decides nothing. Deep dive: test-plan_companion.md section 3 (Risk-Ranked Approach) and section 5. ASK What could go wrong here, and how bad would it be? Which areas therefore carry the most risk? What technique and depth does each risk tier earn? What order will they be tested in? Where do these risks come from (the PRD, the risk register, a risk-storming session)? PRIORITY Order the table by risk, highest first, and test in that order. Reuse product risks already recorded upstream rather than inventing a parallel list. Risks that threaten the TESTING (not the product) do not belong here; in the full variant they have their own section. ROW HINT A good row names an area, states the risk as a consequence rather than a theme, gives a tier with a reason, and names the technique and depth. A weak row is a feature name with "high" beside it. GOOD | Shared-view permissions | A shared view could expose rows the recipient may not see | High: low likelihood, severe and irreversible impact (data exposure) | Permission matrix across 3 personas x 4 filter scopes; negative cases first; security review gate | 1 | WEAK | Sharing | High | Test it thoroughly | 1 | TRAP Ranking everything "high". If every area is top priority, the ranking has told the team nothing and the order of work is back to guesswork. -->
| Area under test | Risk (what could go wrong, and the consequence) | Tier and why | Technique and depth | Order ||---|---|---|---|---|| {{area}} | {{risk_and_consequence}} | {{tier_and_reason}} | {{technique_and_depth}} | {{order}} |
## Entry and Exit Criteria
<!-- WHAT The conditions that must be true before testing starts, and the conditions that define testing as done, with who agreed the exit criteria and when. WHY Criteria are thresholds, not adjectives. "Environment is ready" cannot be checked; a named build in a named state can. The functional test for an entry criterion is whether failing it would actually stop you: if not, it is not a criterion. Exit criteria are supposed to be agreed with stakeholders, and almost every template omits the evidence of that agreement. Deep dive: test-plan_companion.md section 3 (Entry and Exit Criteria) and section 7. ASK What must exist before testing can start without wasting effort? What would make you stop and say testing is finished? How is each condition measured, and by whom? Who agreed the exit criteria, and on what date? GOOD "Entry: build 2.3.1 deployed to staging with the flag on; three permission personas seeded; smoke suite green. Exit: every High-tier area has its planned cases executed; zero open Sev-1 or Sev-2 defects in scope; the permission matrix 100 percent executed with no failures. Agreed with Priya Nair (PM) and Marcus Bell (Eng lead) on 2026-07-06." WEAK "Entry: environment ready. Exit: 95 percent of tests pass." (the first cannot be checked; the second can be satisfied by writing more trivial tests, and a comfortable percentage can hide a single catastrophic failure) TRAP A pass-rate number as the only exit criterion. Pair every count-based criterion with a risk-coverage criterion and a severity rule, or you have written a target that rewards writing easy tests. -->
{{entry_and_exit_criteria}}
## Environment, Data, and Ownership
<!-- WHAT Where testing runs, what test data it needs and who provides it, and one named owner per area. WHY Test data is where plans quietly fail: the data that exercises a permission boundary or an error path never exists by accident, and manufacturing it is often the longest lead-time item in the whole effort. Naming it early is what stops week one being spent making accounts. Deep dive: test-plan_companion.md section 3 (Environment, Data, and Ownership). ASK Which environments, at which versions and configurations? What data is needed for the negative and boundary cases, and who creates it? What is anonymized or synthetic? Who owns each area, by name? GOOD "Staging, build 2.3.x, flag on, seeded nightly from an anonymized production snapshot. Data: three personas (owner, permitted viewer, restricted viewer) and two dashboards with a restricted filter field, created by Platform (Dana Osei) before entry. Owners: functional and permissions, Anjali Rao; performance, Marcus Bell; accessibility, Sofia Marino." WEAK "Test on staging with test data. QA owns testing." (no versions, no data specifics, no named person; "QA owns it" is owned by nobody) TRAP Naming a team instead of a person. A team owns nothing on a Friday afternoon. -->
{{environment_data_and_ownership}}
## Schedule and Deliverables
<!-- WHAT The timebox and milestones for the testing effort, and what testing will hand over at the end. WHY Deliverables set expectations about what exists when testing stops: executed cases, a defect list, a summary. Naming them is also what keeps this plan prospective; if results start appearing in it, it has quietly become an execution tracker and stopped being a plan. Deep dive: test-plan_companion.md section 3 (Schedule and Deliverables) and section 7. ASK When does testing start and stop, and against which milestones? What gates sit inside the window? What artifacts are handed over, to whom? Where does execution status actually live? GOOD "Window: 2026-07-06 to 2026-07-17, with the security review gate on 2026-07-10 between phase 1 and phase 2. Deliverables: executed case results in the test tool, an open-defect list, and a one-page summary to the release checklist on 2026-07-17." WEAK "Two weeks of testing before release. Deliverable: test results." (no gates, no owner of the handover, no statement of where results live) TRAP Adding status or results columns to this plan. Keep execution state in the tool or the report; a plan that tracks itself becomes stale in both roles. -->
{{schedule_and_deliverables}}test-plan_template-full.md · ~4,550 tokens
---title: "{{release_or_feature}} Test Plan"release_or_feature: "{{release_or_feature}}"plan_owner: "{{plan_owner}}"status: "{{status}}"last_updated: "{{date}}"test_strategy_ref: "{{test_strategy_ref}}"doc_type: test-plansize: fullsource_template: test-plansource_template_version: 0.1.0---
<!--FULL TEST PLAN. The approvable plan: everything the lean variant carries, plus the levels and types breakdown,the rule for stopping and restarting mid-cycle, the risks to the testing itself, and the sign-off record. Useit when the context genuinely demands it: regulated or audited work, several teams or vendors testing onerelease, a formal gate inside the cycle, or a process that requires an approved artifact.
THIS VARIANT IS A STRICT SUPERSET OF THE LEAN ONE. The five lean sections appear here in the same order, withthe same headings and the same placeholders, and four sections are added. If you started lean, you can growinto this without rewriting anything you already filled in. (The guidance comments in the shared sectionscarry a little extra context for the approvable case; the content you wrote does not change.)
LENGTH IS NOT RIGOR. A nine-section plan padded to look complete is the exact failure the critics ofstandardized test documentation name: a document can satisfy every required heading and still say nothing.Every section here should decide something. If one would change nothing by its absence, write "N/A" and oneline of why. See test-plan_companion.md sections 4 and 6.
WHAT A TEST PLAN IS, AND IS NOTIt is the prospective document that scopes and prioritizes a testing effort. It is NOT a test case (that isone executable verification), NOT a test report (that is retrospective, written after), NOT a definition ofdone (that is a standing team invariant), and NOT the thing your test management tool calls a "test plan"(that is an execution container whose only required field is usually a name). Seetest-plan_companion.md section 8.
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into test-plan_companion.md), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid. For tables, PRIORITY explains the ordering rule and ROW HINT says what a good row contains.2. Replace each {{placeholder}} with your content. Fill Scope and Risk-Ranked Approach first; everything else follows from them.3. If a section does not apply, write "N/A" and one line of why, rather than deleting it silently.4. Before you circulate it for approval: self-grade against test-plan_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{release_or_feature}} Test Plan
## Scope and Non-Scope
<!-- WHAT What is being tested (the concrete test items: builds, services, flag states, requirement IDs) and what is explicitly NOT being tested this cycle, with a reason for each exclusion. WHY The exclusion half is the highest-value content in the document and the most commonly deleted. If nothing is out of scope, nothing can be found missing, and the scope argument happens after the release instead of before it. In an approvable plan it is also what the sign-off is actually signing. Deep dive: test-plan_companion.md section 3 (Anatomy > Scope and Non-Scope). ASK What exactly is under test, by build or requirement ID? What is deliberately excluded, and why (deferred, covered elsewhere, owned by another team, accepted risk)? Who agreed to the exclusions? GOOD "In scope: FR-1 to FR-5 of the Saved Views PRD, on build 2.3.x behind the saved_views flag. Out of scope: FR-6 (stale-view indicator), Could priority, deferred to the next release; and the permissions service itself, owned and tested by Platform." WEAK "Testing the Saved Views feature." (no test items, no exclusions, nothing anyone can disagree with or be held to) TRAP Leaving out the exclusions because they feel negative. An unstated exclusion is not a smaller scope, it is an unmanaged expectation. -->
{{scope_and_non_scope}}
## Risk-Ranked Approach
<!-- WHAT The areas under test, ranked by product risk, with the reason for each ranking and the depth of testing it earns. This is the section that makes the plan a plan. WHY Risk-based testing exists to steer effort rather than spread it evenly: high-risk areas earn heavier techniques and are tested FIRST, so that if time runs out what is missing is the least important coverage rather than a random sample. A section that only names test types ("functional, regression") decides nothing. Deep dive: test-plan_companion.md section 3 (Risk-Ranked Approach) and section 5. ASK What could go wrong here, and how bad would it be? Which areas therefore carry the most risk? What technique and depth does each risk tier earn? What order will they be tested in? Where do these risks come from (the PRD, the risk register, a risk-storming session)? PRIORITY Order the table by risk, highest first, and test in that order. Reuse product risks already recorded upstream rather than inventing a parallel list. Risks that threaten the TESTING rather than the product belong in "Risks to the Test Effort" below, not here. ROW HINT A good row names an area, states the risk as a consequence rather than a theme, gives a tier with a reason, and names the technique and depth. A weak row is a feature name with "high" beside it. GOOD | Shared-view permissions | A shared view could expose rows the recipient may not see | High: low likelihood, severe and irreversible impact (data exposure) | Permission matrix across 3 personas x 4 filter scopes; negative cases first; security review gate | 1 | WEAK | Sharing | High | Test it thoroughly | 1 | TRAP Ranking everything "high". If every area is top priority, the ranking has told the team nothing and the order of work is back to guesswork. -->
| Area under test | Risk (what could go wrong, and the consequence) | Tier and why | Technique and depth | Order ||---|---|---|---|---|| {{area}} | {{risk_and_consequence}} | {{tier_and_reason}} | {{technique_and_depth}} | {{order}} |
## Test Levels and Types
<!-- WHAT Which test levels (unit, integration, system, end-to-end) and which types (functional, performance, security, accessibility, compatibility) are in play, who owns each, and which are out. WHY A lean plan omits this because one team already knows. It earns its place the moment more than one group tests, a level belongs to somebody else, or a non-functional type needs its own environment, data or specialist. Writing it down is what stops two teams both assuming the other covered integration. Deep dive: test-plan_companion.md section 3 (Test Levels and Types). ASK Which levels are in scope for this effort, and which are assumed already covered? Which non-functional types are required, and by what evidence (an NFR, a regulation, a past incident)? Who owns each row? Which types are explicitly not being run? PRIORITY List only levels and types that someone will actually run or explicitly waive. A row with no owner is a gap, not a plan. ROW HINT A good row names the level or type, its owner by name, the environment it needs, and the evidence that it is required. A weak row is a checkbox with no owner. GOOD | System (API) | Anjali Rao | Staging | Required: the views endpoints are new in this release | WEAK | Performance | TBD | | Nice to have | TRAP Listing every level and type the organization has ever run. This section is a scope statement, not a catalog; unowned rows read as coverage that will not happen. -->
| Level or type | Owner | Environment | Why it is required (or waived) ||---|---|---|---|| {{level_or_type}} | {{level_owner}} | {{level_environment}} | {{level_rationale}} |
## Entry and Exit Criteria
<!-- WHAT The conditions that must be true before testing starts, and the conditions that define testing as done, with who agreed the exit criteria and when. WHY Criteria are thresholds, not adjectives. "Environment is ready" cannot be checked; a named build in a named state can. The functional test for an entry criterion is whether failing it would actually stop you: if not, it is not a criterion. Exit criteria are supposed to be agreed with stakeholders, and almost every template omits the evidence of that agreement, which is exactly what an approvable plan needs to carry. Deep dive: test-plan_companion.md section 3 (Entry and Exit Criteria) and section 7. ASK What must exist before testing can start without wasting effort? What would make you stop and say testing is finished? How is each condition measured, and by whom? Who agreed the exit criteria, and on what date? GOOD "Entry: build 2.3.1 deployed to staging with the flag on; three permission personas seeded; smoke suite green. Exit: every High-tier area has its planned cases executed; zero open Sev-1 or Sev-2 defects in scope; the permission matrix 100 percent executed with no failures. Agreed with Priya Nair (PM) and Marcus Bell (Eng lead) on 2026-07-06." WEAK "Entry: environment ready. Exit: 95 percent of tests pass." (the first cannot be checked; the second can be satisfied by writing more trivial tests, and a comfortable percentage can hide a single catastrophic failure) TRAP A pass-rate number as the only exit criterion. Pair every count-based criterion with a risk-coverage criterion and a severity rule, or you have written a target that rewards writing easy tests. -->
{{entry_and_exit_criteria}}
## Suspension and Resumption Criteria
<!-- WHAT What stops testing mid-cycle, who makes that call, and what has to be true to restart. WHY This is the most commonly dropped section of the classic standard and the one teams wish they had written when a blocking defect lands on a Thursday afternoon. Deciding it calmly in advance is worth more than deciding it under pressure, particularly where a gate exists or an environment is shared. Deep dive: test-plan_companion.md section 3 (Suspension and Resumption Criteria). ASK What class of finding should stop the cycle rather than just be logged? Who decides, and who is told? What must be true to resume, and does resuming require re-running anything already passed? Which failures pause only part of the effort rather than all of it? GOOD "A confirmed permission leak suspends all phase 2 (sharing) testing immediately; Anjali Rao calls it and notifies Priya Nair and Marcus Bell the same day. Resumption requires a fix, a passing permission matrix re-run in full, and the security review signed. Phase 1 testing continues throughout." WEAK "Testing will be suspended if there are too many defects." (no threshold, no decider, no resumption condition, and no statement of what keeps running) TRAP Writing a suspension rule with no named decider. In practice the cycle then continues by default while people wait for someone to say stop. -->
{{suspension_and_resumption}}
## Environment, Data, and Ownership
<!-- WHAT Where testing runs, what test data it needs and who provides it, and one named owner per area. WHY Test data is where plans quietly fail: the data that exercises a permission boundary or an error path never exists by accident, and manufacturing it is often the longest lead-time item in the whole effort. Naming it early is what stops week one being spent making accounts. Deep dive: test-plan_companion.md section 3 (Environment, Data, and Ownership). ASK Which environments, at which versions and configurations? What data is needed for the negative and boundary cases, and who creates it? What is anonymized or synthetic, and under what policy? Who owns each area, by name? GOOD "Staging, build 2.3.x, flag on, seeded nightly from an anonymized production snapshot. Data: three personas (owner, permitted viewer, restricted viewer) and two dashboards with a restricted filter field, created by Platform (Dana Osei) before entry. Owners: functional and permissions, Anjali Rao; performance, Marcus Bell; accessibility, Sofia Marino." WEAK "Test on staging with test data. QA owns testing." (no versions, no data specifics, no named person; "QA owns it" is owned by nobody) TRAP Naming a team instead of a person. A team owns nothing on a Friday afternoon. -->
{{environment_data_and_ownership}}
## Schedule and Deliverables
<!-- WHAT The timebox and milestones for the testing effort, and what testing will hand over at the end. WHY Deliverables set expectations about what exists when testing stops: executed cases, a defect list, a summary, and any sign-off artifact. Naming them is also what keeps this plan prospective; if results start appearing in it, it has quietly become an execution tracker and stopped being a plan. Deep dive: test-plan_companion.md section 3 (Schedule and Deliverables) and section 7. ASK When does testing start and stop, and against which milestones? What gates sit inside the window? What artifacts are handed over, to whom, and by when? Where does execution status actually live? GOOD "Window: 2026-07-06 to 2026-07-17, with the security review gate on 2026-07-10 between phase 1 and phase 2. Deliverables: executed case results in the test tool, an open-defect list, and a one-page summary to the release checklist on 2026-07-17." WEAK "Two weeks of testing before release. Deliverable: test results." (no gates, no owner of the handover, no statement of where results live) TRAP Adding status or results columns to this plan. Keep execution state in the tool or the report; a plan that tracks itself becomes stale in both roles. -->
{{schedule_and_deliverables}}
## Risks to the Test Effort
<!-- WHAT The risks that threaten the TESTING, with an owner and a contingency for each. Not the product risks that shape coverage; those are in the Risk-Ranked Approach above. WHY These are two different lists, and merging them is why "risks" sections read as noise. Product risk: the permission check might leak data, so test it hardest. Project risk: the staging environment is shared and might be unavailable in week two, so the schedule needs a contingency. The first shapes what you test; the second shapes whether you get to test at all. Deep dive: test-plan_companion.md section 3 (Risks to the Test Effort). ASK What could delay, block or invalidate the testing itself? How likely is it, and what would it cost in days or coverage? Who owns it? What is the contingency, and what is the trigger for invoking it? PRIORITY Order by expected damage to the effort. Every row needs a named owner and a contingency that someone could actually execute; a risk with no contingency is just an anxiety. ROW HINT A good row names the threat to testing, the impact in concrete terms (days, coverage lost), a named owner, and a contingency with a trigger. A weak row names a worry with no owner. GOOD | Staging shared with the Billing migration in week 2 | Up to 3 days of blocked execution, losing the performance runs | Dana Osei | Book the window now; if lost, run performance against the pre-prod replica and note the deviation in the summary | WEAK | Environment issues | Could delay testing | QA | Escalate | TRAP Copying the product risk table into this section. If a row would change what you test rather than whether you can test, it belongs above. -->
| Risk to the testing | Impact if it happens | Owner | Contingency and trigger ||---|---|---|---|| {{effort_risk}} | {{effort_risk_impact}} | {{effort_risk_owner}} | {{effort_risk_contingency}} |
## Approvals and Change Control
<!-- WHAT Who approved this plan and when, and how it changes once approved. WHY This section exists for regulated, audited and multi-team contexts, and can be cut everywhere else. Its honest purpose is not ceremony: a plan that changes silently after sign-off is worse than one never signed, because it carries borrowed authority. The change rule is the part people forget, and it is the part that matters on day nine. Deep dive: test-plan_companion.md section 3 (Approvals and Change Control) and section 9. ASK Who must approve this, in what role, and what are they actually attesting to? What kind of change requires re-approval rather than an edit? Where is the version history kept? Who is told when the plan changes? PRIORITY Approvers are named individuals with roles, not distribution lists. State plainly which changes need re-approval; "material change" without a definition is not a rule. ROW HINT A good row names a person, a role, what they are approving, and a date. A weak row is a role with no name and no date. GOOD | Priya Nair | PM, Reporting | Scope, exclusions and exit criteria | 2026-07-06 | WEAK | Product | | Approved | | TRAP Collecting approvals on a plan nobody has read. A signature on an unread document transfers blame rather than creating agreement. -->
| Approver | Role | What they are approving | Date ||---|---|---|---|| {{approver}} | {{approver_role}} | {{approval_scope}} | {{approval_date}} |
{{change_control_rule}}---title: "Saved Views for Dashboards Test Plan"release_or_feature: "Saved Views for Dashboards (phases 1 and 2)"plan_owner: "Anjali Rao (QA Lead, Reporting)"status: "approved"last_updated: "2026-07-06"test_strategy_ref: "None. Acme Reporting has no standing test strategy document; the Risk-Ranked Approach below is the strategy for this release."related: - "../prd/prd_example.md (Saved Views for Dashboards PRD, the scope this plan tests)" - "../sdd/sdd_example.md (Saved Views design, the technical surface under test)" - "../acceptance-criteria/acceptance-criteria_example.md (acceptance criteria for the default-view story)" - "../risk-register/risk-register_example.md (program risk register; R-05, R-06 and R-02 are inherited below)"doc_type: test-plansize: fullsource_template: test-plansource_template_version: 0.1.0---
<!--Worked example for the test-plan bundle: a realistic, fully filled full-variant test plan for one feature.It chains onto the delivery-docs examples (same company, same feature, same cast) so the reader can followone thread from PRD through acceptance criteria into verification. Figures marked "illustrative" are made upfor the example and would be real data in a live plan.
Why the FULL variant here: this release has a formal gate inside the cycle (the security review that the PRDmakes a precondition of phase 2), three groups testing (Reporting QA, Platform, Design Systems), and amigration whose rollback needs rehearsing. Any one of those would justify it. A single-team feature with nogate should use the lean variant.-->
# Saved Views for Dashboards Test Plan
## Scope and Non-Scope
**In scope.** Functional requirements FR-1 to FR-5 of the [Saved Views PRD](../prd/prd_example.md), on build2.3.x with the `saved_views` flag enabled, covering both rollout phases: phase 1 (private views: save, list,switch, set default, rename, delete) and phase 2 (sharing). The concrete test items are the five`ViewsController` REST endpoints described in the [design document](../sdd/sdd_example.md), the `saved_view`table migration and its rollback, the `default_view_id` addition to the per-user preferences record, and theViews control in the dashboard frontend. Non-functional coverage in scope: the shared-view entitlementboundary, view-list load performance, graceful degradation when a saved config references a deleted filterfield, and WCAG 2.2 AA keyboard and screen-reader operation of the Views control.
**Out of scope, and why.**
- **FR-6, the stale-view change indicator.** Could priority in the PRD and not built in this release. Deferred with the requirement; there is nothing to test.- **The dashboard permissions service itself.** Owned and tested by Platform. This plan tests *our use of it* (that a shared view re-checks the recipient's access on read), not the service's own correctness.- **Adoption.** Risk R-04 on the program register is an adoption risk, and adoption is measured after launch on the KPI dashboard, not verified by testing. Naming it here stops the recurring question of whether QA is covering it.- **The legacy report flow.** Unchanged by this release; covered by the existing regression suite, which runs but is not re-planned here.
Exclusions agreed with Priya Nair (PM, Reporting) and Marcus Bell (Staff Engineer, Reporting) on 2026-07-06.
## Risk-Ranked Approach
Product risks are inherited from the [program risk register](../risk-register/risk-register_example.md) andthe PRD's own risk and non-functional tables rather than invented here, so a change in either flows into thisplan instead of diverging from it. Areas are tested in the order shown: if the window is cut short, what islost is the bottom of this table, by design.
| Area under test | Risk (what could go wrong, and the consequence) | Tier and why | Technique and depth | Order ||---|---|---|---|---|| Shared-view entitlement | A shared view embeds filter values, so a recipient without entitlement could see a segment (and the PII in it) they may not access, causing a reportable incident | **High.** Register R-05, escalated: residual 8 exceeds the program's PII appetite line of 6. Low likelihood, severe and irreversible impact | Full permission matrix: 3 personas (owner, permitted viewer, restricted viewer) x 4 filter scopes, negative cases executed first; re-check asserted on *recipient read*, not at share time, per the design; PII-in-filter scan on every share path | 1 || Config migration and rollback | Saved-view configs move from the legacy key-value store to the `saved_view` schema, so a silent conversion failure loses analysts' views at cutover | **High.** Register R-02. Data loss is not user-recoverable and the blast radius is every existing view | Migration dry run with count reconciliation (expect 0 mismatch); rollback rehearsal to the read-only legacy store; `config` schema version 1 validation on read of a v0 row | 2 || Stale-field degradation | A saved config references a filter field that no longer exists, so the dashboard fails to load instead of loading what it can | **Medium-high.** PRD reliability NFR and the design's `stale_fields` path. Moderate likelihood (fields do get removed), contained impact | Error-path testing: delete a referenced field, assert the resolvable parts load, `stale_fields` is populated, and the PRD's "some filters no longer exist" message appears with a route to re-save | 3 || View-list load performance | A dashboard accumulates many views, so the view list exceeds the 500ms budget and degrades the speed the program is selling | **Medium.** Register R-06, residual 6. Degrades rather than breaks | Load test at 3x the expected view count (illustrative target: 150 views on one dashboard), p95 measured separately on the view-list endpoint (register R-06 budget: under 500ms) and on view switch (PRD non-functional requirement: under 1s at p95) | 4 || Default-view resolution | The per-user default resolves wrongly, so a user opens someone else's view or the generic default | **Medium.** Touches the acceptance criteria directly, but failures are visible and recoverable | State and precedence cases from the [acceptance criteria](../acceptance-criteria/acceptance-criteria_example.md): setting a new default clears the previous one; a default that is a shared view the user can no longer read falls back cleanly; one user's default never changes another's | 5 || Accessibility of the Views control | The control is not keyboard-operable or not labeled, so the feature is unusable with assistive technology and breaches the PRD's stated WCAG 2.2 AA target | **Medium.** Certain to matter if wrong; caught late is expensive to fix | Keyboard-only traversal of the full menu, screen-reader label and state announcement (NVDA and VoiceOver), focus management on open and close | 6 || Rename and delete | A rename or delete fails or affects the wrong row | **Low.** Simple owner-only operations on a single row, easily observed | Equivalence partitioning only: owner and non-owner, existing and deleted row | 7 |
## Test Levels and Types
| Level or type | Owner | Environment | Why it is required (or waived) ||---|---|---|---|| API (system) | Anjali Rao | Staging | The four views endpoints are new in this release; the entitlement re-check is only observable here || End-to-end (browser) | Anjali Rao | Staging | The Views control, the default-on-open path and the stale-field message are user-visible behaviors || Migration and rollback | Lee Zhang (Data Eng) | Migration sandbox, then staging | Register R-02; a rollback that has never been rehearsed is not a rollback || Performance | Marcus Bell | Pre-prod replica | Register R-06 and the PRD's p95 targets; staging is too small to be representative || Security review | Sam Okafor (Security) | n/a (review, not execution) | Made a precondition of phase 2 by the PRD rollout plan; gates sharing || Accessibility | Sofia Marino (Design Systems) | Staging | PRD non-functional requirement: WCAG 2.2 AA || Unit and component | Marcus Bell | CI | Waived as a planned activity here: covered by the team's existing CI gate, which must be green as an entry criterion below || Localization | n/a | n/a | Waived: this release adds no user-facing copy beyond two strings already in the translation pipeline |
## Entry and Exit Criteria
**Entry criteria** (all must hold before execution starts; each would genuinely stop the work if it failed):
1. Build 2.3.1 or later deployed to staging with the `saved_views` flag enabled.2. The CI unit and component gate is green on that build.3. Three permission personas seeded (owner, permitted viewer, restricted viewer) and two dashboards provisioned, one of which carries a restricted filter field. Provided by Platform (Dana Osei).4. The migration dry-run dataset is loaded in the migration sandbox.5. The smoke suite passes on staging.
**Exit criteria** (all must hold before testing is called done):
1. Every High and Medium-high tier area in the Risk-Ranked Approach has its planned cases executed. Not "most cases": these three areas, complete.2. The permission matrix is 100 percent executed with zero failures. This one is absolute; a single failure here is a suspension event, not a defect to triage.3. Zero open Sev-1 or Sev-2 defects against in-scope requirements.4. Migration reconciliation shows a zero-row mismatch on the dry run, and the rollback rehearsal has been completed once end to end.5. p95 view switch under 1s and p95 view-list load under 500ms on the pre-prod replica at 3x expected view count (illustrative thresholds, taken from the PRD and the register).6. Accessibility findings at AA level are either fixed or accepted in writing by Priya Nair.
Deliberately **not** an exit criterion: an overall pass-rate percentage. A pass rate can be raised by writingmore shallow cases and would hide a single catastrophic permission failure inside a comfortable number. Everycriterion above is a coverage, severity or threshold statement instead.
**Agreed** with Priya Nair (PM, Reporting), Marcus Bell (Staff Engineer, Reporting) and Sam Okafor (Security)on 2026-07-06. Any change to criterion 2 or 3 requires re-agreement by all three.
## Suspension and Resumption Criteria
**Suspension.** A confirmed entitlement failure (any case in which a recipient can see data through a sharedview that they cannot see directly) suspends **all phase 2 sharing testing** immediately. Anjali Rao makes thecall and notifies Priya Nair, Marcus Bell and Sam Okafor the same day. Phase 1 (private views) testingcontinues, because it does not exercise the sharing path.
Testing is also suspended, wholly, if the staging environment loses the seeded permission personas, sinceevery High-tier case depends on them and results produced without them would be misleading rather than merelyabsent.
**Resumption.** Sharing testing resumes when: the defect is fixed and deployed; the **entire** permissionmatrix is re-run from the start, not just the failing case, because an entitlement bug invalidates theassumption behind every passing result in that matrix; and Sam Okafor confirms the security review isunblocked. Environment loss resumes on re-seeding plus a green smoke run.
## Environment, Data, and Ownership
**Environments.** Staging on build 2.3.x with the flag enabled, reseeded nightly from an anonymized productionsnapshot, for functional, end-to-end and accessibility work. A pre-prod replica sized to production forperformance, because staging carries roughly a tenth of the data (illustrative) and would produce reassuringnumbers that mean nothing. A separate migration sandbox holds a copy of the legacy key-value store for thedry run and rollback rehearsal.
**Test data**, and this is the long-lead item:
- Three permission personas: an owner, a permitted viewer, and a restricted viewer who lacks access to one filter field used in a shared view. Created by Platform (Dana Osei) before entry criterion 3 can be met.- Two dashboards, one carrying a restricted filter field, one carrying a field that will be deleted mid-cycle to exercise the stale-field path.- 150 saved views on a single dashboard for the performance run (illustrative: 3x the expected ceiling).- A legacy-store extract with 12 known-bad configs (illustrative) for migration reconciliation, including two that reference deleted fields.
The anonymized snapshot does **not** contain a restricted filter field by default; that is why persona data ismanufactured rather than sampled, and why it is called out as a risk to the effort below.
**Ownership.** Functional, end-to-end and the permission matrix: Anjali Rao. Migration and rollback: LeeZhang. Performance: Marcus Bell. Accessibility: Sofia Marino. Security review: Sam Okafor. Test dataprovisioning: Dana Osei. Plan ownership and the suspension call: Anjali Rao.
## Schedule and Deliverables
**Window:** 2026-07-06 to 2026-07-17 (illustrative), aligned to the PRD's phased rollout.
| Milestone | Date | Gate ||---|---|---|| Entry criteria met, phase 1 execution starts | 2026-07-06 | Smoke green, personas seeded || Migration dry run and rollback rehearsal complete | 2026-07-09 | Zero-row reconciliation mismatch || Security review | 2026-07-10 | **Gates phase 2**: no sharing testing before it, per the PRD rollout plan || Phase 2 (sharing) execution | 2026-07-13 to 2026-07-16 | Permission matrix first || Performance run on pre-prod replica | 2026-07-15 | 3x view count loaded || Exit review and handover | 2026-07-17 | Exit criteria assessed with the three agreeing stakeholders |
**Deliverables.** Executed case results in the test tool (the tool's own "test plan" record for sprint 14, notthis document); an open-defect list with severities; a one-page test summary against the exit criteria,delivered to the release checklist on 2026-07-17; and the migration rollback rehearsal record, which therelease checklist requires separately.
Execution status lives in the test tool, not in this plan. This document is the intent; it is revised when theintent changes, not when a case passes.
## Risks to the Test Effort
These threaten the testing rather than the product; product risks are in the Risk-Ranked Approach above.
| Risk to the testing | Impact if it happens | Owner | Contingency and trigger ||---|---|---|---|| Staging is shared with the Billing migration in week 2 | Up to 3 days of blocked execution (illustrative), landing on the phase 2 window when the permission matrix runs | Dana Osei | Window booked 2026-06-29 for 13-16 July. Trigger: Billing requests the environment. Fallback: run the permission matrix against the pre-prod replica and record the environment deviation in the summary || The anonymized snapshot has no restricted filter field, so persona data must be manufactured | Entry criterion 3 unmet; every High-tier case blocked from day one | Dana Osei | Manufactured persona set built and verified by 2026-07-03, ahead of entry. Trigger: verification fails. Fallback: hand-built fixtures in the migration sandbox, accepting reduced realism || Sam Okafor is the only security reviewer and is single-threaded across two programs | The 2026-07-10 gate slips, and phase 2 cannot start; the whole sharing scope is at risk of leaving the window | Marta Reyes | Review slot confirmed in writing. Trigger: no confirmation by 2026-07-08. Fallback: escalate to the security lead for a second reviewer; if unavailable, ship phase 1 alone and re-plan phase 2 || The performance replica is refreshed mid-window, changing the data profile | Performance results across the window are not comparable, and R-06 stays unverified | Marcus Bell | Freeze the replica for the window. Trigger: a refresh is scheduled. Fallback: re-run the whole performance set after the refresh rather than comparing across it |
## Approvals and Change Control
| Approver | Role | What they are approving | Date ||---|---|---|---|| Priya Nair | PM, Reporting | Scope, the exclusions, and the exit criteria | 2026-07-06 || Marcus Bell | Staff Engineer, Reporting | Technical scope, environments, and the entry criteria | 2026-07-06 || Sam Okafor | Security | The entitlement approach, the suspension rule, and the phase 2 gate | 2026-07-06 || Marta Reyes | Program Manager | Schedule, and the contingencies for risks to the effort | 2026-07-06 |
**Change control.** Edits to wording, dates within the agreed window, and additions to the test data list aremade in place by the plan owner and noted in the version history. The following require re-approval by thenamed approver before they take effect: any change to the in-scope requirement list or the exclusions (PriyaNair); any relaxation of exit criteria 2 or 3 (all three signatories to the criteria); any change to thesuspension rule or the phase 2 gate (Sam Okafor). A change made without the required re-approval leaves theplan unapproved until it is obtained, and the release checklist treats an unapproved plan as a blocker.Provenance
Section titled “Provenance”The reasoning, the history and every source, in the repository:
- Companion - the long-form argument: why these sections, where the sources disagree, and what the bundle refuses to claim
- History - what changed in this bundle, and when
- Research log - every source consulted, with what each one actually supports
- Catalog metadata - the machine-readable record this page is generated from
Catalog record: 9 sections across 1 format(s), methodology methodology-agnostic, typically owned by QA Lead.