Spike Report
beta · Family decision-docs · Phase develop · Sizes lean · ~4,550 tokens
The written output of a time-boxed technical investigation: the uncertainty being reduced, what was actually tried, what was actually found, a proceed-or-not recommendation, and the edge of what that answer covers. Investigates the question that precedes a decision, as distinct from an RFC (which proposes one), an ADR (which records one), and an SDD (which describes the design that implements one). Whether a spike produces a written document at all is contested: the term’s Extreme Programming inventors describe its output as throwaway code, and one named source publishes it as a document.
The short card. Why the document is shaped this way, and the argument behind every rule here, is in
spike-report_companion.md. A fully worked instance is
spike-report_example.md.
When to use
Section titled “When to use”Read this first, because it changes what the card is claiming. This is the one document type in this library whose own canon argues against writing it: in Extreme Programming a spike produces code you throw away, and Ward Cunningham’s founding account says “We plan to throw away the code, although sometimes something is salvaged.” Exactly one named source in this bundle’s research publishes the deliverable as a document, Microsoft’s Code with Engineering Playbook: “Generally the deliverable from a Technical Spike should be a document detailing what was evaluated and the outcome of that evaluation.” So nothing here says that writing up your spike is settled practice. It says how to write a good one if you are writing one. Both sides: companion section 6.
- A named decision is waiting on an answer nobody has, and someone other than you will act on it.
- The uncertainty is knowledge-limited, not time-limited, which is the founding XP wiki’s own condition for reaching for a spike at all: “Spikes are good when you are knowledge-limited, not time-limited”. More hours of the current approach will not produce the answer; a different, bounded experiment will.
- The result has to be checkable by someone who was not there, rather than taken on your word.
- The evidence will outlive your memory of it and get cited. A later decision record, proposal or design document will lean on this rather than reconstruct it.
- The answer might be no, and a no needs to be as findable as a yes. A refuted hypothesis is a successful spike.
When NOT to use
Section titled “When NOT to use”Write something else if:
| You actually need | Because |
|---|---|
| an ADR | the decision is being recorded, not investigated. It has already been made, and you are writing down what was chosen and what it costs. This is the drift to watch for: a spike report that arrives at a recommendation without recording what was tried is a bad ADR, and an ADR that shows its working is not a spike report. If the call is genuinely made, write the record and stop |
| an RFC | you are proposing a change and asking people to shape it, not investigating whether it is possible. An RFC argues a case and requests input; a spike report hands over evidence and gets out of the way. They are a sequence rather than a choice: what a spike hands over is what an RFC’s motivation can then argue from |
| an SDD | you are describing a design: how the thing will be built, its components and their interactions. That document assumes the feasibility question is behind you. If it is not, the spike comes first and the design document cites it |
| nothing | the question can be answered by reading the documentation for ten minutes. Then read it, and spend the box on something that is actually uncertain. A spike is for an uncertainty that reading cannot close, and a spike with no answerable question is a research project wearing a time box |
Write nothing at all if no named decision is waiting on the answer. A ticket and a write-up created to demonstrate activity rather than to reduce uncertainty is a failure mode this bundle calls the performative spike, and the test for it is simple: somebody should be blocked on the result, and you should be able to say who.
Pick a variant
Section titled “Pick a variant”There is no choice to make. This bundle ships one file, spike-report_template-lean.md, and that is a
finding rather than a shortcut: no source read for this bundle publishes two weights of a spike report, and
the one named source that publishes the document at all ships exactly one template. The variation the
research did find runs across genre, not weight. The longest templates found were written for an AI
coding agent to fill in, and their section counts should not be read as normal for a human team; at the other
end, one real published engineering process ships no document at all and closes the spike issue with a single
summary comment instead. The case is in
companion section 4.
The practice’s own discipline argues against a heavier variant anyway. A spike is bounded at a couple of days at the canonical end and one iteration at the outer end, and a document whose ceremony costs a meaningful fraction of the investigation that produced it has inverted the point of the practice. (That last sentence is this bundle’s judgment; no source read makes the argument directly.)
Quality rubric (self-grade before you hand it over)
Section titled “Quality rubric (self-grade before you hand it over)”Score each 0, 1 or 2. Under 13 out of 18 and the report hands its reader a verdict they cannot check, so they either take your word for it or spend their own box repeating your work. Both outcomes waste the one you just spent.
| # | Criterion | 0 | 1 | 2 |
|---|---|---|---|---|
| 1 | Question answerable | Names a topic, or an activity to perform | A question, but no result could have made it false | A stranger can say what result would have refuted it, and who was waiting on the answer |
| 2 | Box reported both ways | Only the word timeboxed, or an allotment with no spend | Both figures are present, but an overrun or an early finish is left unexplained | Allotted and spent are both there, and where they differ the document says what happened |
| 3 | Non-scope is explicit | Nothing says what was left alone | An exclusion is named, but a branch abandoned partway is not | A reader can name what was deliberately not attempted, including anything dropped mid-investigation, and the signal that stopped it |
| 4 | Method reproducible | Approaches or vendors named with no versions, data or commands | Enough detail to guess at the approach, not enough to repeat it | A colleague could re-run it from what is written: versions, environment and data pinned, and any evidence reused from earlier work flagged as reused |
| 5 | Facts apart from inference | Observation and interpretation share the same sentences | Separated by heading, but at least one inference is still stated as an observation | Every claim is visibly one or the other, and the confidence in an interpretive step is stated wherever it is not obvious |
| 6 | Retractions in place | An earlier finding changed and nothing in the document says so | A correction is mentioned, but the superseded claim has been deleted | Anything that turned out wrong is marked where it stood, saying what changed, so a reader can tell which conclusions moved and which never did |
| 7 | Verdict is explicit | The reader has to infer it from tone | Proceed or stop is stated, but nothing says what it rests on | The first line says proceed, do not proceed, or names the next spike, and any condition is written so a reader knows what would flip it |
| 8 | Boundary written down | The section is empty, or marked N/A on a spike that got an answer | It lists what the team plans to do next | It names what a reader must not conclude, and the dataset, version or platform the answer was actually measured on |
| 9 | Askable later | No investigator, no contact, no date | A team is named rather than a person | A person, a way to reach them, a date and a stable identifier, so a later document can cite this instead of paraphrasing it |
The test behind every cell above: could someone satisfy it without improving the document? A row that counted findings, or counted exclusions, would reward padding, and five bulleted things nobody looked at would score the same as one honest sentence. Every cell instead asks whether a specific piece of evidence exists and whether a second person, not the author, could find it.
Named anti-patterns (the usual wrecks)
Section titled “Named anti-patterns (the usual wrecks)”- The ADR with a longer preamble. The sharpest failure for this type, and the easiest to commit. The verdict was reached before the box opened and the evidence sections exist to justify it; the tell is a Recommendation you could have written on day zero. If the decision is genuinely already made, write an ADR and stop pretending the investigation was open.
- The recommendation with no working. A verdict is there, and What Was Tried is missing, vague or unrepeatable, so nobody can check the reasoning that produced it. This is the cost Microsoft’s playbook is warning about when it argues “The goal of a spike should be fact-finding, not decision-making or recommendation.” This template keeps the Recommendation section deliberately, and this anti-pattern is the price of keeping it.
- The unanswerable question. A spike aimed at a topic rather than a question. The box expires, the report describes the terrain, and nothing is decided. Three questions is three spikes; a question no evidence could refute is a research project that a time box will not rescue.
- The silently dropped branch. You abandoned a promising approach at hour three and never mentioned it, so the next team spends their whole box rediscovering that it does not work. You stopped for a reason, and the reason is evidence. It belongs beside the budget that forced it.
- Findings with no anchor. No counts, no versions, no commands, no file references. A reader who disagrees has nothing to check, so the disagreement gets settled by seniority rather than by evidence, and the report has become an opinion with formatting.
- The bounded answer with no boundary. The result held on one dataset, one version and one platform, and the report never said so, so the next reader takes a narrow answer as a general one. An empty What This Does Not Settle on a spike that believes it answered its question is worth a second look before you accept it as clean. Why that section is here at all, and the three caveats that travel with the evidence behind it, are in companion section 6.
Pairing with a skill
Section titled “Pairing with a skill”pairs_with: [develop-spike-summary]. The skill documents a completed spike, and it draws the same boundary
this card does: it sends you to develop-adr for the architecture decision the spike informs, and to
discover-interview-synthesis when the exploration was user research rather than technical feasibility. That
agreement is worth noticing, because it is one edge two independently built artifacts describe the same way.
Be aware that the two shapes differ. The skill’s description names the four things it captures - the original question, the approach, evidence-backed findings, and a proceed-or-not recommendation - and this template keeps all four. What its own output template looks like is not recorded in this bundle’s research, so no claim is made here about its section list; if you are moving content between the two, read the skill and compare for yourself. The difference worth knowing in advance is the one this template is opinionated about:
- The time box. The skill makes documenting allocated against actual time its own instruction step. This bundle folds the same content into Scope and Time Box, on the evidence that a dedicated time-box heading belongs to pre-spike planning artifacts rather than to the report of a completed spike.
- Open questions and follow-ups are two sections there, and one here. This bundle ships What This Does Not Settle and deliberately keeps it out of roadmap territory. The open questions transfer; the follow-up items mostly belong somewhere else, in the backlog rather than in this document.
- Neither shape asks what was deliberately not attempted. The skill has no non-scope section, which is exactly consistent with this bundle’s research: not one blank template found asks for one, and the good filled reports supply it anyway. That half of Scope and Time Box is content you will have to write without prompting from either tool.
The artifacts
Section titled “The artifacts”spike-report_template-lean.md · ~4,550 tokens
---title: "{{title}}"spike_id: "{{spike_id}}"status: "{{status}}"conducted_by: ["{{investigator_and_contact}}"]date: "{{date}}"work_item: "{{work_item_link}}"doc_type: spike-reportsize: leansource_template: spike-reportsource_template_version: 0.1.0---
<!--LEAN SPIKE REPORT. This bundle ships one size, and the evidence for that is unusually clean: no sourceread for this bundle publishes two weights of a spike report, and the one named source that publishesthe document at all ships exactly one template. A document whose ceremony costs a meaningful fractionof the investigation that produced it has inverted the point of the practice (this bundle's judgment,not a sourced claim). See spike-report_companion.md section 4 (Variants and sizing).
READ THIS BEFORE YOU FILL IT IN, BECAUSE IT CHANGES WHAT THIS TEMPLATE CLAIMS.The term's own inventors describe a spike's output as throwaway code, not a document. WardCunningham's founding account, crediting Kent Beck with the name, says plainly: "We plan to throw awaythe code, although sometimes something is salvaged." Mike Cohn describes an activity, not an artifact.Exactly one named source found in this bundle's research pass publishes the spike's deliverable as awritten document, Microsoft's Code with Engineering Playbook: "Generally the deliverable from aTechnical Spike should be a document detailing what was evaluated and the outcome of that evaluation."So this template does not tell you that writing up your spike is settled practice, because theevidence does not support that sentence. It gives you a good shape for the write-up if you are doingone. See spike-report_companion.md sections 1, 2 and 6.1 for the full argument and its sources.
THIS IS NOT AN ADR, AND THAT DRIFT IS THE EASIEST MISTAKE TO MAKE HERE.A spike report that recommends without recording what was tried is a bad ADR; an ADR that shows itsworking is not a spike report. This document investigates the question that precedes the decision: ithands over evidence and a proceed-or-not recommendation, and it does not propose, decide, or design.If the decision is genuinely already made and you are writing this to show your working, write an ADRand stop. See spike-report_companion.md sections 7 and 8.
WHERE THIS FILE GOES: wherever your team keeps them. This bundle prescribes no location, and it doesnot assume a file at all. One real published process (GitLab's) has none: the spike issue is theartifact, closed out with a single summary comment carrying the learnings and the recommended paths.If that is your culture, use the six sections below as the shape of that comment rather than fightingfor a file. See spike-report_companion.md section 9 (Adaptations).
STATUS is a short vocabulary, and there is no third state that means "gave up": in-progress -> complete the box closes whether or not the question was answered complete -> superseded once the decision this fed is recorded, the report is historyA spike whose box expired without an answer is still complete. Write the report anyway and make theRecommendation a named next spike with a narrower question.
FRONTMATTER: name a person and a way to reach them, not a team. A spike report generates questions,and a report nobody can be asked about decays into an assertion. The stable id and the date are whatlet a later design doc or ADR cite this by name instead of by memory. Seespike-report_companion.md section 3 (Anatomy > Frontmatter and status).
HOW TO FILL THIS IN1. Read the comment under each heading: WHAT it wants, WHY it matters (with a pointer into spike-report_companion.md for the deep reasoning), guiding questions to ASK, a GOOD and a WEAK example, and the TRAP to avoid.2. Replace each {{placeholder}} with your content.3. If a section does not apply, write "N/A" and one line of why, rather than deleting it.4. Before you ship it: self-grade against spike-report_guide.md, then DELETE every HTML comment. They are guidance, not content.-->
# {{title}}
<!-- WHAT A title naming the question the spike went after and, where you can, the answer it came back with. A statement, not a topic. WHY These are read later, by someone deciding whether your box already answered their question before they spend theirs. A title naming the area rather than the question makes them read the whole document to find out. Deep dive: spike-report_companion.md section 3 (Anatomy > Frontmatter and status). ASK What was the one uncertainty? What did the box come back with? Scanning a folder of these six months from now, would a stranger know whether to open this one? GOOD "Hosted OCR reads our supplier invoice headers well enough to drop manual entry, and does not read line items" WEAK "OCR investigation" (names the area, not the question, and carries no answer) TRAP Titling the technology instead of the question. "Vendor X evaluation" tells a reader nothing about what was asked or what came back, and ages into a filename nobody opens. -->
## The Question
<!-- WHAT The single uncertainty this spike existed to reduce, written so evidence can answer it yes or no. One question. Three questions is three spikes, or a research project. WHY The question is what makes the box meaningful: it is what the box was spent on, and it is what a reader checks the findings against. A spike aimed at a topic rather than a question ends with the terrain described and nothing decided, which is one of the named failure modes for this format. Deep dive: spike-report_companion.md section 3 (Anatomy > 1. The Question). ASK What one thing did we not know? Who is waiting on the answer and what will they do with it? Could the evidence have come back "no"? Was this written before the work started? GOOD "Can a hosted OCR service extract supplier, invoice number, date and line-item totals from our scanned supplier invoices accurately enough for finance to stop keying them by hand? The answer goes to AP operations, who will either schedule the integration this quarter or renew the data-entry contract for another year." WEAK "Investigate OCR options for invoice processing." (a topic and an activity, not a question; no evidence could make it false, so the box cannot close on an answer) TRAP Writing the question after the findings, shaped to the answer you got. State it as something the evidence is allowed to kill. A refuted hypothesis is a successful spike; a question that cannot be refuted is not a spike question, and the box will not save it. -->
{{the_question}}
## Scope and Time Box
<!-- WHAT Three things: what the investigation was allotted, what it actually spent, and what was deliberately not attempted. All three, including the branch you abandoned partway. WHY The box is what separates a spike from open exploration, and reporting it honestly in both directions is information: overrunning says one thing, finishing in half a day says another. The non-scope half sits here rather than under What Was Tried because what you chose not to try is a scope decision, and it reads best beside the budget that forced it. This bundle folds the time box into this section instead of giving it its own heading, and that is a departure argued from evidence rather than taste: a dedicated time-box heading turns out to live in pre-spike planning artifacts (a ticket field, a plan's deadline) and to be usually absent from the report of a completed spike. Deep dive: spike-report_companion.md section 3 (Anatomy > 2. Scope and Time Box), which carries the evidence for that departure. ASK What was the box, and agreed with whom? What did it actually cost? What did we choose not to look at, and why? What did we abandon partway, at what point, and on what signal? GOOD "Allotted: three days, agreed at sprint planning. Spent: two days. Vendor B's trial tier rate-limited us on day two and we cut its evaluation short rather than buy a larger key. Deliberately not attempted: handwritten annotations on delivery notes, which are a separate workflow and were excluded by agreement; and self-hosting an open-source model, which we stopped at hour four once the first run needed a GPU the AP environment does not have, so the shape was unusable before it got interesting." WEAK "Timeboxed to a sprint." (no actual spend and no scope boundary; the reader learns only that a box was mentioned once) TRAP Silently dropping the branch you abandoned. You gave up on it for a reason, and that reason is evidence. Leave it out and the next team spends their whole box rediscovering that it does not work. -->
* Allotted: {{time_allotted}}* Spent: {{time_spent}}* Deliberately not attempted: {{not_attempted}}
## What Was Tried
<!-- WHAT The approach: what you built, ran, measured or read, in enough detail that a colleague could repeat it and get the same result. Pin the versions, the environment and the data. WHY This is the section that separates a spike report from an opinion. Repeatability is what lets a reader disagree with your conclusion by checking your work rather than by trading intuitions, and the one named source that publishes this document makes repeatability an explicit property of it. Deep dive: spike-report_companion.md section 3 (Anatomy > 3. What Was Tried). ASK What exactly did we run, on what data, at which versions? What would a colleague need to reproduce it? What did we reuse from earlier work rather than re-run, and why is that still valid? GOOD "Sampled 200 invoices from last quarter, stratified across our five highest-volume suppliers. Ran each through vendor A's document API (v3, eu-west endpoint, default invoice model) and vendor B's (v2024.2, trial tier). Scored extraction field by field against the values already keyed into the ledger. Scoring script and sample manifest are in spikes/ocr/, commit 4f1c9ab. Vendor B's numbers cover 60 invoices only, for the rate-limit reason above." WEAK "Tried a couple of OCR services and compared the results." (nothing named, nothing versioned, no data described; a second person cannot repeat any of it) TRAP Reusing evidence from an earlier spike without saying so. Reusing it is legitimate; reusing it silently makes a stale measurement look fresh and hides the version it was actually taken on. -->
{{what_was_tried}}
## What Was Found
<!-- WHAT The evidence, with the facts you observed kept visibly apart from what you think they imply. Anchor every fact to something a reader can open: a count, a command, a file and line, a version. WHY That separation is this document's whole defence. An observation and an inference read identically in prose, and once they are blended a reader cannot tell which parts survive if your interpretation turns out to be wrong. The one named source that publishes this document argues the goal of a spike is fact-finding rather than decision-making or recommendation; this template keeps a Recommendation section anyway, and the separation here is what pays for that. Deep dive: spike-report_companion.md section 3 (Anatomy > 4. What Was Found). ASK What did we actually observe, in numbers or output? What does each observation imply, and how confident are we in that step? What surprised us? Did anything we wrote earlier turn out to be wrong, and did we mark it or quietly edit it away? GOOD "Observed: header fields (supplier, invoice number, date, total) correct on 197 of 200 invoices for vendor A; line-item totals correct on 138 of 200. All 62 line-item failures were invoices carrying handwritten corrections in the line-item block, or a scan skew above roughly 5 degrees. Implies: header automation is viable now for the three suppliers who send clean PDFs; line-item automation is not, and because the failure mode is input quality rather than the model, a different vendor is unlikely to move it much. Confidence: high on the header count, medium on the skew threshold, which we eyeballed rather than measured." WEAK "OCR worked well for most invoices but struggled with the messier ones, so it looks promising." (no counts, no anchor, and observation and inference are the same sentence) TRAP Deleting a finding that turned out to be wrong instead of marking it. Track the retraction in place, saying what changed and when, so a later reader can tell which conclusions moved and which never did. A quietly edited report cannot be trusted on the parts that held. -->
* Observed: {{observed_facts}}* Which implies: {{what_it_implies}}
## Recommendation
<!-- WHAT Proceed, do not proceed, or a named next spike with a narrower question. The verdict in the first line, in words, followed by the reasoning and any condition it depends on. WHY An investigator who spent the box and will not say what they now believe has pushed the hardest part of the work onto a reader with less context who was not there. Worth knowing as you write it: the single named source that publishes this document argues the opposite, that a spike should be fact-finding rather than recommendation. This template keeps the section deliberately, and it pays for it with the trap below. Deep dive: spike-report_companion.md section 6.2 (Fact-finding or recommendation?) for both sides of that, and section 3 (Anatomy > 5. Recommendation) for how to write one. ASK Proceed, stop, or spike again? Under what condition, and what would flip it? Who decides, and what do they need that is not in this document? Was the decision rule written down before the results came in? GOOD "Proceed, narrowly. Automate header extraction with vendor A for the three suppliers who send clean PDFs, about 60 percent of monthly volume, and keep manual entry for the rest. Conditional on the eu-west endpoint clearing the data-residency review already open with legal. If that comes back no, this becomes a do-not-proceed, not a smaller version of the same plan." WEAK "The results were encouraging and we think this is worth pursuing further." (no verdict a reader can act on, no condition, and "further" names nothing) TRAP Writing this section first. A verdict reached before the box opened turns the evidence sections into a preamble that justifies it, and that document is an ADR wearing a spike report's headings. A conditional recommendation can be more honest than a clean one; a recommendation invented after the evidence arrived is hard to tell from a preference. -->
**{{verdict}}** - {{recommendation_detail}}
## What This Does Not Settle
<!-- WHAT The open questions this spike did not close, and what a reader must not conclude from it. What you would still not bet on. Then stop: this is not a roadmap. WHY A spike buys a bounded answer, and the boundary is part of the answer. Without it the next reader takes a result that held on one dataset, one version and one platform as a general one. This section is a deliberate departure from the four-part shape (question, approach, findings, recommendation), argued from research rather than taste: real filled spike reports supply an explicit non-scope statement under four different headings across four unrelated projects, and not one blank template found asks for it. Three caveats travel with that finding and it should not be quoted without them. Deep dive: spike-report_companion.md section 3 (Anatomy > 6. What This Does Not Settle), and section 6.4 for exactly what that finding does and does not establish. ASK What would we still not bet on? What did we measure on one dataset, one version or one platform that may not generalise? What question is sharper now but still open? What might a reader wrongly conclude from this document? GOOD "Nothing here covers credit notes or multi-currency invoices; the sample contained none. The 197 of 200 header figure is measured on last quarter's suppliers, two of whom are mid-migration to a billing system that will change their layouts. Cost at production volume is unknown: we ran 260 documents on trial pricing and never reached a tiered rate. And we did not test what happens when extraction is confidently wrong rather than visibly failing, which is the risk finance actually cares about." WEAK "Next steps: evaluate pricing, test credit notes, book a follow-up with finance." (a to-do list, not a boundary; it says what you plan to do, not what your answer fails to cover) TRAP Turning this into a roadmap, or leaving it empty because the spike "answered the question". A report that gives its answer but not the edge of its answer is the failure this whole format exists to prevent, and an empty section here is worth a second look before you accept it as clean. -->
* Still open: {{open_questions}}* Do not conclude from this: {{what_this_does_not_show}}---title: "Over-the-air updates fit the overnight window, and today's bootloader cannot survive one that fails"spike_id: "SPK-2026-014"status: "complete"conducted_by: ["Hana Okafor, Firmware (hana.okafor@riverlandsensing.example)"]date: "2026-09-04"work_item: "OPS-1187, Can deployed catchment nodes be updated over the air?"doc_type: spike-reportsize: leansource_template: spike-reportsource_template_version: 0.1.0---
> **Worked example.** A filled `spike-report` for a four-day technical investigation at Riverland Sensing,> a fictional operator of water-quality monitoring nodes deployed in remote catchments. The bundle ships> one variant, so there is no size choice to demonstrate here. Every figure, version, serial number,> commit hash, identifier and date below is illustrative.>> It is deliberately independent of the [`adr`](../adr/adr_example.md), [`rfc`](../rfc/rfc_example.md) and> [`sdd`](../sdd/sdd_example.md) examples rather than chained to them. The `decision-docs` family teaches> the difference between investigating a question, proposing an answer, recording a decision and> describing a design, so each member's example is a different piece of work.>> Read it alongside [`spike-report_guide.md`](spike-report_guide.md), the rubric it was graded against.> Three things it does are the ones this template's own TRAP notes warn against skipping: it names the> branch the investigation abandoned and the signal that stopped it, it corrects an early finding in> place instead of deleting it, and it> spends its closing section on what its own answer fails to cover. Whether a spike should produce a> written document at all is contested;> [`spike-report_companion.md`](spike-report_companion.md) section 6.1 carries both sides of that, and> this example demonstrates the shape rather than a settled practice.
# Over-the-air updates fit the overnight window, and today's bootloader cannot survive one that fails
## The Question
Can a complete application firmware image be delivered to a deployed RS-4 catchment node over its existinglow-power radio link, inside one overnight maintenance window, without drawing the node's battery below thereserve it needs to survive until its next scheduled calibration visit?
One question, one node, one window. Field operations owns the site-visit budget for the 2027 season andholds a line in it for unscheduled trips, of which firmware faults are the largest single cause. They aredeciding in October whether the firmware share of that line can come out. Two results would have closedthis as a no: a node that cannot finish a transfer before the window ends, or a node that finishes andlands under reserve. Both were live possibilities on day one, and one of them happened.
## Scope and Time Box
* Allotted: four working days, agreed with the field operations lead and the firmware lead at the August planning session, under one standing condition: nothing gets transmitted to a node that is actually deployed. Bench only.* Allotted time is **calendar** time, not desk time, and that distinction is what makes the run count possible. The bench ran unattended around the clock: the sixteen scored runs, one node at a time, at four to eight hours each are roughly ninety hours of radio time, which fits three and a half calendar days only because nobody had to watch it. Attended work was the setup, the two stalls, and the analysis.* Spent: three and a half days. The last half-day was planned for a fourth link profile and was not used. Once the bootloader constraint in What Was Found turned up, a fourth set of transfer times could not have changed what this report recommends, so the box was closed early rather than spent on a number nobody would act on.* Deliberately not attempted: * **Gen-1 nodes**, the 60 units carrying the smaller flash part. Excluded by agreement at the start: they are already on the retirement schedule and will be off the network before any rollout could begin. Nothing here applies to them. * **Differential updates.** Planned for day three, abandoned at the end of day two. We had intended to measure a binary delta against the full image, and stopped after reading the bootloader source and linker map, which show a single application slot with nowhere to stage and verify a patched image before it overwrites the running one. The signal that ended this branch was the memory map, not a measurement, and it is the reason the recommendation below reads the way it does. * **More than one node updating behind a gateway at a time.** Never inside a four-day box. It is the largest item in What This Does Not Settle, and it is the one that decides how long a fleet rollout takes. * **The calibration firmware**, which is a separately signed image on its own update path and was out of scope by agreement.
## What Was Tried
Everything below ran on a bench rig, never on the deployed fleet.
**Hardware and versions.** Six RS-4 nodes, serials RS4-0031 through RS4-0036, all on application firmware3.8.2 and bootloader 1.4.0. One production-configuration gateway, GW-C, on build 2.11.4-rs. A programmableRF attenuator sat between the nodes and the gateway so that link conditions could be set on purpose insteadof waited for. Transfer used the fragmented-downlink mechanism already shipping in 3.8.2; no new transportwas written for this.
**Link profiles, taken from logs rather than invented.** Three profiles were built from twelve months ofgateway logs: a **strong** profile derived from the 40 deployed nodes with the best median link margin, a**fleet-median** profile, and a **weak** profile matched to the 20 worst-performing nodes on the network.Each profile is a fixed attenuation plus a scripted fade replayed from a recorded month.
**What was measured.** Wall-clock time to transfer the full 214 KB image and acknowledge every fragment,the fragment-acknowledgement percentage at the point a run was stopped, and charge drawn. Charge was takenwith a coulomb counter in series on three of the six nodes, RS4-0031, RS4-0033 and RS4-0035, sampled onceper second.
**Where the evidence lives.** Harness, attenuation profiles and raw per-run logs are in `spikes/ota-window/`at commit `9d41c2e`. Runs are numbered in `runs/manifest.csv`, and every figure in the next section namesthe runs it comes from. Nineteen runs are recorded: runs 1 and 19 are rig setup and teardown checks, run 2is excluded (see the correction below), and runs 3 to 18 are the sixteen scored runs.
**Reused rather than re-measured, and flagged as such.** The idle-current baseline each charge figure issubtracted from is the one captured during 3.8.2 release qualification in June, not re-taken this week. Thefirmware did not change between then and now, so the reuse is sound, but the baseline is three months oldand a reader should know that before treating the charge numbers as fresh.
## What Was Found
* Observed: * **Strong profile, runs 3 to 8.** Five of six transfers completed, in 4 h 10 m to 4 h 35 m, against a seven-hour overnight window. The sixth, run 7, ended in the downlink stall described two bullets below rather than in anything to do with the link. * **Fleet-median profile, runs 9 to 14.** Four of six completed, in 6 h 05 m to 6 h 50 m. Run 13 stood at 84 percent of fragments acknowledged when the window closed. Run 12 ended in the stall. * **Weak profile, runs 15 to 18.** No run completed. Three were deliberately allowed to run an hour past the seven-hour window, to see how close the weakest link came, and were stopped at the eight-hour mark with between 61 and 74 percent of fragments acknowledged; run 17 ended in the stall at 52 percent. * **Charge.** 41 to 46 mAh per completed strong-profile transfer; 58 mAh for the slowest median-profile run that finished. The illustrative annual energy budget for an RS-4 is 1,100 mAh per node. * **A repeated failure with no cause attached.** On runs 7, 12 and 17, the gateway stopped serving downlink to a node that was awake and still requesting fragments. The node kept asking; the queue never moved again. Service resumed only after the gateway process was restarted. Same symptom, three of sixteen scored runs, and once on each of the three link profiles. * **The bootloader has one application slot.** Read directly from bootloader 1.4.0 source and the linker map. There is no second region to write into and no known-good image to return to, so a write interrupted partway leaves a node holding an incomplete image. No node was deliberately interrupted to demonstrate this.* Which implies: * **The window binds where the link is worst, and only there.** For well-linked nodes the overnight window is comfortably long enough. Airtime grows with retries, retries are what the weak-profile nodes spend their time doing, and the nodes that are hardest to reach by radio are the same ones that are most expensive to reach by road. So the saving this work was chasing is concentrated in exactly the part of the fleet where it is worth least. * **Energy is not the reason to stop.** At 58 mAh, the worst completed transfer is roughly five percent of a node's annual budget, for an update we would run at most twice a year. Nothing measured here threatens the calibration-visit reserve. * **The downlink stall is an open defect, not a known limit.** It is not a link-quality effect, because it happened once on each of the three profiles, which rules out the explanation we would otherwise have reached for first. We can describe the symptom precisely and cannot say what causes it: gateway configuration, a node-side request pattern, and a resource ceiling on GW-C are all still live. Any schedule built on the transfer times above is built on a path that stopped working three times this week. * **The single slot is the finding that changes the answer.** With no fallback image, a transfer that fails at the wrong moment does not cost a retry. It costs a technician standing in a catchment with a programmer, which is the cost this whole idea exists to remove. This is an inference from code rather than from an observed failure, and the confidence is high for that reason rather than in spite of it: the recovery path is absent from the source, not merely unreliable. * **Confidence, stated where it is uneven.** High on the transfer times, which come from sixteen scored runs with one variable moved at a time. Lower on the charge figures, which rest on three nodes from a single cell batch against a three-month-old baseline. Lowest on anything to do with the stall, where we have a symptom and nothing else.
**Correction, kept in place.** The day-one note in `runs/README.md` recorded that a weak-profile transfer"completes in about nine hours". That was wrong and it is corrected rather than deleted. It came from run 2,where the attenuator script had not yet been applied, so the node was effectively running on the strongprofile. Run 2 is excluded from every figure above and is marked excluded in the manifest. The correctstatement is that no weak-profile run completed at all.
## Recommendation
**Proceed to a bounded field pilot, and only after a bootloader that can fall back has shipped.** - Run thepilot on the 40 strong-profile nodes behind GW-C and GW-F, for one season, before anyone plans anythingwider. The precondition is not a preference and it is not a nice-to-have sequencing note: on bootloader1.4.0 a failed transfer strands a node until a person reaches it, so piloting first would risk the exactexpense the pilot is meant to prove we can avoid. If a recoverable bootloader is not on the firmwareroadmap for this year, then the honest answer for the 2027 budget is no, and the unscheduled-visit lineshould be renewed in October at its current size.
Two things this recommendation deliberately does not ask for. It does not ask for a differential updatescheme: once the window fits, nothing measured suggests the image needs to be smaller. And it does not askfor a gateway change, because the stall has a symptom and no diagnosis, and buying a fix for anunexplained fault is how teams end up with two problems.
Who decides what: the firmware lead owns the bootloader call, field operations owns the budget line. Bothneed the stall triaged before either can commit, and that triage is the next spike if the bootloader workgets scheduled. One day, one question, bench rig already standing: what makes the downlink queue stopserving a node that is still awake and asking?
## What This Does Not Settle
* Still open: * **Concurrency, which is the big one.** Every number in this report is one node at a time. There are 380 nodes behind nine gateways. Whether a gateway can carry several transfers in a single window is the difference between a rollout that takes one season and one that takes three, and this spike does not even bound it. * **The cause of the stall on runs 7, 12 and 17.** Named as a symptom, unexplained as a mechanism. Until someone can say what triggers it, nobody can say how often it would happen across a fleet, and the pilot's schedule inherits that ignorance. * **Cold.** Charge was measured on three nodes from one cell batch at bench temperature, roughly 19 to 22 C. Several catchment sites sit below freezing for weeks at a time, cell behaviour there is not something a bench at room temperature can tell you, and the coldest sites are the ones already running the thinnest energy margin. * **Whether last year predicts next winter.** The weak profile is a replay of a recorded month from the past twelve. It reconstructs conditions that happened; it does not forecast conditions that have not. * **Whether a node can abandon a transfer cleanly.** We never tried commanding one to stop partway, so a transfer that is visibly not going to finish currently just runs until the window ends. If the stall turns out to be node-side, this matters more than it currently sounds.* Do not conclude from this: * **That over-the-air updating removes the site visit.** It removes one reason for one. Annual calibration still puts a person next to every node, and the budget line this question came from covers both reasons, so what is on offer is a share of that line rather than the line. * **That 214 KB is a fixed quantity.** It is the size of today's release. The slowest median-profile run that finished had ten minutes of window left, so a larger image does not degrade these results gradually; it moves nodes across the line from completing to not completing. * **That the strong-profile results describe the fleet.** They describe 40 nodes out of 380. What matters for a rollout plan is the distribution of link margin across the whole network, and this spike sampled three points on that distribution rather than measuring it. * **That the energy result holds for a fleet doing this routinely.** It was measured for one successful update. Nothing here covers a node that retries a failing update night after night because no alert fires when a transfer never completes.Provenance
Section titled “Provenance”The reasoning, the history and every source, in the repository:
- Companion - the long-form argument: why these sections, where the sources disagree, and what the bundle refuses to claim
- History - what changed in this bundle, and when
- Research log - every source consulted, with what each one actually supports
- Catalog metadata - the machine-readable record this page is generated from
Catalog record: 7 sections across 1 format(s).