D4 build week: the table turns green
420-302-VA · WEEK 14 · FALL 2026

Stage 4 of 6 · The truth · one artifact, three jobs

The scoreboard

The verification table has been in the repo since the project opened, mostly red. This is the week it earns its keep, doing three jobs at once: compass for the daily loop (the reddest high-value row is the next rung), honest record of where the system stands, and, come Monday, most of the demonstration itself, since D4's brief makes each row demonstrable on request. A table updated the moment evidence exists costs seconds; a table reconstructed Sunday night is fiction with dates.

Why the week needs a scoreboard

Feelings about progress are exactly the instrument the ninety-ninety rule defeats: a week of "basically done" rungs feels productive right up until the joins come due. The table is the countermeasure because it only accepts the currency that cannot be faked, evidence against requirements. This is also where Week 5's discipline pays its dividend: requirements written testable and unambiguous back then are rows that can turn green with a stopwatch now, which was the entire point of writing them that way. A row that cannot be tested cannot be verified, and a row that cannot be verified cannot be demonstrated on request; if you find one this week, the rung is to rewrite the requirement first, Week 5 style, then prove it.

The traceability chain

Every green cell makes a four-link claim, and the links are checkable in both directions. Backward: this green exists because a requirement demanded it. Forward: this green is backed by evidence anyone can re-run. The chain is what "demonstrable on request" means mechanically:

Four linked cards left to right. One, the requirement: R3, dashboard shows PV at most three seconds old, written testable in Week 5. Two, the rung: staleness guard, tagged and on main. Three, the evidence: stopwatch run 2.4 seconds, journal entry December 3, re-run command in docs slash verification. Four, the table cell: R3, green check, December 3. Arrows connect them forward; a reverse dashed arrow over the top reads: on request, walk it backward and re-run it live. RequirementRungEvidenceTable cell R3: PV on page ≤ 3 s old testable, unambiguous: Week 5 paying off staleness guard on main, joined, old rungs re-run stopwatch: 2.4 s journal · Dec 3 re-run cmd in docs/ R3 ✓ Dec 3 "show me R3": walk it backward, then re-run it live A green any link of which is missing is not evidence yet; it is a claim wearing green.
Traceability: the chain behind every green. This is the standard engineering idea of requirements traceability scaled to a bench: nothing proven that was not required, nothing claimed that cannot be re-shown. Monday's "show me R3" is this figure performed, which is why filling it live each day is also demo rehearsal.

Green, amber, red: what each claims

CellIt claimsIt requiresOn request, you can
Green ✓ + dateRequirement met, verified on that date, still true on current mainThe full chain: rung joined, evidence recorded, re-run command writtenRe-run it live inside a minute or two
Amber ◐ + notePartially met, honestly scoped: "holds for one node", "manual restart needed"The note naming exactly what holds and what does notDemonstrate the part that holds, state the gap unprompted
Red ✗ (or "cut + date")Not attempted, or deliberately descoped at triage with the reason journaledNothing, which is its honesty: red hides nothingSay why it ranked where it did; a named cut is method, not failure

The one forbidden state

A green without its chain. One optimistic cell discovered under a marker's "show me" does not cost one row; it re-prices every green in the table at the worst moment, the exact mechanism the narrated-demo risk row describes. Ambers are cheap all week; a false green on Monday is the most expensive cell in the course.

The table across the week

Three snapshots of an eight-row table across the build week. Monday: all rows red, the list ranked beside it; one green from D3. Thursday, triage day: four greens with dates, one amber, three reds, caption majors landing, cut decided from the bottom. Sunday: six greens, one amber with its note, one row marked cut with its date, caption green with honest edges, nothing pending but the demo. Mon · rankedThu · triageSun · demo-ready R1 ✓ (D3)R2 ✗R3 ✗R4 ✗R5 ✗R6 ✗R7 ✗R8 ✗ R1 ✓ (D3)R2 ✓ Dec 1R3 ✓ Dec 3R4 ✓ Dec 3R5 ◐ one nodeR6 ✗R7 ✗R8 ✗ R1 ✓ (D3)R2 ✓ Dec 1R3 ✓ Dec 3R4 ✓ Dec 3R5 ✓ Dec 5R6 ✓ Dec 6R7 ◐ notedR8 · cut Dec 4 Dates accumulate as they happen. A table whose greens all share one date was written, not earned, and markers know it.
The honest trajectory. Monday red is correct, Thursday's mix is correct, and Sunday's green-with-edges is what strong looks like: one amber with its note, one cut with its date, nothing pretending. The daily loop's station 5 is what draws this figure one cell at a time.

Checklist for this stage

Check yourself

A requirement reads "the dashboard should be responsive". Walk it through the chain and show where it jams, then repair it.
It jams at the first link: "responsive" is neither testable nor unambiguous (responsive to what, within how long, measured how?), so no rung can target it, no evidence can satisfy it, and no cell can honestly turn green; any green on it is opinion. Repair with Week 5's qualities: "the dashboard reflects a PV change within 3 s" (testable with a stopwatch and a shaded sensor) or "the setpoint control acknowledges within 1 s of the click". Then the chain runs: the staleness rung targets it, the stopwatch run is evidence, the cell gets a date. The general lesson: when verification jams, the fault is as often in the requirement as in the system, and rewriting the row is a legitimate rung this week.
Why does the brief say "demonstrable on request" rather than just "verified"? What extra work does that one phrase impose, and where in this hub is it paid?
"Verified" is a claim about the past: it passed once, for you. "Demonstrable on request" is a claim about the present that a stranger can audit: the marker picks any row, and the chain must run live, backward to the requirement and forward through a re-run. The extra work is exactly the chain's third link kept executable: a written re-run command per row (docs/verification), evidence that survives re-execution, and a system still passing on current main, which is what the daily regression re-runs protect. It is paid at station 4 and 5 of the loop and in this page's checklist; and it is why the table doubles as the demo script, the next page's starting point.
Sunday night: R7 is amber, "alerts work but a restart loses the threshold setting". Your partner wants to colour it green since "the demo won't restart anything". Adjudicate, with mechanisms.
Amber, and the note stays. Two mechanisms say so. First, the false-green repricing: if any "show me" or mishap exposes the gap, every other green in your table loses credibility at once, trading one cell's optics for the whole artifact's trustworthiness. Second, the demo's own conditions make the gap live, not hypothetical: Week 12's services restart on failure by design (Restart=on-failure), so a mid-demo crash would restart the alert service and silently drop the threshold, turning a cosmetic amber into a live surprise. The honest play is strong anyway: demonstrate alerts working, state the restart caveat unprompted, and show the one-line journal DEBT entry planning the fix. That performance scores as method; the painted green risks everything to avoid saying one true sentence.