Stage 4 of 6 · The truth · one artifact, three jobs
The scoreboard
The verification table has been in the repo since the project opened, mostly red. This is the week it earns its keep, doing three jobs at once: compass for the daily loop (the reddest high-value row is the next rung), honest record of where the system stands, and, come Monday, most of the demonstration itself, since D4's brief makes each row demonstrable on request. A table updated the moment evidence exists costs seconds; a table reconstructed Sunday night is fiction with dates.
Why the week needs a scoreboard
Feelings about progress are exactly the instrument the ninety-ninety rule defeats: a week of "basically done" rungs feels productive right up until the joins come due. The table is the countermeasure because it only accepts the currency that cannot be faked, evidence against requirements. This is also where Week 5's discipline pays its dividend: requirements written testable and unambiguous back then are rows that can turn green with a stopwatch now, which was the entire point of writing them that way. A row that cannot be tested cannot be verified, and a row that cannot be verified cannot be demonstrated on request; if you find one this week, the rung is to rewrite the requirement first, Week 5 style, then prove it.
The traceability chain
Every green cell makes a four-link claim, and the links are checkable in both directions. Backward: this green exists because a requirement demanded it. Forward: this green is backed by evidence anyone can re-run. The chain is what "demonstrable on request" means mechanically:
Traceability: the chain behind every green. This is the standard engineering idea of requirements traceability scaled to a bench: nothing proven that was not required, nothing claimed that cannot be re-shown. Monday's "show me R3" is this figure performed, which is why filling it live each day is also demo rehearsal.
Green, amber, red: what each claims
Cell
It claims
It requires
On request, you can
Green ✓ + date
Requirement met, verified on that date, still true on current main
The full chain: rung joined, evidence recorded, re-run command written
Re-run it live inside a minute or two
Amber ◐ + note
Partially met, honestly scoped: "holds for one node", "manual restart needed"
The note naming exactly what holds and what does not
Demonstrate the part that holds, state the gap unprompted
Red ✗ (or "cut + date")
Not attempted, or deliberately descoped at triage with the reason journaled
Nothing, which is its honesty: red hides nothing
Say why it ranked where it did; a named cut is method, not failure
The one forbidden state
A green without its chain. One optimistic cell discovered under a marker's "show me" does not cost one row; it re-prices every green in the table at the worst moment, the exact mechanism the narrated-demo risk row describes. Ambers are cheap all week; a false green on Monday is the most expensive cell in the course.
The table across the week
The honest trajectory. Monday red is correct, Thursday's mix is correct, and Sunday's green-with-edges is what strong looks like: one amber with its note, one cut with its date, nothing pretending. The daily loop's station 5 is what draws this figure one cell at a time.
Checklist for this stage
Check yourself
A requirement reads "the dashboard should be responsive". Walk it through the chain and show where it jams, then repair it.
It jams at the first link: "responsive" is neither testable nor unambiguous (responsive to what, within how long, measured how?), so no rung can target it, no evidence can satisfy it, and no cell can honestly turn green; any green on it is opinion. Repair with Week 5's qualities: "the dashboard reflects a PV change within 3 s" (testable with a stopwatch and a shaded sensor) or "the setpoint control acknowledges within 1 s of the click". Then the chain runs: the staleness rung targets it, the stopwatch run is evidence, the cell gets a date. The general lesson: when verification jams, the fault is as often in the requirement as in the system, and rewriting the row is a legitimate rung this week.
Why does the brief say "demonstrable on request" rather than just "verified"? What extra work does that one phrase impose, and where in this hub is it paid?
"Verified" is a claim about the past: it passed once, for you. "Demonstrable on request" is a claim about the present that a stranger can audit: the marker picks any row, and the chain must run live, backward to the requirement and forward through a re-run. The extra work is exactly the chain's third link kept executable: a written re-run command per row (docs/verification), evidence that survives re-execution, and a system still passing on current main, which is what the daily regression re-runs protect. It is paid at station 4 and 5 of the loop and in this page's checklist; and it is why the table doubles as the demo script, the next page's starting point.
Sunday night: R7 is amber, "alerts work but a restart loses the threshold setting". Your partner wants to colour it green since "the demo won't restart anything". Adjudicate, with mechanisms.
Amber, and the note stays. Two mechanisms say so. First, the false-green repricing: if any "show me" or mishap exposes the gap, every other green in your table loses credibility at once, trading one cell's optics for the whole artifact's trustworthiness. Second, the demo's own conditions make the gap live, not hypothetical: Week 12's services restart on failure by design (Restart=on-failure), so a mid-demo crash would restart the alert service and silently drop the threshold, turning a cosmetic amber into a live surprise. The honest play is strong anyway: demonstrate alerts working, state the restart caveat unprompted, and show the one-line journal DEBT entry planning the fix. That performance scores as method; the painted green risks everything to avoid saying one true sentence.