Stage 2 of 6 · The list · theory behind Monday night's ten minutes
Rank the rungs
The ranked list exists; you made it leaving D3. This page is the theory that makes it trustworthy: what the ordering optimizes, why D4's marking makes breadth the right axis, and the two empirical laws about estimation that guarantee the list will need surgery midweek. A list you re-rank on Wednesday without guilt is a plan; a list you obey after it stops being true is a liability.
Why ordering is the whole game
A build week is a fixed budget of hours spent against an unbounded wish list, and under a fixed budget the order of work determines the outcome almost by itself. The reasoning is blunt: whatever the week actually delivers will be some prefix of the list, because the week ends mid-list wherever it ends. Put the wrong things first and the week delivers the wrong prefix, worked just as hard. This is why the course keeps returning to the ranked list since Week 11: not as paperwork, but because ordering is the one decision that controls every hour downstream of it.
Value density and the quadrants
The sort key is value density: v = value ÷ effort, the milestone marks a rung protects per hour it costs. Two rungs of equal value are not equal if one takes an evening and the other takes three; the evening one buys the same green sooner and leaves hours for the next. Estimating both inputs is rough, and rough is enough, because the list only needs the order right, not the numbers. The classic picture:
Sort by v = value ÷ effort and the quadrants fall out. "Value" has a definition this week, the three faces of D4: does the rung land a requirement, strengthen the evidence, or improve replicability? A rung that does none of those has a density of zero at any effort, which is what the bottom-right is for.
What counts as value is written down
Resist inventing a private definition of value. D4's evaluation is weighted toward working breadth, verified requirements and a replicable repo, so those are the axes, and the ranked list should be defensible out loud: "this rung is third because it turns R4 green and nothing cheaper does." Documentation rungs score here too; "replicable by a stranger" is a face of the milestone, so an hour on the README can outrank an hour of code.
Breadth beats depth, and why
The deepest ordering rule follows from diminishing returns: the first rung that makes a requirement work at all is worth more than the fifth rung polishing a requirement that already works. Value concentrates in firsts, a pattern the Pareto principle names in general: a small fraction of the work carries most of the value, and for D4 that fraction is "each requirement's first proof". The marking agrees on purpose, weighted toward breadth across the table rather than shine on one row:
Depth feels like craftsmanship and marks like absence. The left pair worked just as hard; the table does not record effort, it records proof. The second coat of paint on R1 was always worth less than the first coat on R4, and this week the difference is 15 % of a course.
Estimates lie, on schedule
Two empirical laws, both older than you and both verified every term on this project, say your effort estimates are wrong in a known direction. Hofstadter's law (from Gödel, Escher, Bach, 1979): it always takes longer than you expect, even when you take into account Hofstadter's law; the recursion is the point, because knowing about optimism bias does not switch it off. And the ninety-ninety rule (Tom Cargill, Bell Labs): the first 90 % of the work takes 90 % of the time, and the remaining 10 % takes the other 90 %, because the last slice is where integration, edge cases and "done done" live, exactly the part a happy-path demo skips. The practical consequences for this week are arithmetic, not attitude:
Plan the top half. If estimates err by a factor near two, a six-day week honestly funds about half the rungs you believe it does. Rank so that the half that survives is the half that matters.
Count a rung done only when it is "done done": joined, old rungs re-run, evidence recorded. "Basically working" is the ninety-ninety rule's favourite meal; the definition of done on the next page is the fence around it.
Front-load the scary join. The rung whose integration you fear belongs early, while there is week left to absorb its second 90 %. Week 12's integration-order argument, now applied to the calendar.
Re-ranking without shame
Given the laws above, the Wednesday where reality and the list disagree is scheduled, not hypothetical, which is why the triage checkpoint exists. The move is always the same and always in the same direction: cut scope from the bottom of the list, never quality from the rungs you keep. A shorter list of done-done rungs beats a full list of almosts by the ninety-ninety rule alone, and a named cut ("R6 descoped Wednesday, here is why") is evidence of method, the thing D4's evaluation explicitly reads the repo for. The scope cliff row in the risk table describes the pairs who cut quality instead; midweek is exactly when that cliff recruits.
Checklist for this stage
Check yourself
Rung A turns R5 green and takes an estimated 2 h; rung B adds a second disturbance demo to already-green R2 and takes 1 h. Rank them, with the arithmetic and the principle.
By raw density B looks strong: cheap hour, visible payoff. But value is defined by D4's faces, and R2 is already proven; a second proof adds little where the first added everything, the diminishing-returns point. A's value is a whole requirement moving from unproven to proven, worth far more than B's garnish even at twice the cost, so A ranks first: v_A = (a requirement's first proof)/2 h beats v_B = (a flourish)/1 h. B may still happen later as a fill-in. The principle: firsts before seconds, breadth before shine, because the table records proof, not polish.
Your pair estimated eight rungs for the week. Using Hofstadter's law and the ninety-ninety rule, what should the plan actually assume, and what follows for the list?
Assume roughly four land at done-done, because estimates on complex, integration-heavy work run optimistic by about a factor of two even when you correct for it (that correction failing is Hofstadter's recursion), and the ninety-ninety rule says the late, invisible slice of each rung, joining, edge cases, evidence, costs about as much as the visible slice did. Consequences: the top four rungs must be the four that matter most (ranking is load-bearing precisely because the week ends mid-list); "done" must mean the full definition of done so progress is not an illusion; and Wednesday's triage is in the plan from Monday, since the laws predict the gap, they just cannot say which rung opens it.
A pair argues: "rewriting the bridge with async would make it so much cleaner, and we know how." Place the rung on the quadrant and make the case, including the one condition that would move it.
Bottom-right, the money pit: real effort (a rewrite of a working component plus re-verifying everything that touches it, which by n(n−1)/2 is most of the system) against no face of D4, since the bridge already works, no requirement turns green, no evidence strengthens, no stranger rebuilds easier. Worse, a rewrite resets a proven component to unproven in the exact week proof is the product. The one condition that moves it: if the bridge's current form blocks a high-value rung (it demonstrably cannot carry a needed feature), then the smallest rewrite that unblocks that rung becomes part of that rung's effort and is ranked as such. Otherwise it is journaled as technical debt, the next page's discipline, and left for January.