Project studio: the system comes together
420-302-VA · WEEK 12 · FALL 2026

Stage 4 of 6 · Theory · finding the fault

Debugging the whole

A dead dashboard tells you one bit: something, somewhere, is wrong. Your system has eight segments where that something can hide, and the difference between a ten-minute fix and a lost evening is rarely cleverness; it is search strategy. Electronics technicians named the strategy half-splitting decades ago, and computer science knows it as binary search. Today it becomes your reflex.

Eight segments, one symptom

Lay your project flat and it is a chain: sensor hardware, node code, Wi-Fi, broker, rule program, bridge, Flask route, browser. Every symptom you will ever see, a frozen number, a silent actuator, a blank page, is the visible end of a fault in exactly one segment (usually), and the segments are ordered: data flows one way, so a fault at segment k poisons everything after it and nothing before it. That monotone structure is what makes smart search possible. "It broke somewhere" over eight ordered segments is a question with a known optimal algorithm.

Half-splitting: log2 beats guessing

Probe the middle of the chain. If the data is healthy there, the fault is downstream and four segments are acquitted in one observation; if unhealthy, it is upstream and the other four walk. Repeat on the surviving half. Each probe halves the suspects, so locating one faulty segment among n takes at most ⌈log2 n⌉ probes: for your n = 8 chain, 3 probes, every time, regardless of luck. Checking segments one by one from either end averages about n/2 = 4 and can cost 7; checking them in the order you suspect them costs whatever your hunch costs, which at 11 pm is a lot. The gap widens with system size, which is why the habit is worth installing on a system still small enough to forgive you.

The eight segment chain from sensor to browser drawn as boxes, with the fault hidden in the Wi-Fi segment. Probe one lands mid chain at the broker, using the firehose, and reads silent, acquitting the four downstream segments, shown faded. Probe two lands at the node's log, reads healthy, acquitting sensor and node code. Probe three lands on the Wi-Fi join and finds it, with the caption three probes, guaranteed, log two of eight. sensor node code Wi-Fi broker rule bridge route browser 1 firehose at the broker: silent → fault is upstream 2 node log: readings healthy → sensor and code acquitted 3 Wi-Fi join: found it faded = acquitted by probe 1, four segments in one observation ⌈log2 8⌉ = 3 probes, guaranteed · one-by-one averages 4 and can cost 7
Binary search on a bench. The middle probe acquits half the chain whatever it finds; that indifference to luck is the method's power. The worked example is Week 9's oldest fault, the node that never joined Wi-Fi, found in three looks.

Why the maths holds here

Half-splitting assumes the chain is ordered and the fault is single, both usually true on this bench because the incremental habit keeps faults from accumulating. When two faults coexist, the method still works, it just finds them one at a time: fix the upstream one, re-run, split again. What breaks the method is probing with an instrument you have not trusted yet, which is the next section's subject.

The probe at every point

A probe is only as good as its instrument, and the course has quietly equipped every point of the chain already. This table is eleven weeks of troubleshoot pages compressed into one column each:

Probe pointInstrumentHealthy looks likeTaught in
Sensor, electricallyThe node's REPL: read the raw value directlyNumbers that move when the world movesWeek 8
Node codeprint in the loop, watched over the REPLReading, converting, publishing lines in rhythmWeek 8
Wi-Fi joinThe bounded connect_wifi wait's own reportAn IP address, once, at bootWeek 9
Broker, everything it carriesThe firehose: mosquitto_sub -t "#" -vYour topics, your payloads, at your periodWeek 9
Rule programIts own log lines; journalctl -u once it is a serviceParsed values and decisions, statedWeek 10, today
Bridge and routecurl http://localhost:5000/api/...Fresh JSON, right keys, not nullWeek 11
Browser sideThe console and the Network tab, F12Requests firing, 200s, no redWeek 11

One discipline binds the table: trust a probe only after it has shown you a healthy system. Run the firehose and the curl once while everything works, today, and save what healthy looked like in your journal; a baseline turns every future probe from "is this normal?" into a comparison.

The debugging loop

Half-splitting locates; the loop around it fixes without collateral damage. Four steps, in order, every time:

Four boxes in a cycle: reproduce, make it fail on demand; isolate, half-split to one segment; fix, the smallest change that addresses the cause; verify, the end to end check plus the old rungs. An arrow returns from verify to reproduce labelled still failing, split again, and a green exit arrow from verify is labelled green: commit, name the cause. Reproduce Isolate Fix Verify fail on demand half-split to one segment smallest change, the cause end to end + old rungs still failing: split again, with what the probe taught you green: commit, name the cause
Reproduce before you touch anything. A fault you can trigger on demand can be measured; one that comes and goes can only be feared. Verify re-runs the old rungs too, because the second-worst bug is the regression your fix just planted.

Two companions make the loop faster. The first is the journal the course has kept since Week 1: a line per fault, symptom, segment, cause, costs a minute and builds the personal troubleshoot table that D4's pace will depend on. The second is older than software: explain the fault aloud to your partner, or failing that to any object on the bench, before changing code. Hunt and Thomas canonized it as rubber-duck debugging in The Pragmatic Programmer (1999), and it works because articulating forces the assumptions you have been skipping into the open, where one of them is usually the bug.

Checklist for this stage

Check yourself

The dashboard shows a number that never changes. Give the first probe, both possible readings, and what each acquits.
First probe: the firehose on the broker, mosquitto_sub -t "#" -v, the middle of the chain. If messages flow with changing values, everything from sensor to broker is acquitted and the fault lives downstream: rule, bridge (Week 11's stale latest or null bridge), route, or the browser's cache. If the firehose is silent or frozen too, the downstream four are acquitted and the fault is upstream: sensor, node code, Wi-Fi, or the broker itself. One observation, four suspects either way.
For a chain of 32 segments, compare worst-case probes for half-splitting versus one-by-one, and say what property of the chain the advantage depends on.
Half-splitting: ⌈log2 32⌉ = 5 probes worst case. One-by-one: up to 31, averaging about 16. The advantage depends on the chain being ordered with one-way data flow, so that a single healthy or faulty observation at a point classifies every segment on one side of it. In an unordered tangle, a mid-point probe acquits only itself, and the log2 collapses back toward linear, which is one more argument for pipeline-shaped architectures like yours.
Why does the loop demand reproduction before isolation, when the fault seems obvious?
Because without a trigger you cannot run the verify step: a fault that appears on its own schedule also disappears on its own schedule, so any fix "works" until it does not, and you cannot distinguish repaired from dormant. Reproduction also sharpens isolation itself, since the triggering condition (only after reboot, only when both laptops join, only past 60 readings) is evidence about the segment, often better evidence than the symptom. If you truly cannot reproduce it, say so in the journal and instrument the suspect segments with logging; the next occurrence then arrives with data attached.