Stage 4 of 6 · Theory · finding the fault
Debugging the whole
A dead dashboard tells you one bit: something, somewhere, is wrong. Your system has eight segments where that something can hide, and the difference between a ten-minute fix and a lost evening is rarely cleverness; it is search strategy. Electronics technicians named the strategy half-splitting decades ago, and computer science knows it as binary search. Today it becomes your reflex.
Eight segments, one symptom
Lay your project flat and it is a chain: sensor hardware, node code, Wi-Fi, broker, rule program, bridge, Flask route, browser. Every symptom you will ever see, a frozen number, a silent actuator, a blank page, is the visible end of a fault in exactly one segment (usually), and the segments are ordered: data flows one way, so a fault at segment k poisons everything after it and nothing before it. That monotone structure is what makes smart search possible. "It broke somewhere" over eight ordered segments is a question with a known optimal algorithm.
Half-splitting: log2 beats guessing
Probe the middle of the chain. If the data is healthy there, the fault is downstream and four segments are acquitted in one observation; if unhealthy, it is upstream and the other four walk. Repeat on the surviving half. Each probe halves the suspects, so locating one faulty segment among n takes at most ⌈log2 n⌉ probes: for your n = 8 chain, 3 probes, every time, regardless of luck. Checking segments one by one from either end averages about n/2 = 4 and can cost 7; checking them in the order you suspect them costs whatever your hunch costs, which at 11 pm is a lot. The gap widens with system size, which is why the habit is worth installing on a system still small enough to forgive you.
Why the maths holds here
Half-splitting assumes the chain is ordered and the fault is single, both usually true on this bench because the incremental habit keeps faults from accumulating. When two faults coexist, the method still works, it just finds them one at a time: fix the upstream one, re-run, split again. What breaks the method is probing with an instrument you have not trusted yet, which is the next section's subject.
The probe at every point
A probe is only as good as its instrument, and the course has quietly equipped every point of the chain already. This table is eleven weeks of troubleshoot pages compressed into one column each:
| Probe point | Instrument | Healthy looks like | Taught in |
|---|---|---|---|
| Sensor, electrically | The node's REPL: read the raw value directly | Numbers that move when the world moves | Week 8 |
| Node code | print in the loop, watched over the REPL | Reading, converting, publishing lines in rhythm | Week 8 |
| Wi-Fi join | The bounded connect_wifi wait's own report | An IP address, once, at boot | Week 9 |
| Broker, everything it carries | The firehose: mosquitto_sub -t "#" -v | Your topics, your payloads, at your period | Week 9 |
| Rule program | Its own log lines; journalctl -u once it is a service | Parsed values and decisions, stated | Week 10, today |
| Bridge and route | curl http://localhost:5000/api/... | Fresh JSON, right keys, not null | Week 11 |
| Browser side | The console and the Network tab, F12 | Requests firing, 200s, no red | Week 11 |
One discipline binds the table: trust a probe only after it has shown you a healthy system. Run the firehose and the curl once while everything works, today, and save what healthy looked like in your journal; a baseline turns every future probe from "is this normal?" into a comparison.
The debugging loop
Half-splitting locates; the loop around it fixes without collateral damage. Four steps, in order, every time:
Two companions make the loop faster. The first is the journal the course has kept since Week 1: a line per fault, symptom, segment, cause, costs a minute and builds the personal troubleshoot table that D4's pace will depend on. The second is older than software: explain the fault aloud to your partner, or failing that to any object on the bench, before changing code. Hunt and Thomas canonized it as rubber-duck debugging in The Pragmatic Programmer (1999), and it works because articulating forces the assumptions you have been skipping into the open, where one of them is usually the bug.
Checklist for this stage
Check yourself
The dashboard shows a number that never changes. Give the first probe, both possible readings, and what each acquits.
mosquitto_sub -t "#" -v, the middle of the chain. If messages flow with changing values, everything from sensor to broker is acquitted and the fault lives downstream: rule, bridge (Week 11's stale latest or null bridge), route, or the browser's cache. If the firehose is silent or frozen too, the downstream four are acquitted and the fault is upstream: sensor, node code, Wi-Fi, or the broker itself. One observation, four suspects either way.