Project studio: the system comes together
420-302-VA · WEEK 12 · FALL 2026

Utility · keep open during the studio

Troubleshoot

Integration week has its own failure signature: components that each work alone and a system that does not. The ladder below orders the search the way the whole hub taught, pieces, then pairs, then the whole, then time, because a fault's level decides its instrument, and climbing in order is half-splitting by another name.

Debug in layers

A four rung ladder. Rung A, pieces: each component alone, REPL, print lines, curl against a stub; prove each piece against an instrument, not a teammate. Rung B, pairs: one join at a time in the bench order, firehose between every pair. Rung C, whole: the end to end check, cause, travel, name; half-split any failure. Rung D, time: reboots, restarts, hours of running; journalctl, the crash loop, the regression. A · Pieces B · Pairs C · Whole D · Time Each component alone:REPL, print lines,curl against a stub.Prove it with aninstrument, never"my partner's programseemed happy". One join at a time,in the bench order:node↔broker first.The firehose sitsbetween every pair;payloads read exactlyas docs/topics.md says. The end-to-end check:cause, travel, name.Any failure here ishalf-split, middle first.Three probes locate it;the loop fixes it withoutplanting a regression. Reboots, restarts,hours of running.journalctl reads thenights; the crash loophides in uptime, theregression in old rungs.Re-run them weekly.
Pieces, pairs, whole, time. A symptom at one level with all lower levels unproven is not yet a diagnosis; the studio queue moves faster for pairs who arrive saying which rung they are stuck on.

Symptoms and fixes

RungSymptomLikely cause and fix
AComponent passes alone, fails the moment it is joinedAlmost always a contract mismatch, not a component bug: diff what each side actually sends and expects (firehose, curl -v) against docs/topics.md; the file, or one side, is wrong
APartner's piece "works on their laptop", fails on the benchEnvironment, not code: missing package in the bench venv, hard-coded path or address, or a config value living outside config_example.py's pattern; reproduce on the bench with their stub commands
BFirehose shows the topic twice with different spellingsTwo programs disagree with the contract (Light vs light, s07 vs s12): fix the deviant side to match docs/topics.md, then clear any retained ghost with mosquitto_pub -r -n
BRule sees messages but crashes parsing some of themA stub or an old publisher is sending a second payload format; the firehose's -v shows which; kill the stray publisher and make the rule log, not die, on a malformed payload
CEnd-to-end check fails; everything passed last sessionOne-suspect rule: what changed since green? git diff and the journal answer in a minute; if truly nothing changed, suspect rung D, something died overnight
CDashboard frozen; components all look aliveHalf-split at the broker first (Week 11's table covers the downstream half: null bridge, stale latest, cache); if the firehose is healthy, the fault is bridge-or-later, go to curl
Dsystemctl status: status=203/EXECExecStart path is not an executable: a ~ systemd will not expand, a typo, or a missing venv; use absolute paths, then daemon-reload and start again
DService log: ModuleNotFoundError at the first importThe system python ran, not the venv's: point ExecStart at /home/username/project/venv/bin/python explicitly; a service has no activate
DService active (running), but port 5000 refusesThe process lives but is not serving: read journalctl -u for the missing Flask banner or a traceback; a banner on 127.0.0.1 is Week 11's bind wall inside a unit
DUptime looks fine, dashboard blinks stale every few minutesThe crash-loop disguise: Restart= is resurrecting a program that keeps dying; journalctl -u shows the repeated trace; fix the cause the log names, keep the restart
DWorked before the reboot, dead afterStart-order or enablement: is the unit enabled? Does it declare After=network-online.target mosquitto.service? Boot is the one time order is not up to you
DYesterday's rung broke while building today'sA regression: the verify step was skipped; git diff against the last tag that passed, and from now on re-run old rungs after every merge, two minutes of insurance
anygit pull greets the studio with conflict markersWeek 1's ritual, under deadline now: open the file, keep the right lines between the markers, delete the markers, commit; prevention is the stand-up's job, agree who owns which file today

Before you ask for help

  • Which ladder rung you are on, and what the rungs below it showed (the instrument's actual output, not a summary).
  • The half-split so far: which probe, which reading, which half survived.
  • What changed since the last green end-to-end check: git diff, or the honest "we skipped the check".
  • For service faults, the last 30 lines of journalctl -u your-unit, not a screenshot of status alone.

The studio queue rewards exactly this preparation: a pair arriving with rung, probe and diff gets unstuck in two minutes, which is the latency the whole week is designed around.