Utility · keep open during the studio
Troubleshoot
Integration week has its own failure signature: components that each work alone and a system that does not. The ladder below orders the search the way the whole hub taught, pieces, then pairs, then the whole, then time, because a fault's level decides its instrument, and climbing in order is half-splitting by another name.
Debug in layers
Symptoms and fixes
| Rung | Symptom | Likely cause and fix |
|---|---|---|
| A | Component passes alone, fails the moment it is joined | Almost always a contract mismatch, not a component bug: diff what each side actually sends and expects (firehose, curl -v) against docs/topics.md; the file, or one side, is wrong |
| A | Partner's piece "works on their laptop", fails on the bench | Environment, not code: missing package in the bench venv, hard-coded path or address, or a config value living outside config_example.py's pattern; reproduce on the bench with their stub commands |
| B | Firehose shows the topic twice with different spellings | Two programs disagree with the contract (Light vs light, s07 vs s12): fix the deviant side to match docs/topics.md, then clear any retained ghost with mosquitto_pub -r -n |
| B | Rule sees messages but crashes parsing some of them | A stub or an old publisher is sending a second payload format; the firehose's -v shows which; kill the stray publisher and make the rule log, not die, on a malformed payload |
| C | End-to-end check fails; everything passed last session | One-suspect rule: what changed since green? git diff and the journal answer in a minute; if truly nothing changed, suspect rung D, something died overnight |
| C | Dashboard frozen; components all look alive | Half-split at the broker first (Week 11's table covers the downstream half: null bridge, stale latest, cache); if the firehose is healthy, the fault is bridge-or-later, go to curl |
| D | systemctl status: status=203/EXEC | ExecStart path is not an executable: a ~ systemd will not expand, a typo, or a missing venv; use absolute paths, then daemon-reload and start again |
| D | Service log: ModuleNotFoundError at the first import | The system python ran, not the venv's: point ExecStart at /home/username/project/venv/bin/python explicitly; a service has no activate |
| D | Service active (running), but port 5000 refuses | The process lives but is not serving: read journalctl -u for the missing Flask banner or a traceback; a banner on 127.0.0.1 is Week 11's bind wall inside a unit |
| D | Uptime looks fine, dashboard blinks stale every few minutes | The crash-loop disguise: Restart= is resurrecting a program that keeps dying; journalctl -u shows the repeated trace; fix the cause the log names, keep the restart |
| D | Worked before the reboot, dead after | Start-order or enablement: is the unit enabled? Does it declare After=network-online.target mosquitto.service? Boot is the one time order is not up to you |
| D | Yesterday's rung broke while building today's | A regression: the verify step was skipped; git diff against the last tag that passed, and from now on re-run old rungs after every merge, two minutes of insurance |
| any | git pull greets the studio with conflict markers | Week 1's ritual, under deadline now: open the file, keep the right lines between the markers, delete the markers, commit; prevention is the stand-up's job, agree who owns which file today |
Before you ask for help
- Which ladder rung you are on, and what the rungs below it showed (the instrument's actual output, not a summary).
- The half-split so far: which probe, which reading, which half survived.
- What changed since the last green end-to-end check:
git diff, or the honest "we skipped the check". - For service faults, the last 30 lines of
journalctl -u your-unit, not a screenshot ofstatusalone.
The studio queue rewards exactly this preparation: a pair arriving with rung, probe and diff gets unstuck in two minutes, which is the latency the whole week is designed around.