Control: closing the loop
420-302-VA · WEEK 10 · FALL 2026

Stage 5 of 6 · Lab · about 30 minutes, plus the ladder

Tuning & the big loop

A PID with wrong gains is worse than no controller, and there is no formula that reads your plant's mind. Tuning is an ordered experiment with a written log, and this page is the method, plus the week's second act: lifting the loop off the board and onto the network, where the real architectural question of IoT control lives.

Tuning is a logged experiment

The manual recipe, one knob at a time, in this order, logging every trial:

  1. Zero everything but P. Raise Kp from small until the step response is brisk and just beginning to ring, then back it off to roughly half or two-thirds of that edge. You now have speed and a known offset.
  2. Add I until the offset dies politely. Raise Ki from small until ess reaches zero within a second or two of settling. Too far shows up as a slow, rolling overshoot after arrival.
  3. Add D only if overshoot needs shaving. Small doses of Kd; stop at the first sign of output shimmer. Kd = 0 is a finding, not a failure.
  4. Re-grade all four verification tests from the build page, both steps and both disturbances, because a tuning that only looks good on the test you tuned with is not tuned.

The log is the deliverable. docs/tuning.md gets one row per trial, gains, what you observed, what you changed and why, and Week 9's CSV logger rung becomes an instrument: log timestamp,sp,pv,u during each step test and the curve is yours to grade offline instead of squinting at a terminal.

Reading the curve

Experienced hands diagnose a loop from the shape alone. The three shapes you will actually meet:

Three small panels, each a setpoint step on the same plant. Too timid: the curve crawls upward and sits below the setpoint. About right: a brisk rise, one small hump, then flat on the setpoint. Too hot: fast rise then sustained ringing around the setpoint.Too timidSPcrawls, arrives late if everAbout rightSPbrisk, one hump, doneToo hotSPrings and keeps ringing
Three shapes you will actually meet. Same plant, same step, three tunings. Learn these silhouettes the way you learned Week 4's tracebacks: the shape names the fault before you touch a knob, and the table below gives the move for each.
The curve showsDiagnosisMove
Slow crawl, settles short of SPToo timid: Kp low; offset says Ki missing or tinyRaise Kp; then bring in Ki
Brisk rise, one modest hump, settles clean at SPAbout rightLog it and stop; resist one more "improvement"
Ringing that persists or growsToo hot: gain past what the loop's delay toleratesCut Kp (and Ki); re-approach the edge slowly
Arrives, then a slow roll past SP and lazy returnIntegral overdone or windup tailLower Ki; confirm anti-windup is active
Fuzzy, trembling output at restD amplifying noiseLower Kd; confirm averaging; consider PI

Ziegler–Nichols: the classical recipe

In 1942, Ziegler and Nichols published the tuning rules the industry still names first. The ultimate-gain method: with I and D off, raise Kp until the loop oscillates steadily, neither dying nor growing. Call that gain Ku and the oscillation's period Tu, then set:

ControllerKpKiKd
P0.5·Ku——
PI0.45·Ku0.54·Ku/Tu—
PID0.6·Ku1.2·Ku/Tu0.075·Ku·Tu

Honesty about the classic: it was designed for fast disturbance rejection on sluggish industrial plants and lands deliberately aggressive, about a quarter-decay ringing per oscillation, so treat its numbers as a starting point to soften, not gospel. Its real value for you is conceptual: it names the edge (Ku, where your P-sweep went unstable) and scales every gain from two measured properties of your plant, which is the same empiricism as the manual recipe, formalized. Try it on the bench once; compare to your hand tuning in the log.

The big loop: control over the network

Everything so far ran on one board: sensor, controller and actuator sharing a chip, a local loop. The course architecture puts the control app on the Pi, so now the loop stretches across the bench, Week 9's channels carrying it:

The distributed control loop. On the ESP32 node: sensor and actuator. Telemetry flows node to broker to Pi on the light topic; the Pi's controller computes duty and publishes on the cmd topic; commands flow broker to node, which applies them. The loop's ring now includes two network hops, adding delay both ways. A clock annotation marks each hop. ESP32 node sensor · read_pct() actuator · set_brightness() Broker routes both ways Raspberry Pi controller PID in LightController …/light → (PV, once a second) ← …/cmd (duty) + hop delay + hop delay + hop delay + hop delay The ring is the same ring; it is just longer now, and every extra millisecond of it is delay inside the loop.
The same block diagram, stretched over Wi-Fi. Sensor and actuator stay on the node; the controller moves to the Pi; the broker carries both halves of the ring. The telemetry period (1 s) now is the sample period, and both hops add dead time, the exact quantity the stability warning priced.

Building it is Week 9 reuse: a LightController(LightMonitor) whose react() runs u = pid.update(sp, pv, dt) (with dt measured by time.monotonic() between messages) and publishes the duty on the cmd topic; the node's command handler applies it. It works, and it visibly works differently: with a 1 s sample period and network jitter, the central loop is softer and slower, and aggressive gains that the local loop tolerated will ring here. That comparison is the lab's point, not a defect.

Where should control live?

You have now run the same law in two places, which earns you the real question, one of the central architecture decisions of industrial IoT:

Local (on the node)

Milliseconds of loop delay, immune to network loss, keeps working unplugged from everything. This is why fast, safety-relevant loops, motor drives, flight controllers, your PLC's own PID blocks, run at the edge, next to the plant. The modern name for the principle is edge computing; your program has called it a PLC for two years.

Central (on the Pi)

Sees every node at once, coordinates across them, logs everything, retunes without reflashing. This is the supervisory layer, SCADA in industrial language, and it is the right home for slow decisions, coordination and oversight, exactly because it is the wrong home for fast ones.

Industry's standard answer uses both: fast loops local, supervision central, the supervisor adjusting the local loops' setpoints rather than their outputs. The ladder's rung 3 builds precisely that pattern on your bench, and it is the architecture your LIA project should default to.

The ladder

  1. Tune by the book. P, then PI, then (maybe) PID by the recipe, with a CSV log per step test and a docs/tuning.md row per trial. Finish with the four verification tests and your final gains justified in one line each.
  2. The Ziegler–Nichols detour. Find Ku and Tu on your plant, compute the table's PID gains, run the same step test, and write two sentences comparing it to your hand tuning. You have now tuned a loop two ways and can argue about it, which is the skill.
  3. Supervisory control. Node keeps its local PID; teach its command handler sp:60 so the Pi adjusts the setpoint over MQTT. Demonstrate the Pi walking the setpoint 30 → 60 → 45 while the local loop does the holding. This is the professional pattern from the cards above, running on your bench.
  4. The fully central loop. Node falls back to publish-and-obey (Week 9's rung 3); Pi runs LightController with the PID. Grade it against the local loop with the same step test, and note what the 1 s sample period and the two hops cost in overshoot and settling.
  5. Stretch: the step-test instrument. A small Pi script that publishes a setpoint step, logs sp,pv,u to CSV for ten seconds, then computes rise time, overshoot and settling time from the data automatically. You have built the thing a loop-tuning consultant actually carries, and Week 11's dashboard will happily display its verdicts.

Checklist for this stage

Check yourself

Why does the recipe tune P before I, and I before D?
Each term's symptom is only readable against the previous one's baseline. Kp sets the loop's basic speed and stability margin; Ki's job (killing the offset) and its overdose symptom (slow overshoot) are only visible once Kp is sane; Kd exists to shave the overshoot the first two produce. Tuning D first is adjusting the brakes of a car with no engine.
Your Ku came out at 9 with Tu = 0.9 s. Give the ZN PID gains and say what response character to expect.
Kp = 0.6·9 = 5.4; Ki = 1.2·9/0.9 = 12; Kd = 0.075·9·0.9 ≈ 0.61. Expect it fast and aggressive, quarter-decay ringing on steps, strong against disturbances. A comfortable bench tuning typically softens from there, which is the published rules working as intended: a principled starting point.
The same gains that ran clean locally ring in the central loop. Explain with the stability sentence.
Gain buys speed; delay sells stability. Moving the controller across the network added dead time (two hops) and stretched the sample period to the telemetry rate, so the loop's total delay grew while the gains stayed, pushing it past the edge the local loop sat safely inside. Central loops must be tuned softer, or sampled faster, or both.
In the supervisory pattern, the Pi never publishes a duty. What does it publish, and why is that the robust division of labour?
It publishes setpoints; the node's local PID does the holding. The fast ring (sense, decide, act) stays milliseconds long and survives network loss, the node just keeps holding the last setpoint, while the slow, smart decisions travel the network where their latency is harmless. Fast where physics lives, supervision where the overview lives.
Why does the central controller measure dt with time.monotonic() between messages instead of assuming 1.0 s?
The telemetry period is nominal, not guaranteed: Wi-Fi retries, broker scheduling and publish jitter stretch and bunch the arrivals. The integral and derivative scale with Δt, so assuming 1.0 s injects the jitter straight into the math, while measuring between callbacks keeps both terms honest, the build page's rule surviving the move to the Pi.