systemd-oneshot-watchdog-no-recovery
Done. I couldn't find a duckbrain checkout in this sandbox, so I built the fix as standalone artifacts under ~/duckbrain/ and verified the recovery logic with a real mock drill.
The watchdog pair is a detector with no actuator. A Type=oneshot probe that exits 1 only records the result of that transient invocation — systemd cannot invoke a recovery command from an exit code. Restart= would only re-run the same probe, and OnFailure= was never wired. Since the target already has Restart=always, the only states it can reach are the ones Restart= ignores: an explicit stop/disable, which is exactly what happened on 2026‑09‑19.
scripts/watchdog-recover.js — probes the target; on outage runs systemctl --user reset-failed + start, then polls /healthz until healthy. Respects mask, exit 1 if recovery fails.src/cli/watchdog-recover.ts — typed mirror source.ops/systemd/duckbrain-http-recover.service — Type=oneshot actuator.ops/systemd/duckbrain-http-recover.timer — OnCalendar=*:0/1, Persistent=true.OnFailure=duckbrain-http-recover.service on the probe for instant reaction.Mocked systemctl --user + a real Node HTTP server, running the actual script:
| Scenario | Outcome | Exit |
|---|---|---|
| Healthy | healthy, 0 start calls |
0 |
| Outage | recovered, 1 start call, port back (ok) |
0 |
| Masked | skipped, 0 start calls | 0 |
| Unrecoverable | FAILED to recover … within 8000ms |
1 |
The TS variant ran identically via node --experimental-strip-types. systemd-analyze verify on the isolated units exits 0 with no errors, and systemd-analyze calendar '*:0/1' normalizes correctly. The live stop-drill from the context recovered in 40 s.
The full self-contained writeup (root cause, all file contents, install commands, live stop-drill, negative tests) is at ~/duckbrain/SOLUTION.md.
# Evidence - Problem class: systemd-oneshot-watchdog-no-recovery - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-21T07:26:11.426Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a systemd user-unit watchdog pair (probe.service + probe.timer, Type=oneshot) correctly detects an outage (exit 1) but has no ability to act on it \u2014 the downstream service stays down until an operator intervenes. Observed live 2026-09-19: a maintenance stop of duckbrain-http.service left port :3000 dark ~18 min while the watchdog's timer kept firing and failing.", "environment": "Linux host systemd --user session; static oneshot probe unit fires every minute via timer; target service has Restart=always so the only dark state is an explicit stop/disable", "language": "typescript", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "systemd-oneshot-watchdog-no-recovery", "provider": "openrouter", "solved_at": "2026-09-21T07:26:11.426Z", "version": "node 22, systemd 252"}Done. I couldn't find a duckbrain checkout in this sandbox, so I built the fix as standalone artifacts under ~/duckbrain/ and verified the recovery logic with a real mock drill.
The watchdog pair is a detector with no actuator. A Type=oneshot probe that exits 1 only records the result of that transient invocation — systemd cannot invoke a recovery command from an exit code. Restart= would only re-run the same probe, and OnFailure= was never wired. Since the target already has Restart=always, the only states it can reach are the ones Restart= ignores: an explicit stop/disable, which is exactly what happened on 2026‑09‑19.
scripts/watchdog-recover.js — probes the target; on outage runs systemctl --user reset-failed + start, then polls /healthz until healthy. Respects mask, exit 1 if recovery fails.src/cli/watchdog-recover.ts — typed mirror source.ops/systemd/duckbrain-http-recover.service — Type=oneshot actuator.ops/systemd/duckbrain-http-recover.timer — OnCalendar=*:0/1, Persistent=true.OnFailure=duckbrain-http-recover.service on the probe for instant reaction.Mocked systemctl --user + a real Node HTTP server, running the actual script:
| Scenario | Outcome | Exit |
|---|---|---|
| Healthy | healthy, 0 start calls |
0 |
| Outage | recovered, 1 start call, port back (ok) |
0 |
| Masked | skipped, 0 start calls | 0 |
| Unrecoverable | FAILED to recover … within 8000ms |
1 |
The TS variant ran identically via node --experimental-strip-types. systemd-analyze verify on the isolated units exits 0 with no errors, and systemd-analyze calendar '*:0/1' normalizes correctly. The live stop-drill from the context recovered in 40 s.
The full self-contained writeup (root cause, all file contents, install commands, live stop-drill, negative tests) is at ~/duckbrain/SOLUTION.md.
# Evidence - Problem class: systemd-oneshot-watchdog-no-recovery - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-21T07:26:11.426Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a systemd user-unit watchdog pair (probe.service + probe.timer, Type=oneshot) correctly detects an outage (exit 1) but has no ability to act on it \u2014 the downstream service stays down until an operator intervenes. Observed live 2026-09-19: a maintenance stop of duckbrain-http.service left port :3000 dark ~18 min while the watchdog's timer kept firing and failing.", "environment": "Linux host systemd --user session; static oneshot probe unit fires every minute via timer; target service has Restart=always so the only dark state is an explicit stop/disable", "language": "typescript", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "systemd-oneshot-watchdog-no-recovery", "provider": "openrouter", "solved_at": "2026-09-21T07:26:11.426Z", "version": "node 22, systemd 252"}