Class: graceful SIGTERM on a Restart=on-failure user service permanently downs the daemon; an over-broad pkill -f stop script makes any scratch cleanup fatal.
Saved to ~/SOLUTION.md. Full contents:
systemd-on-failure-graceful-sigterm-dark-port (OPS-001)Class: graceful SIGTERM on a Restart=on-failure user service permanently downs the daemon; an over-broad pkill -f stop script makes any scratch cleanup fatal.
Repo: wojons/duckbrain · Commit: b7d9914 · Judge: 8292fe3d FAIL → 3a5ff2a0 PASS
Restart=on-failure treats exit 0 as successsystemd restarts a unit under Restart=on-failure only when the main process:
| Termination | on-failure behavior |
|---|---|
| Exits with non-zero status | restart |
| Dies from an unhandled signal (e.g. SIGTERM not caught) | restart |
| Times out / watchdog fires | restart |
| Exits with status 0 (a handled SIGTERM/SIGINT) | no restart — unit goes to inactive (dead) |
DuckBrain installs a graceful shutdown handler that traps SIGTERM/SIGINT, drains the HTTP listener, closes the DuckDB handle, then calls process.exit(0). That is correct application behavior but interacts fatally with Restart=on-failure: systemd sees a clean exit, marks the unit inactive (dead), and never schedules a restart. NRestarts stays put and the journal shows Deactivated successfully instead of Stopping ... — exactly why the incident looked like a mysterious 7-minute outage with no systemctl stop in the logs.
Any out-of-band SIGTERM becomes permanent downtime: a stray pattern kill, judge/worker cleanup, an OOM-adjacent TERM, an operator kill <pid>. Recovery requires a manual systemctl --user start (or reboot with lingering).
pkill -f 'duckbrain.*http'The production unit's ExecStart cmdline matches that pattern exactly, so any cleanup (npm run stop, CI teardown, scratch stop) TERMs production, not the scratch instance. Combined with §1.1 it is a one-command permanent outage.
The real ~/.config/systemd/user/duckbrain-http.service carries the authoritative ExecStart (auth flags, unix-socket path, env). Rewriting it from the repo clobbers operator config. Durability directives belong in a drop-in that overrides only Restart* / StartLimit*.
A tier-2 "live systemd verification" criterion cannot be met by committed files. The foreman must install the drop-in, daemon-reload, enable/start the timer, and run the live SIGTERM probe before re-judging. 8292fe3d failed because the unit was never installed live; 3a5ff2a0 passed once it was.
Restart=always drop-in~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf:
[Unit]
# Start-rate limiting lives in [Unit] on modern systemd (>= 230).
StartLimitIntervalSec=300
StartLimitBurst=10
[Service]
# Restart regardless of exit status, including the graceful exit(0) path.
Restart=always
RestartSec=3
Install without touching the base unit:
mkdir -p ~/.config/systemd/user/duckbrain-http.service.d
cat > ~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf <<'EOF'
[Unit]
StartLimitIntervalSec=300
StartLimitBurst=10
[Service]
Restart=always
RestartSec=3
EOF
systemctl --user daemon-reload
systemctl --user restart duckbrain-http.service
systemctl --user show -p Restart -p RestartUSec -p StartLimitIntervalUSec -p StartLimitBurst duckbrain-http.service
# Restart=always RestartUSec=3s StartLimitIntervalUSec=5min StartLimitBurst=10
Confirm the real ExecStart survived:
systemctl --user cat duckbrain-http.service
systemctl --user show -p ExecStart duckbrain-http.service
Enable lingering if not already:
loginctl enable-linger "$USER"
Replace every pkill -f 'duckbrain.*http' with a helper that can only signal the PID named by a pidfile, after validating liveness, format, daemon identity, and port. Refuses stale / malformed / wrong-port pidfiles.
scripts/scoped-stop.js:
#!/usr/bin/env node
'use strict';
const fs = require('fs');
const path = require('path');
function die(msg, code = 2) {
console.error(`scoped-stop: ${msg}`);
process.exit(code);
}
function parseArgs(argv) {
const out = { port: null, pidfile: null, timeoutSec: 15, force: false };
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
if (a === '--port' || a === '-p') out.port = argv[++i];
else if (a === '--pidfile') out.pidfile = argv[++i];
else if (a === '--timeout') out.timeoutSec = Number(argv[++i]);
else if (a === '--force') out.force = true;
else die(`unknown argument: ${a}`);
}
if (!out.port || !/^\d+$/.test(String(out.port))) die('--port <number> is required');
if (!Number.isFinite(out.timeoutSec) || out.timeoutSec < 0) die('--timeout must be >= 0');
out.port = Number(out.port);
return out;
}
function defaultPidfile(port) {
const base = process.env.XDG_RUNTIME_DIR || '/tmp';
return path.join(base, `duckbrain-http-${port}.pid`);
}
function readPid(pidfile) {
let raw;
try {
raw = fs.readFileSync(pidfile, 'utf8');
} catch (e) {
die(`cannot read pidfile ${pidfile}: ${e.message}`);
}
const trimmed = raw.trim();
if (!/^\d+$/.test(trimmed)) die(`malformed pidfile ${pidfile}: ${JSON.stringify(raw)}`);
const pid = Number(trimmed);
if (!Number.isInteger(pid) || pid <= 1) die(`malformed pidfile ${pidfile}: bad pid ${trimmed}`);
return pid;
}
function isAlive(pid) {
try {
process.kill(pid, 0);
return true;
} catch (e) {
return e.code === 'EPERM';
}
}
function cmdline(pid) {
try {
return fs.readFileSync(`/proc/${pid}/cmdline`, 'utf8').split('\0').filter(Boolean);
} catch (e) {
return [];
}
}
function isDuckbrainHttp(args, port) {
const joined = args.join(' ');
if (!/duckbrain/i.test(joined)) return false;
for (let i = 0; i < args.length; i++) {
const a = args[i];
if ((a === '--port' || a === '-p') && args[i + 1] === String(port)) return true;
if (a === `--port=${port}`) return true;
if (a === `:${port}`) return true;
if (a === `--listen=:${port}`) return true;
}
return false;
}
async function waitForExit(pid, timeoutSec) {
const deadline = Date.now() + timeoutSec * 1000;
while (Date.now() < deadline) {
if (!isAlive(pid)) return true;
await new Promise((r) => setTimeout(r, 250));
}
return !isAlive(pid);
}
(async () => {
const opts = parseArgs(process.argv.slice(2));
const pidfile = opts.pidfile || defaultPidfile(opts.port);
if (!fs.existsSync(pidfile)) die(`no pidfile ${pidfile} (refusing to pattern-kill)`);
const pid = readPid(pidfile);
if (!isAlive(pid)) die(`stale pidfile ${pidfile}: pid ${pid} is not running`, 3);
const args = cmdline(pid);
if (args.length === 0) die(`cannot read cmdline for pid ${pid} (permission or race)`, 3);
if (!isDuckbrainHttp(args, opts.port)) {
die(
`wrong-port/mismatched pidfile ${pidfile}: pid ${pid} is not a duckbrain http daemon ` +
`for port ${opts.port}\n cmdline: ${args.join(' ')}`,
4,
);
}
console.log(`scoped-stop: SIGTERM -> pid ${pid} (port ${opts.port})`);
process.kill(pid, 'SIGTERM');
const gone = await waitForExit(pid, opts.timeoutSec);
if (!gone) {
if (!opts.force) die(`pid ${pid} did not exit within ${opts.timeoutSec}s (use --force to SIGKILL)`, 5);
console.warn(`scoped-stop: SIGKILL -> pid ${pid}`);
process.kill(pid, 'SIGKILL');
if (!(await waitForExit(pid, 5))) die(`pid ${pid} survived SIGKILL`, 6);
}
try { fs.unlinkSync(pidfile); } catch (_) {}
console.log(`scoped-stop: pid ${pid} stopped cleanly`);
})().catch((e) => die(e.stack || String(e), 1));
TypeScript source at src/cli/scoped-stop.ts; scripts/scoped-stop.js is the compiled artifact. Wire into package.json:
{ "scripts": { "stop": "node scripts/scoped-stop.js --port 3000" } }
Remove every pkill -f 'duckbrain.*http', pkill -f duckbrain, and bare kill $(pgrep -f ...) from scripts, Makefiles, and CI.
The daemon writes the pidfile (atomically) at startup and unlinks it on shutdown:
const pidfile = path.join(process.env.XDG_RUNTIME_DIR || '/tmp', `duckbrain-http-${port}.pid`);
fs.writeFileSync(`${pidfile}.${process.pid}.tmp`, String(process.pid));
fs.renameSync(`${pidfile}.${process.pid}.tmp`, pidfile);
const cleanup = () => { try { fs.unlinkSync(pidfile); } catch (_) {} };
process.on('SIGTERM', () => { cleanup(); server.close(() => process.exit(0)); });
process.on('SIGINT', () => { cleanup(); server.close(() => process.exit(0)); });
/health probe timerMake /health auth-exempt (register before the auth guard):
// src/http/router.ts
router.get('/health', healthHandler); // must be BEFORE auth middleware
app.use(authMiddleware); // keys apply only below this line
scripts/health-check.js:
#!/usr/bin/env node
'use strict';
const http = require('http');
const args = process.argv.slice(2);
function opt(name, dflt) {
const i = args.indexOf(name);
return i >= 0 ? args[i + 1] : dflt;
}
const port = Number(opt('--port', '3000'));
const host = opt('--url', `http://<ip-address>:${port}/health`);
const timeoutMs = Number(opt('--timeout-ms', '5000'));
const req = http.get(host, { timeout: timeoutMs }, (res) => {
res.resume();
// 200 = healthy; 503 = degraded but ALIVE (dependency-health contract).
if (res.statusCode === 200 || res.statusCode === 503) {
console.log(`health: alive status=${res.statusCode}`);
process.exit(0);
}
console.error(`health: DARK status=${res.statusCode}`);
process.exit(1);
});
req.on('timeout', () => { req.destroy(new Error(`timeout after ${timeoutMs}ms`)); });
req.on('error', (e) => {
console.error(`health: DARK connection failure: ${e.message}`);
process.exit(1);
});
~/.config/systemd/user/duckbrain-http-health.service:
[Unit]
Description=DuckBrain HTTP /health probe (oneshot)
After=duckbrain-http.service
[Service]
Type=oneshot
# Do NOT set Restart=always here; the timer is the retry mechanism.
ExecStart=/usr/bin/node %h/duckbrain/scripts/health-check.js --port 3000
~/.config/systemd/user/duckbrain-http-health.timer:
[Unit]
Description=Probe DuckBrain /health every minute
[Timer]
OnBootSec=60
OnUnitActiveSec=60
AccuracySec=5
Persistent=true
Unit=duckbrain-http-health.service
[Install]
WantedBy=timers.target
Install and enable:
cp ops/systemd/duckbrain-http-health.service ~/.config/systemd/user/
cp ops/systemd/duckbrain-http-health.timer ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now duckbrain-http-health.timer
systemctl --user list-timers duckbrain-http-health.timer
Why 503 counts as alive: /health may report degraded (a dependency down) while the process serves. Treating 503 as dark causes false restart/alert loops. Only transport failure (ECONNREFUSED / timeout) or an unexpected status is dark.
Run on the host that actually runs the unit.
systemctl --user is-active duckbrain-http.service
systemctl --user show -p Restart --value duckbrain-http.service # always
systemctl --user is-enabled duckbrain-http-health.timer
loginctl show-user "$USER" -p Linger # Linger=yes
OLD_PID=$(systemctl --user show -p MainPID --value duckbrain-http.service)
OLD_RESTARTS=$(systemctl --user show -p NRestarts --value duckbrain-http.service)
echo "old pid=$OLD_PID restarts=$OLD_RESTARTS"
kill -TERM "$OLD_PID" # out-of-band, not systemctl stop
NEW_PID=0
NEW_RESTARTS=$OLD_RESTARTS
for i in $(seq 1 30); do
NEW_PID=$(systemctl --user show -p MainPID --value duckbrain-http.service)
NEW_RESTARTS=$(systemctl --user show -p NRestarts --value duckbrain-http.service)
if [ -n "$NEW_PID" ] && [ "$NEW_PID" != "0" ] && [ "$NEW_PID" != "$OLD_PID" ]; then
break
fi
sleep 1
done
echo "old=$OLD_PID new=$NEW_PID restarts_before=$OLD_RESTARTS restarts_after=$NEW_RESTARTS"
[ "$NEW_PID" != "0" ] && [ "$NEW_PID" != "$OLD_PID" ] && [ "$NEW_RESTARTS" -gt "$OLD_RESTARTS" ] \
&& echo "PASS: restarted" || { echo "FAIL: not restarted"; exit 1; }
Expected journal evidence:
journalctl --user -u duckbrain-http.service --since '2 minutes ago' --no-pager | tail -n 20
# ... Deactivated successfully (graceful handler ran)
# ... Scheduled restart job, restart counter is at N
# ... Started DuckBrain HTTP ...
Scheduled restart job + incremented counter is conclusive proof Restart=always caught exit 0.
curl -s -o /dev/null -w 'health=%{http_code}\n' http://<ip-address>:3000/health
# health=200 (503 is also acceptable = degraded-but-alive)
curl -s -o /dev/null -w 'authed=%{http_code}\n' \
-H "Authorization: Bearer $DUCKBRAIN_API_KEY" \
http://<ip-address>:3000/v1/ping
# authed=200
systemctl --user start duckbrain-http-health.service
journalctl --user -u duckbrain-http-health.service --since '1 minute ago' --no-pager | tail
systemctl --user list-timers duckbrain-http-health.timer --no-pager
# 1. Normal scoped stop.
node scripts/scoped-stop.js --port 3000
# 2. Malformed pidfile is refused.
printf 'not-a-pid\n' > /tmp/duckbrain-http-3000.pid
node scripts/scoped-stop.js --port 3000; echo "exit=$?" # exit=2
# 3. Stale pidfile is refused.
node -e 'require("fs").writeFileSync("/tmp/duckbrain-http-3000.pid","999999\n")'
node scripts/scoped-stop.js --port 3000; echo "exit=$?" # exit=3
# 4. Wrong-port pidfile is refused instead of killing the daemon. exit=4.
# 5. Production daemon survives an over-broad scratch cleanup.
grep -rn "pkill" package.json scripts/ Makefile .github/ 2>/dev/null || echo "no pkill stop paths"
| File | Change |
|---|---|
~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf |
Restart=always, RestartSec=3, StartLimitIntervalSec=300, StartLimitBurst=10 (drop-in; base unit untouched) |
ops/systemd/duckbrain-http.service |
repo reference copy; must not clobber live ExecStart |
ops/systemd/duckbrain-http-health.service + .timer |
oneshot /health probe, 60 s interval, enabled + started live |
scripts/scoped-stop.js, src/cli/scoped-stop.ts |
pidfile-scoped SIGTERM with live/format/daemon/port validation |
scripts/health-check.js, src/cli/health-check.ts |
200 and 503 = alive; connection failure = dark |
| HTTP router | GET /health registered before auth middleware (auth-exempt) |
| daemon startup/shutdown | write/unlink port-keyed pidfile atomically |
| all stop scripts / CI | delete pkill -f 'duckbrain.*http' in favor of the scoped helper |
A tier-2 criterion demanding live systemd verification cannot be satisfied by repo-owned unit files. The foreman must actually:
systemctl --user daemon-reload,enable --now the health timer,kill -TERM $MainPID probe and confirm a new MainPID and incremented NRestarts,curl /health = 200 (or 503) and an authed GET = 200,before re-judging. Judge 8292fe3d returned FAIL because the unit was still uninstalled on the host; 3a5ff2a0 passed only after this install-and-probe sequence ran. Repo files are necessary but not sufficient — the live state is the artifact under test.
# Evidence - Problem class: systemd-on-failure-graceful-sigterm-dark-port - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-12T12:56:15.432Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a production daemon under a systemd user unit with Restart=on-failure received an out-of-band SIGTERM, logged its graceful shutdown, exited status=0/SUCCESS, and stayed dark for ~7 minutes; NRestarts never incremented and the journal showed no Stopping lines (proving it was not systemctl stop). Root cause: Restart=on-failure treats exit status 0 as success and never restarts, so ANY graceful SIGTERM/SIGINT (stray pkill pattern kill, judge/worker cleanup, OOM-adjacent TERM) permanently downs the service until manual intervention. A second latent hazard: the repo stop script was pkill -f 'duckbrain.*http', which matches the production daemon cmdline exactly, so any scratch cleanup takes production down. Fix (three parts): (1) durability \u2014 add a drop-in (e.g. 10-restart-always.conf) with Restart=always, RestartSec=3, StartLimitIntervalSec=300, StartLimitBurst=10, then daemon-reload; do NOT overwrite the base unit when drop-ins carry the real ExecStart (auth flags, unix socket). (2) blast radius \u2014 replace pkill stop scripts with a pidfile-scoped SIGTERM helper that validates the PID is live and its cmdline matches a daemon for the requested port, refusing stale/malformed/wrong-port pidfiles. (3) monitoring \u2014 oneshot health-check service + systemd timer probing /health every minute, treating HTTP 200 AND 503 as alive (503 degraded may be an intentional dependency-health contract) and connection failure as dark; /health must be auth-exempt so the probe needs no keys. Verification: kill -TERM <MainPID>, poll systemctl show -p MainPID,NRestarts until a new PID appears and NRestarts increments, then curl /health = 200 and an authed GET works. GitReins lesson: when a tier2 criterion demands 'live systemd verification', repo-owned unit files alone FAIL the judge \u2014 the foreman must install the unit/drop-in, daemon-reload, enable the timer, and run the live SIGTERM probe before re-judging.", "environment": "systemd user service, node daemon, DuckBrain HTTP :3000, Linux user lingering", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "systemd-on-failure-graceful-sigterm-dark-port", "provider": "openrouter", "solved_at": "2026-09-12T12:56:15.432Z", "version": ""}Saved to ~/SOLUTION.md. Full contents:
systemd-on-failure-graceful-sigterm-dark-port (OPS-001)Class: graceful SIGTERM on a Restart=on-failure user service permanently downs the daemon; an over-broad pkill -f stop script makes any scratch cleanup fatal.
Repo: wojons/duckbrain · Commit: b7d9914 · Judge: 8292fe3d FAIL → 3a5ff2a0 PASS
Restart=on-failure treats exit 0 as successsystemd restarts a unit under Restart=on-failure only when the main process:
| Termination | on-failure behavior |
|---|---|
| Exits with non-zero status | restart |
| Dies from an unhandled signal (e.g. SIGTERM not caught) | restart |
| Times out / watchdog fires | restart |
| Exits with status 0 (a handled SIGTERM/SIGINT) | no restart — unit goes to inactive (dead) |
DuckBrain installs a graceful shutdown handler that traps SIGTERM/SIGINT, drains the HTTP listener, closes the DuckDB handle, then calls process.exit(0). That is correct application behavior but interacts fatally with Restart=on-failure: systemd sees a clean exit, marks the unit inactive (dead), and never schedules a restart. NRestarts stays put and the journal shows Deactivated successfully instead of Stopping ... — exactly why the incident looked like a mysterious 7-minute outage with no systemctl stop in the logs.
Any out-of-band SIGTERM becomes permanent downtime: a stray pattern kill, judge/worker cleanup, an OOM-adjacent TERM, an operator kill <pid>. Recovery requires a manual systemctl --user start (or reboot with lingering).
pkill -f 'duckbrain.*http'The production unit's ExecStart cmdline matches that pattern exactly, so any cleanup (npm run stop, CI teardown, scratch stop) TERMs production, not the scratch instance. Combined with §1.1 it is a one-command permanent outage.
The real ~/.config/systemd/user/duckbrain-http.service carries the authoritative ExecStart (auth flags, unix-socket path, env). Rewriting it from the repo clobbers operator config. Durability directives belong in a drop-in that overrides only Restart* / StartLimit*.
A tier-2 "live systemd verification" criterion cannot be met by committed files. The foreman must install the drop-in, daemon-reload, enable/start the timer, and run the live SIGTERM probe before re-judging. 8292fe3d failed because the unit was never installed live; 3a5ff2a0 passed once it was.
Restart=always drop-in~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf:
[Unit]
# Start-rate limiting lives in [Unit] on modern systemd (>= 230).
StartLimitIntervalSec=300
StartLimitBurst=10
[Service]
# Restart regardless of exit status, including the graceful exit(0) path.
Restart=always
RestartSec=3
Install without touching the base unit:
mkdir -p ~/.config/systemd/user/duckbrain-http.service.d
cat > ~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf <<'EOF'
[Unit]
StartLimitIntervalSec=300
StartLimitBurst=10
[Service]
Restart=always
RestartSec=3
EOF
systemctl --user daemon-reload
systemctl --user restart duckbrain-http.service
systemctl --user show -p Restart -p RestartUSec -p StartLimitIntervalUSec -p StartLimitBurst duckbrain-http.service
# Restart=always RestartUSec=3s StartLimitIntervalUSec=5min StartLimitBurst=10
Confirm the real ExecStart survived:
systemctl --user cat duckbrain-http.service
systemctl --user show -p ExecStart duckbrain-http.service
Enable lingering if not already:
loginctl enable-linger "$USER"
Replace every pkill -f 'duckbrain.*http' with a helper that can only signal the PID named by a pidfile, after validating liveness, format, daemon identity, and port. Refuses stale / malformed / wrong-port pidfiles.
scripts/scoped-stop.js:
#!/usr/bin/env node
'use strict';
const fs = require('fs');
const path = require('path');
function die(msg, code = 2) {
console.error(`scoped-stop: ${msg}`);
process.exit(code);
}
function parseArgs(argv) {
const out = { port: null, pidfile: null, timeoutSec: 15, force: false };
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
if (a === '--port' || a === '-p') out.port = argv[++i];
else if (a === '--pidfile') out.pidfile = argv[++i];
else if (a === '--timeout') out.timeoutSec = Number(argv[++i]);
else if (a === '--force') out.force = true;
else die(`unknown argument: ${a}`);
}
if (!out.port || !/^\d+$/.test(String(out.port))) die('--port <number> is required');
if (!Number.isFinite(out.timeoutSec) || out.timeoutSec < 0) die('--timeout must be >= 0');
out.port = Number(out.port);
return out;
}
function defaultPidfile(port) {
const base = process.env.XDG_RUNTIME_DIR || '/tmp';
return path.join(base, `duckbrain-http-${port}.pid`);
}
function readPid(pidfile) {
let raw;
try {
raw = fs.readFileSync(pidfile, 'utf8');
} catch (e) {
die(`cannot read pidfile ${pidfile}: ${e.message}`);
}
const trimmed = raw.trim();
if (!/^\d+$/.test(trimmed)) die(`malformed pidfile ${pidfile}: ${JSON.stringify(raw)}`);
const pid = Number(trimmed);
if (!Number.isInteger(pid) || pid <= 1) die(`malformed pidfile ${pidfile}: bad pid ${trimmed}`);
return pid;
}
function isAlive(pid) {
try {
process.kill(pid, 0);
return true;
} catch (e) {
return e.code === 'EPERM';
}
}
function cmdline(pid) {
try {
return fs.readFileSync(`/proc/${pid}/cmdline`, 'utf8').split('\0').filter(Boolean);
} catch (e) {
return [];
}
}
function isDuckbrainHttp(args, port) {
const joined = args.join(' ');
if (!/duckbrain/i.test(joined)) return false;
for (let i = 0; i < args.length; i++) {
const a = args[i];
if ((a === '--port' || a === '-p') && args[i + 1] === String(port)) return true;
if (a === `--port=${port}`) return true;
if (a === `:${port}`) return true;
if (a === `--listen=:${port}`) return true;
}
return false;
}
async function waitForExit(pid, timeoutSec) {
const deadline = Date.now() + timeoutSec * 1000;
while (Date.now() < deadline) {
if (!isAlive(pid)) return true;
await new Promise((r) => setTimeout(r, 250));
}
return !isAlive(pid);
}
(async () => {
const opts = parseArgs(process.argv.slice(2));
const pidfile = opts.pidfile || defaultPidfile(opts.port);
if (!fs.existsSync(pidfile)) die(`no pidfile ${pidfile} (refusing to pattern-kill)`);
const pid = readPid(pidfile);
if (!isAlive(pid)) die(`stale pidfile ${pidfile}: pid ${pid} is not running`, 3);
const args = cmdline(pid);
if (args.length === 0) die(`cannot read cmdline for pid ${pid} (permission or race)`, 3);
if (!isDuckbrainHttp(args, opts.port)) {
die(
`wrong-port/mismatched pidfile ${pidfile}: pid ${pid} is not a duckbrain http daemon ` +
`for port ${opts.port}\n cmdline: ${args.join(' ')}`,
4,
);
}
console.log(`scoped-stop: SIGTERM -> pid ${pid} (port ${opts.port})`);
process.kill(pid, 'SIGTERM');
const gone = await waitForExit(pid, opts.timeoutSec);
if (!gone) {
if (!opts.force) die(`pid ${pid} did not exit within ${opts.timeoutSec}s (use --force to SIGKILL)`, 5);
console.warn(`scoped-stop: SIGKILL -> pid ${pid}`);
process.kill(pid, 'SIGKILL');
if (!(await waitForExit(pid, 5))) die(`pid ${pid} survived SIGKILL`, 6);
}
try { fs.unlinkSync(pidfile); } catch (_) {}
console.log(`scoped-stop: pid ${pid} stopped cleanly`);
})().catch((e) => die(e.stack || String(e), 1));
TypeScript source at src/cli/scoped-stop.ts; scripts/scoped-stop.js is the compiled artifact. Wire into package.json:
{ "scripts": { "stop": "node scripts/scoped-stop.js --port 3000" } }
Remove every pkill -f 'duckbrain.*http', pkill -f duckbrain, and bare kill $(pgrep -f ...) from scripts, Makefiles, and CI.
The daemon writes the pidfile (atomically) at startup and unlinks it on shutdown:
const pidfile = path.join(process.env.XDG_RUNTIME_DIR || '/tmp', `duckbrain-http-${port}.pid`);
fs.writeFileSync(`${pidfile}.${process.pid}.tmp`, String(process.pid));
fs.renameSync(`${pidfile}.${process.pid}.tmp`, pidfile);
const cleanup = () => { try { fs.unlinkSync(pidfile); } catch (_) {} };
process.on('SIGTERM', () => { cleanup(); server.close(() => process.exit(0)); });
process.on('SIGINT', () => { cleanup(); server.close(() => process.exit(0)); });
/health probe timerMake /health auth-exempt (register before the auth guard):
// src/http/router.ts
router.get('/health', healthHandler); // must be BEFORE auth middleware
app.use(authMiddleware); // keys apply only below this line
scripts/health-check.js:
#!/usr/bin/env node
'use strict';
const http = require('http');
const args = process.argv.slice(2);
function opt(name, dflt) {
const i = args.indexOf(name);
return i >= 0 ? args[i + 1] : dflt;
}
const port = Number(opt('--port', '3000'));
const host = opt('--url', `http://<ip-address>:${port}/health`);
const timeoutMs = Number(opt('--timeout-ms', '5000'));
const req = http.get(host, { timeout: timeoutMs }, (res) => {
res.resume();
// 200 = healthy; 503 = degraded but ALIVE (dependency-health contract).
if (res.statusCode === 200 || res.statusCode === 503) {
console.log(`health: alive status=${res.statusCode}`);
process.exit(0);
}
console.error(`health: DARK status=${res.statusCode}`);
process.exit(1);
});
req.on('timeout', () => { req.destroy(new Error(`timeout after ${timeoutMs}ms`)); });
req.on('error', (e) => {
console.error(`health: DARK connection failure: ${e.message}`);
process.exit(1);
});
~/.config/systemd/user/duckbrain-http-health.service:
[Unit]
Description=DuckBrain HTTP /health probe (oneshot)
After=duckbrain-http.service
[Service]
Type=oneshot
# Do NOT set Restart=always here; the timer is the retry mechanism.
ExecStart=/usr/bin/node %h/duckbrain/scripts/health-check.js --port 3000
~/.config/systemd/user/duckbrain-http-health.timer:
[Unit]
Description=Probe DuckBrain /health every minute
[Timer]
OnBootSec=60
OnUnitActiveSec=60
AccuracySec=5
Persistent=true
Unit=duckbrain-http-health.service
[Install]
WantedBy=timers.target
Install and enable:
cp ops/systemd/duckbrain-http-health.service ~/.config/systemd/user/
cp ops/systemd/duckbrain-http-health.timer ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now duckbrain-http-health.timer
systemctl --user list-timers duckbrain-http-health.timer
Why 503 counts as alive: /health may report degraded (a dependency down) while the process serves. Treating 503 as dark causes false restart/alert loops. Only transport failure (ECONNREFUSED / timeout) or an unexpected status is dark.
Run on the host that actually runs the unit.
systemctl --user is-active duckbrain-http.service
systemctl --user show -p Restart --value duckbrain-http.service # always
systemctl --user is-enabled duckbrain-http-health.timer
loginctl show-user "$USER" -p Linger # Linger=yes
OLD_PID=$(systemctl --user show -p MainPID --value duckbrain-http.service)
OLD_RESTARTS=$(systemctl --user show -p NRestarts --value duckbrain-http.service)
echo "old pid=$OLD_PID restarts=$OLD_RESTARTS"
kill -TERM "$OLD_PID" # out-of-band, not systemctl stop
NEW_PID=0
NEW_RESTARTS=$OLD_RESTARTS
for i in $(seq 1 30); do
NEW_PID=$(systemctl --user show -p MainPID --value duckbrain-http.service)
NEW_RESTARTS=$(systemctl --user show -p NRestarts --value duckbrain-http.service)
if [ -n "$NEW_PID" ] && [ "$NEW_PID" != "0" ] && [ "$NEW_PID" != "$OLD_PID" ]; then
break
fi
sleep 1
done
echo "old=$OLD_PID new=$NEW_PID restarts_before=$OLD_RESTARTS restarts_after=$NEW_RESTARTS"
[ "$NEW_PID" != "0" ] && [ "$NEW_PID" != "$OLD_PID" ] && [ "$NEW_RESTARTS" -gt "$OLD_RESTARTS" ] \
&& echo "PASS: restarted" || { echo "FAIL: not restarted"; exit 1; }
Expected journal evidence:
journalctl --user -u duckbrain-http.service --since '2 minutes ago' --no-pager | tail -n 20
# ... Deactivated successfully (graceful handler ran)
# ... Scheduled restart job, restart counter is at N
# ... Started DuckBrain HTTP ...
Scheduled restart job + incremented counter is conclusive proof Restart=always caught exit 0.
curl -s -o /dev/null -w 'health=%{http_code}\n' http://<ip-address>:3000/health
# health=200 (503 is also acceptable = degraded-but-alive)
curl -s -o /dev/null -w 'authed=%{http_code}\n' \
-H "Authorization: Bearer $DUCKBRAIN_API_KEY" \
http://<ip-address>:3000/v1/ping
# authed=200
systemctl --user start duckbrain-http-health.service
journalctl --user -u duckbrain-http-health.service --since '1 minute ago' --no-pager | tail
systemctl --user list-timers duckbrain-http-health.timer --no-pager
# 1. Normal scoped stop.
node scripts/scoped-stop.js --port 3000
# 2. Malformed pidfile is refused.
printf 'not-a-pid\n' > /tmp/duckbrain-http-3000.pid
node scripts/scoped-stop.js --port 3000; echo "exit=$?" # exit=2
# 3. Stale pidfile is refused.
node -e 'require("fs").writeFileSync("/tmp/duckbrain-http-3000.pid","999999\n")'
node scripts/scoped-stop.js --port 3000; echo "exit=$?" # exit=3
# 4. Wrong-port pidfile is refused instead of killing the daemon. exit=4.
# 5. Production daemon survives an over-broad scratch cleanup.
grep -rn "pkill" package.json scripts/ Makefile .github/ 2>/dev/null || echo "no pkill stop paths"
| File | Change |
|---|---|
~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf |
Restart=always, RestartSec=3, StartLimitIntervalSec=300, StartLimitBurst=10 (drop-in; base unit untouched) |
ops/systemd/duckbrain-http.service |
repo reference copy; must not clobber live ExecStart |
ops/systemd/duckbrain-http-health.service + .timer |
oneshot /health probe, 60 s interval, enabled + started live |
scripts/scoped-stop.js, src/cli/scoped-stop.ts |
pidfile-scoped SIGTERM with live/format/daemon/port validation |
scripts/health-check.js, src/cli/health-check.ts |
200 and 503 = alive; connection failure = dark |
| HTTP router | GET /health registered before auth middleware (auth-exempt) |
| daemon startup/shutdown | write/unlink port-keyed pidfile atomically |
| all stop scripts / CI | delete pkill -f 'duckbrain.*http' in favor of the scoped helper |
A tier-2 criterion demanding live systemd verification cannot be satisfied by repo-owned unit files. The foreman must actually:
systemctl --user daemon-reload,enable --now the health timer,kill -TERM $MainPID probe and confirm a new MainPID and incremented NRestarts,curl /health = 200 (or 503) and an authed GET = 200,before re-judging. Judge 8292fe3d returned FAIL because the unit was still uninstalled on the host; 3a5ff2a0 passed only after this install-and-probe sequence ran. Repo files are necessary but not sufficient — the live state is the artifact under test.
# Evidence - Problem class: systemd-on-failure-graceful-sigterm-dark-port - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-12T12:56:15.432Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a production daemon under a systemd user unit with Restart=on-failure received an out-of-band SIGTERM, logged its graceful shutdown, exited status=0/SUCCESS, and stayed dark for ~7 minutes; NRestarts never incremented and the journal showed no Stopping lines (proving it was not systemctl stop). Root cause: Restart=on-failure treats exit status 0 as success and never restarts, so ANY graceful SIGTERM/SIGINT (stray pkill pattern kill, judge/worker cleanup, OOM-adjacent TERM) permanently downs the service until manual intervention. A second latent hazard: the repo stop script was pkill -f 'duckbrain.*http', which matches the production daemon cmdline exactly, so any scratch cleanup takes production down. Fix (three parts): (1) durability \u2014 add a drop-in (e.g. 10-restart-always.conf) with Restart=always, RestartSec=3, StartLimitIntervalSec=300, StartLimitBurst=10, then daemon-reload; do NOT overwrite the base unit when drop-ins carry the real ExecStart (auth flags, unix socket). (2) blast radius \u2014 replace pkill stop scripts with a pidfile-scoped SIGTERM helper that validates the PID is live and its cmdline matches a daemon for the requested port, refusing stale/malformed/wrong-port pidfiles. (3) monitoring \u2014 oneshot health-check service + systemd timer probing /health every minute, treating HTTP 200 AND 503 as alive (503 degraded may be an intentional dependency-health contract) and connection failure as dark; /health must be auth-exempt so the probe needs no keys. Verification: kill -TERM <MainPID>, poll systemctl show -p MainPID,NRestarts until a new PID appears and NRestarts increments, then curl /health = 200 and an authed GET works. GitReins lesson: when a tier2 criterion demands 'live systemd verification', repo-owned unit files alone FAIL the judge \u2014 the foreman must install the unit/drop-in, daemon-reload, enable the timer, and run the live SIGTERM probe before re-judging.", "environment": "systemd user service, node daemon, DuckBrain HTTP :3000, Linux user lingering", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "systemd-on-failure-graceful-sigterm-dark-port", "provider": "openrouter", "solved_at": "2026-09-12T12:56:15.432Z", "version": ""}