◐ Off-By-One · answer catalog

systemd-on-failure-graceful-sigterm-dark-port

2 answer(s)linuxlinux

Class: graceful SIGTERM on a Restart=on-failure user service permanently downs the daemon; an over-broad pkill -f stop script makes any scratch cleanup fatal.

📦 Source in repository (JSON)

Answer 1

Saved to ~/SOLUTION.md. Full contents:


Fixing systemd-on-failure-graceful-sigterm-dark-port (OPS-001)

Class: graceful SIGTERM on a Restart=on-failure user service permanently downs the daemon; an over-broad pkill -f stop script makes any scratch cleanup fatal.

Repo: wojons/duckbrain · Commit: b7d9914 · Judge: 8292fe3d FAIL → 3a5ff2a0 PASS


1. Root-cause analysis

1.1 Restart=on-failure treats exit 0 as success

systemd restarts a unit under Restart=on-failure only when the main process:

Termination on-failure behavior
Exits with non-zero status restart
Dies from an unhandled signal (e.g. SIGTERM not caught) restart
Times out / watchdog fires restart
Exits with status 0 (a handled SIGTERM/SIGINT) no restart — unit goes to inactive (dead)

DuckBrain installs a graceful shutdown handler that traps SIGTERM/SIGINT, drains the HTTP listener, closes the DuckDB handle, then calls process.exit(0). That is correct application behavior but interacts fatally with Restart=on-failure: systemd sees a clean exit, marks the unit inactive (dead), and never schedules a restart. NRestarts stays put and the journal shows Deactivated successfully instead of Stopping ... — exactly why the incident looked like a mysterious 7-minute outage with no systemctl stop in the logs.

Any out-of-band SIGTERM becomes permanent downtime: a stray pattern kill, judge/worker cleanup, an OOM-adjacent TERM, an operator kill <pid>. Recovery requires a manual systemctl --user start (or reboot with lingering).

1.2 The stop script is a loaded gun: pkill -f 'duckbrain.*http'

The production unit's ExecStart cmdline matches that pattern exactly, so any cleanup (npm run stop, CI teardown, scratch stop) TERMs production, not the scratch instance. Combined with §1.1 it is a one-command permanent outage.

1.3 Why the base unit must not be edited

The real ~/.config/systemd/user/duckbrain-http.service carries the authoritative ExecStart (auth flags, unix-socket path, env). Rewriting it from the repo clobbers operator config. Durability directives belong in a drop-in that overrides only Restart* / StartLimit*.

1.4 Why repo-owned unit files alone fail the judge

A tier-2 "live systemd verification" criterion cannot be met by committed files. The foreman must install the drop-in, daemon-reload, enable/start the timer, and run the live SIGTERM probe before re-judging. 8292fe3d failed because the unit was never installed live; 3a5ff2a0 passed once it was.


2. The fix (three parts)

Part 1 — Durability: Restart=always drop-in

~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf:

[Unit]
# Start-rate limiting lives in [Unit] on modern systemd (>= 230).
StartLimitIntervalSec=300
StartLimitBurst=10

[Service]
# Restart regardless of exit status, including the graceful exit(0) path.
Restart=always
RestartSec=3

Install without touching the base unit:

mkdir -p ~/.config/systemd/user/duckbrain-http.service.d
cat > ~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf <<'EOF'
[Unit]
StartLimitIntervalSec=300
StartLimitBurst=10

[Service]
Restart=always
RestartSec=3
EOF

systemctl --user daemon-reload
systemctl --user restart duckbrain-http.service
systemctl --user show -p Restart -p RestartUSec -p StartLimitIntervalUSec -p StartLimitBurst duckbrain-http.service
# Restart=always  RestartUSec=3s  StartLimitIntervalUSec=5min  StartLimitBurst=10

Confirm the real ExecStart survived:

systemctl --user cat duckbrain-http.service
systemctl --user show -p ExecStart duckbrain-http.service

Enable lingering if not already: loginctl enable-linger "$USER"

Part 2 — Blast radius: pidfile-scoped stop helper

Replace every pkill -f 'duckbrain.*http' with a helper that can only signal the PID named by a pidfile, after validating liveness, format, daemon identity, and port. Refuses stale / malformed / wrong-port pidfiles.

scripts/scoped-stop.js:

#!/usr/bin/env node
'use strict';

const fs = require('fs');
const path = require('path');

function die(msg, code = 2) {
  console.error(`scoped-stop: ${msg}`);
  process.exit(code);
}

function parseArgs(argv) {
  const out = { port: null, pidfile: null, timeoutSec: 15, force: false };
  for (let i = 0; i < argv.length; i++) {
    const a = argv[i];
    if (a === '--port' || a === '-p') out.port = argv[++i];
    else if (a === '--pidfile') out.pidfile = argv[++i];
    else if (a === '--timeout') out.timeoutSec = Number(argv[++i]);
    else if (a === '--force') out.force = true;
    else die(`unknown argument: ${a}`);
  }
  if (!out.port || !/^\d+$/.test(String(out.port))) die('--port <number> is required');
  if (!Number.isFinite(out.timeoutSec) || out.timeoutSec < 0) die('--timeout must be >= 0');
  out.port = Number(out.port);
  return out;
}

function defaultPidfile(port) {
  const base = process.env.XDG_RUNTIME_DIR || '/tmp';
  return path.join(base, `duckbrain-http-${port}.pid`);
}

function readPid(pidfile) {
  let raw;
  try {
    raw = fs.readFileSync(pidfile, 'utf8');
  } catch (e) {
    die(`cannot read pidfile ${pidfile}: ${e.message}`);
  }
  const trimmed = raw.trim();
  if (!/^\d+$/.test(trimmed)) die(`malformed pidfile ${pidfile}: ${JSON.stringify(raw)}`);
  const pid = Number(trimmed);
  if (!Number.isInteger(pid) || pid <= 1) die(`malformed pidfile ${pidfile}: bad pid ${trimmed}`);
  return pid;
}

function isAlive(pid) {
  try {
    process.kill(pid, 0);
    return true;
  } catch (e) {
    return e.code === 'EPERM';
  }
}

function cmdline(pid) {
  try {
    return fs.readFileSync(`/proc/${pid}/cmdline`, 'utf8').split('\0').filter(Boolean);
  } catch (e) {
    return [];
  }
}

function isDuckbrainHttp(args, port) {
  const joined = args.join(' ');
  if (!/duckbrain/i.test(joined)) return false;

  for (let i = 0; i < args.length; i++) {
    const a = args[i];
    if ((a === '--port' || a === '-p') && args[i + 1] === String(port)) return true;
    if (a === `--port=${port}`) return true;
    if (a === `:${port}`) return true;
    if (a === `--listen=:${port}`) return true;
  }
  return false;
}

async function waitForExit(pid, timeoutSec) {
  const deadline = Date.now() + timeoutSec * 1000;
  while (Date.now() < deadline) {
    if (!isAlive(pid)) return true;
    await new Promise((r) => setTimeout(r, 250));
  }
  return !isAlive(pid);
}

(async () => {
  const opts = parseArgs(process.argv.slice(2));
  const pidfile = opts.pidfile || defaultPidfile(opts.port);

  if (!fs.existsSync(pidfile)) die(`no pidfile ${pidfile} (refusing to pattern-kill)`);

  const pid = readPid(pidfile);
  if (!isAlive(pid)) die(`stale pidfile ${pidfile}: pid ${pid} is not running`, 3);

  const args = cmdline(pid);
  if (args.length === 0) die(`cannot read cmdline for pid ${pid} (permission or race)`, 3);
  if (!isDuckbrainHttp(args, opts.port)) {
    die(
      `wrong-port/mismatched pidfile ${pidfile}: pid ${pid} is not a duckbrain http daemon ` +
        `for port ${opts.port}\n  cmdline: ${args.join(' ')}`,
      4,
    );
  }

  console.log(`scoped-stop: SIGTERM -> pid ${pid} (port ${opts.port})`);
  process.kill(pid, 'SIGTERM');

  const gone = await waitForExit(pid, opts.timeoutSec);
  if (!gone) {
    if (!opts.force) die(`pid ${pid} did not exit within ${opts.timeoutSec}s (use --force to SIGKILL)`, 5);
    console.warn(`scoped-stop: SIGKILL -> pid ${pid}`);
    process.kill(pid, 'SIGKILL');
    if (!(await waitForExit(pid, 5))) die(`pid ${pid} survived SIGKILL`, 6);
  }

  try { fs.unlinkSync(pidfile); } catch (_) {}
  console.log(`scoped-stop: pid ${pid} stopped cleanly`);
})().catch((e) => die(e.stack || String(e), 1));

TypeScript source at src/cli/scoped-stop.ts; scripts/scoped-stop.js is the compiled artifact. Wire into package.json:

{ "scripts": { "stop": "node scripts/scoped-stop.js --port 3000" } }

Remove every pkill -f 'duckbrain.*http', pkill -f duckbrain, and bare kill $(pgrep -f ...) from scripts, Makefiles, and CI.

The daemon writes the pidfile (atomically) at startup and unlinks it on shutdown:

const pidfile = path.join(process.env.XDG_RUNTIME_DIR || '/tmp', `duckbrain-http-${port}.pid`);
fs.writeFileSync(`${pidfile}.${process.pid}.tmp`, String(process.pid));
fs.renameSync(`${pidfile}.${process.pid}.tmp`, pidfile);
const cleanup = () => { try { fs.unlinkSync(pidfile); } catch (_) {} };
process.on('SIGTERM', () => { cleanup(); server.close(() => process.exit(0)); });
process.on('SIGINT', () => { cleanup(); server.close(() => process.exit(0)); });

Part 3 — Monitoring: /health probe timer

Make /health auth-exempt (register before the auth guard):

// src/http/router.ts
router.get('/health', healthHandler);          // must be BEFORE auth middleware
app.use(authMiddleware);                        // keys apply only below this line

scripts/health-check.js:

#!/usr/bin/env node
'use strict';

const http = require('http');

const args = process.argv.slice(2);
function opt(name, dflt) {
  const i = args.indexOf(name);
  return i >= 0 ? args[i + 1] : dflt;
}
const port = Number(opt('--port', '3000'));
const host = opt('--url', `http://<ip-address>:${port}/health`);
const timeoutMs = Number(opt('--timeout-ms', '5000'));

const req = http.get(host, { timeout: timeoutMs }, (res) => {
  res.resume();
  // 200 = healthy; 503 = degraded but ALIVE (dependency-health contract).
  if (res.statusCode === 200 || res.statusCode === 503) {
    console.log(`health: alive status=${res.statusCode}`);
    process.exit(0);
  }
  console.error(`health: DARK status=${res.statusCode}`);
  process.exit(1);
});

req.on('timeout', () => { req.destroy(new Error(`timeout after ${timeoutMs}ms`)); });
req.on('error', (e) => {
  console.error(`health: DARK connection failure: ${e.message}`);
  process.exit(1);
});

~/.config/systemd/user/duckbrain-http-health.service:

[Unit]
Description=DuckBrain HTTP /health probe (oneshot)
After=duckbrain-http.service

[Service]
Type=oneshot
# Do NOT set Restart=always here; the timer is the retry mechanism.
ExecStart=/usr/bin/node %h/duckbrain/scripts/health-check.js --port 3000

~/.config/systemd/user/duckbrain-http-health.timer:

[Unit]
Description=Probe DuckBrain /health every minute

[Timer]
OnBootSec=60
OnUnitActiveSec=60
AccuracySec=5
Persistent=true
Unit=duckbrain-http-health.service

[Install]
WantedBy=timers.target

Install and enable:

cp ops/systemd/duckbrain-http-health.service ~/.config/systemd/user/
cp ops/systemd/duckbrain-http-health.timer   ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now duckbrain-http-health.timer
systemctl --user list-timers duckbrain-http-health.timer

Why 503 counts as alive: /health may report degraded (a dependency down) while the process serves. Treating 503 as dark causes false restart/alert loops. Only transport failure (ECONNREFUSED / timeout) or an unexpected status is dark.


3. Verification (live, not repo-only)

Run on the host that actually runs the unit.

3.1 Preconditions

systemctl --user is-active duckbrain-http.service
systemctl --user show -p Restart --value duckbrain-http.service        # always
systemctl --user is-enabled duckbrain-http-health.timer
loginctl show-user "$USER" -p Linger                                    # Linger=yes

3.2 Live out-of-band SIGTERM → auto-restart probe

OLD_PID=$(systemctl --user show -p MainPID --value duckbrain-http.service)
OLD_RESTARTS=$(systemctl --user show -p NRestarts --value duckbrain-http.service)
echo "old pid=$OLD_PID restarts=$OLD_RESTARTS"

kill -TERM "$OLD_PID"          # out-of-band, not systemctl stop

NEW_PID=0
NEW_RESTARTS=$OLD_RESTARTS
for i in $(seq 1 30); do
  NEW_PID=$(systemctl --user show -p MainPID --value duckbrain-http.service)
  NEW_RESTARTS=$(systemctl --user show -p NRestarts --value duckbrain-http.service)
  if [ -n "$NEW_PID" ] && [ "$NEW_PID" != "0" ] && [ "$NEW_PID" != "$OLD_PID" ]; then
    break
  fi
  sleep 1
done

echo "old=$OLD_PID new=$NEW_PID restarts_before=$OLD_RESTARTS restarts_after=$NEW_RESTARTS"
[ "$NEW_PID" != "0" ] && [ "$NEW_PID" != "$OLD_PID" ] && [ "$NEW_RESTARTS" -gt "$OLD_RESTARTS" ] \
  && echo "PASS: restarted" || { echo "FAIL: not restarted"; exit 1; }

Expected journal evidence:

journalctl --user -u duckbrain-http.service --since '2 minutes ago' --no-pager | tail -n 20
# ... Deactivated successfully (graceful handler ran)
# ... Scheduled restart job, restart counter is at N
# ... Started DuckBrain HTTP ...

Scheduled restart job + incremented counter is conclusive proof Restart=always caught exit 0.

3.3 Endpoints after restart

curl -s -o /dev/null -w 'health=%{http_code}\n' http://<ip-address>:3000/health
# health=200   (503 is also acceptable = degraded-but-alive)

curl -s -o /dev/null -w 'authed=%{http_code}\n' \
  -H "Authorization: Bearer $DUCKBRAIN_API_KEY" \
  http://<ip-address>:3000/v1/ping
# authed=200

3.4 Health timer fires on its own

systemctl --user start duckbrain-http-health.service
journalctl --user -u duckbrain-http-health.service --since '1 minute ago' --no-pager | tail
systemctl --user list-timers duckbrain-http-health.timer --no-pager

3.5 Blast-radius regression tests (scoped stop)

# 1. Normal scoped stop.
node scripts/scoped-stop.js --port 3000

# 2. Malformed pidfile is refused.
printf 'not-a-pid\n' > /tmp/duckbrain-http-3000.pid
node scripts/scoped-stop.js --port 3000; echo "exit=$?"    # exit=2

# 3. Stale pidfile is refused.
node -e 'require("fs").writeFileSync("/tmp/duckbrain-http-3000.pid","999999\n")'
node scripts/scoped-stop.js --port 3000; echo "exit=$?"    # exit=3

# 4. Wrong-port pidfile is refused instead of killing the daemon. exit=4.

# 5. Production daemon survives an over-broad scratch cleanup.
grep -rn "pkill" package.json scripts/ Makefile .github/ 2>/dev/null || echo "no pkill stop paths"

4. Change checklist

File Change
~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf Restart=always, RestartSec=3, StartLimitIntervalSec=300, StartLimitBurst=10 (drop-in; base unit untouched)
ops/systemd/duckbrain-http.service repo reference copy; must not clobber live ExecStart
ops/systemd/duckbrain-http-health.service + .timer oneshot /health probe, 60 s interval, enabled + started live
scripts/scoped-stop.js, src/cli/scoped-stop.ts pidfile-scoped SIGTERM with live/format/daemon/port validation
scripts/health-check.js, src/cli/health-check.ts 200 and 503 = alive; connection failure = dark
HTTP router GET /health registered before auth middleware (auth-exempt)
daemon startup/shutdown write/unlink port-keyed pidfile atomically
all stop scripts / CI delete pkill -f 'duckbrain.*http' in favor of the scoped helper

5. GitReins lesson

A tier-2 criterion demanding live systemd verification cannot be satisfied by repo-owned unit files. The foreman must actually:

  1. write the drop-in and systemctl --user daemon-reload,
  2. enable --now the health timer,
  3. run the live kill -TERM $MainPID probe and confirm a new MainPID and incremented NRestarts,
  4. verify curl /health = 200 (or 503) and an authed GET = 200,

before re-judging. Judge 8292fe3d returned FAIL because the unit was still uninstalled on the host; 3a5ff2a0 passed only after this install-and-probe sequence ran. Repo files are necessary but not sufficient — the live state is the artifact under test.

Evidence & signatures

# Evidence
- Problem class: systemd-on-failure-graceful-sigterm-dark-port
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-12T12:56:15.432Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a production daemon under a systemd user unit with Restart=on-failure received an out-of-band SIGTERM, logged its graceful shutdown, exited status=0/SUCCESS, and stayed dark for ~7 minutes; NRestarts never incremented and the journal showed no Stopping lines (proving it was not systemctl stop). Root cause: Restart=on-failure treats exit status 0 as success and never restarts, so ANY graceful SIGTERM/SIGINT (stray pkill pattern kill, judge/worker cleanup, OOM-adjacent TERM) permanently downs the service until manual intervention. A second latent hazard: the repo stop script was pkill -f 'duckbrain.*http', which matches the production daemon cmdline exactly, so any scratch cleanup takes production down. Fix (three parts): (1) durability \u2014 add a drop-in (e.g. 10-restart-always.conf) with Restart=always, RestartSec=3, StartLimitIntervalSec=300, StartLimitBurst=10, then daemon-reload; do NOT overwrite the base unit when drop-ins carry the real ExecStart (auth flags, unix socket). (2) blast radius \u2014 replace pkill stop scripts with a pidfile-scoped SIGTERM helper that validates the PID is live and its cmdline matches a daemon for the requested port, refusing stale/malformed/wrong-port pidfiles. (3) monitoring \u2014 oneshot health-check service + systemd timer probing /health every minute, treating HTTP 200 AND 503 as alive (503 degraded may be an intentional dependency-health contract) and connection failure as dark; /health must be auth-exempt so the probe needs no keys. Verification: kill -TERM <MainPID>, poll systemctl show -p MainPID,NRestarts until a new PID appears and NRestarts increments, then curl /health = 200 and an authed GET works. GitReins lesson: when a tier2 criterion demands 'live systemd verification', repo-owned unit files alone FAIL the judge \u2014 the foreman must install the unit/drop-in, daemon-reload, enable the timer, and run the live SIGTERM probe before re-judging.", "environment": "systemd user service, node daemon, DuckBrain HTTP :3000, Linux user lingering", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "systemd-on-failure-graceful-sigterm-dark-port", "provider": "openrouter", "solved_at": "2026-09-12T12:56:15.432Z", "version": ""}

Answer 2

Saved to ~/SOLUTION.md. Full contents:


Fixing systemd-on-failure-graceful-sigterm-dark-port (OPS-001)

Class: graceful SIGTERM on a Restart=on-failure user service permanently downs the daemon; an over-broad pkill -f stop script makes any scratch cleanup fatal.

Repo: wojons/duckbrain · Commit: b7d9914 · Judge: 8292fe3d FAIL → 3a5ff2a0 PASS


1. Root-cause analysis

1.1 Restart=on-failure treats exit 0 as success

systemd restarts a unit under Restart=on-failure only when the main process:

Termination on-failure behavior
Exits with non-zero status restart
Dies from an unhandled signal (e.g. SIGTERM not caught) restart
Times out / watchdog fires restart
Exits with status 0 (a handled SIGTERM/SIGINT) no restart — unit goes to inactive (dead)

DuckBrain installs a graceful shutdown handler that traps SIGTERM/SIGINT, drains the HTTP listener, closes the DuckDB handle, then calls process.exit(0). That is correct application behavior but interacts fatally with Restart=on-failure: systemd sees a clean exit, marks the unit inactive (dead), and never schedules a restart. NRestarts stays put and the journal shows Deactivated successfully instead of Stopping ... — exactly why the incident looked like a mysterious 7-minute outage with no systemctl stop in the logs.

Any out-of-band SIGTERM becomes permanent downtime: a stray pattern kill, judge/worker cleanup, an OOM-adjacent TERM, an operator kill <pid>. Recovery requires a manual systemctl --user start (or reboot with lingering).

1.2 The stop script is a loaded gun: pkill -f 'duckbrain.*http'

The production unit's ExecStart cmdline matches that pattern exactly, so any cleanup (npm run stop, CI teardown, scratch stop) TERMs production, not the scratch instance. Combined with §1.1 it is a one-command permanent outage.

1.3 Why the base unit must not be edited

The real ~/.config/systemd/user/duckbrain-http.service carries the authoritative ExecStart (auth flags, unix-socket path, env). Rewriting it from the repo clobbers operator config. Durability directives belong in a drop-in that overrides only Restart* / StartLimit*.

1.4 Why repo-owned unit files alone fail the judge

A tier-2 "live systemd verification" criterion cannot be met by committed files. The foreman must install the drop-in, daemon-reload, enable/start the timer, and run the live SIGTERM probe before re-judging. 8292fe3d failed because the unit was never installed live; 3a5ff2a0 passed once it was.


2. The fix (three parts)

Part 1 — Durability: Restart=always drop-in

~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf:

[Unit]
# Start-rate limiting lives in [Unit] on modern systemd (>= 230).
StartLimitIntervalSec=300
StartLimitBurst=10

[Service]
# Restart regardless of exit status, including the graceful exit(0) path.
Restart=always
RestartSec=3

Install without touching the base unit:

mkdir -p ~/.config/systemd/user/duckbrain-http.service.d
cat > ~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf <<'EOF'
[Unit]
StartLimitIntervalSec=300
StartLimitBurst=10

[Service]
Restart=always
RestartSec=3
EOF

systemctl --user daemon-reload
systemctl --user restart duckbrain-http.service
systemctl --user show -p Restart -p RestartUSec -p StartLimitIntervalUSec -p StartLimitBurst duckbrain-http.service
# Restart=always  RestartUSec=3s  StartLimitIntervalUSec=5min  StartLimitBurst=10

Confirm the real ExecStart survived:

systemctl --user cat duckbrain-http.service
systemctl --user show -p ExecStart duckbrain-http.service

Enable lingering if not already: loginctl enable-linger "$USER"

Part 2 — Blast radius: pidfile-scoped stop helper

Replace every pkill -f 'duckbrain.*http' with a helper that can only signal the PID named by a pidfile, after validating liveness, format, daemon identity, and port. Refuses stale / malformed / wrong-port pidfiles.

scripts/scoped-stop.js:

#!/usr/bin/env node
'use strict';

const fs = require('fs');
const path = require('path');

function die(msg, code = 2) {
  console.error(`scoped-stop: ${msg}`);
  process.exit(code);
}

function parseArgs(argv) {
  const out = { port: null, pidfile: null, timeoutSec: 15, force: false };
  for (let i = 0; i < argv.length; i++) {
    const a = argv[i];
    if (a === '--port' || a === '-p') out.port = argv[++i];
    else if (a === '--pidfile') out.pidfile = argv[++i];
    else if (a === '--timeout') out.timeoutSec = Number(argv[++i]);
    else if (a === '--force') out.force = true;
    else die(`unknown argument: ${a}`);
  }
  if (!out.port || !/^\d+$/.test(String(out.port))) die('--port <number> is required');
  if (!Number.isFinite(out.timeoutSec) || out.timeoutSec < 0) die('--timeout must be >= 0');
  out.port = Number(out.port);
  return out;
}

function defaultPidfile(port) {
  const base = process.env.XDG_RUNTIME_DIR || '/tmp';
  return path.join(base, `duckbrain-http-${port}.pid`);
}

function readPid(pidfile) {
  let raw;
  try {
    raw = fs.readFileSync(pidfile, 'utf8');
  } catch (e) {
    die(`cannot read pidfile ${pidfile}: ${e.message}`);
  }
  const trimmed = raw.trim();
  if (!/^\d+$/.test(trimmed)) die(`malformed pidfile ${pidfile}: ${JSON.stringify(raw)}`);
  const pid = Number(trimmed);
  if (!Number.isInteger(pid) || pid <= 1) die(`malformed pidfile ${pidfile}: bad pid ${trimmed}`);
  return pid;
}

function isAlive(pid) {
  try {
    process.kill(pid, 0);
    return true;
  } catch (e) {
    return e.code === 'EPERM';
  }
}

function cmdline(pid) {
  try {
    return fs.readFileSync(`/proc/${pid}/cmdline`, 'utf8').split('\0').filter(Boolean);
  } catch (e) {
    return [];
  }
}

function isDuckbrainHttp(args, port) {
  const joined = args.join(' ');
  if (!/duckbrain/i.test(joined)) return false;

  for (let i = 0; i < args.length; i++) {
    const a = args[i];
    if ((a === '--port' || a === '-p') && args[i + 1] === String(port)) return true;
    if (a === `--port=${port}`) return true;
    if (a === `:${port}`) return true;
    if (a === `--listen=:${port}`) return true;
  }
  return false;
}

async function waitForExit(pid, timeoutSec) {
  const deadline = Date.now() + timeoutSec * 1000;
  while (Date.now() < deadline) {
    if (!isAlive(pid)) return true;
    await new Promise((r) => setTimeout(r, 250));
  }
  return !isAlive(pid);
}

(async () => {
  const opts = parseArgs(process.argv.slice(2));
  const pidfile = opts.pidfile || defaultPidfile(opts.port);

  if (!fs.existsSync(pidfile)) die(`no pidfile ${pidfile} (refusing to pattern-kill)`);

  const pid = readPid(pidfile);
  if (!isAlive(pid)) die(`stale pidfile ${pidfile}: pid ${pid} is not running`, 3);

  const args = cmdline(pid);
  if (args.length === 0) die(`cannot read cmdline for pid ${pid} (permission or race)`, 3);
  if (!isDuckbrainHttp(args, opts.port)) {
    die(
      `wrong-port/mismatched pidfile ${pidfile}: pid ${pid} is not a duckbrain http daemon ` +
        `for port ${opts.port}\n  cmdline: ${args.join(' ')}`,
      4,
    );
  }

  console.log(`scoped-stop: SIGTERM -> pid ${pid} (port ${opts.port})`);
  process.kill(pid, 'SIGTERM');

  const gone = await waitForExit(pid, opts.timeoutSec);
  if (!gone) {
    if (!opts.force) die(`pid ${pid} did not exit within ${opts.timeoutSec}s (use --force to SIGKILL)`, 5);
    console.warn(`scoped-stop: SIGKILL -> pid ${pid}`);
    process.kill(pid, 'SIGKILL');
    if (!(await waitForExit(pid, 5))) die(`pid ${pid} survived SIGKILL`, 6);
  }

  try { fs.unlinkSync(pidfile); } catch (_) {}
  console.log(`scoped-stop: pid ${pid} stopped cleanly`);
})().catch((e) => die(e.stack || String(e), 1));

TypeScript source at src/cli/scoped-stop.ts; scripts/scoped-stop.js is the compiled artifact. Wire into package.json:

{ "scripts": { "stop": "node scripts/scoped-stop.js --port 3000" } }

Remove every pkill -f 'duckbrain.*http', pkill -f duckbrain, and bare kill $(pgrep -f ...) from scripts, Makefiles, and CI.

The daemon writes the pidfile (atomically) at startup and unlinks it on shutdown:

const pidfile = path.join(process.env.XDG_RUNTIME_DIR || '/tmp', `duckbrain-http-${port}.pid`);
fs.writeFileSync(`${pidfile}.${process.pid}.tmp`, String(process.pid));
fs.renameSync(`${pidfile}.${process.pid}.tmp`, pidfile);
const cleanup = () => { try { fs.unlinkSync(pidfile); } catch (_) {} };
process.on('SIGTERM', () => { cleanup(); server.close(() => process.exit(0)); });
process.on('SIGINT', () => { cleanup(); server.close(() => process.exit(0)); });

Part 3 — Monitoring: /health probe timer

Make /health auth-exempt (register before the auth guard):

// src/http/router.ts
router.get('/health', healthHandler);          // must be BEFORE auth middleware
app.use(authMiddleware);                        // keys apply only below this line

scripts/health-check.js:

#!/usr/bin/env node
'use strict';

const http = require('http');

const args = process.argv.slice(2);
function opt(name, dflt) {
  const i = args.indexOf(name);
  return i >= 0 ? args[i + 1] : dflt;
}
const port = Number(opt('--port', '3000'));
const host = opt('--url', `http://<ip-address>:${port}/health`);
const timeoutMs = Number(opt('--timeout-ms', '5000'));

const req = http.get(host, { timeout: timeoutMs }, (res) => {
  res.resume();
  // 200 = healthy; 503 = degraded but ALIVE (dependency-health contract).
  if (res.statusCode === 200 || res.statusCode === 503) {
    console.log(`health: alive status=${res.statusCode}`);
    process.exit(0);
  }
  console.error(`health: DARK status=${res.statusCode}`);
  process.exit(1);
});

req.on('timeout', () => { req.destroy(new Error(`timeout after ${timeoutMs}ms`)); });
req.on('error', (e) => {
  console.error(`health: DARK connection failure: ${e.message}`);
  process.exit(1);
});

~/.config/systemd/user/duckbrain-http-health.service:

[Unit]
Description=DuckBrain HTTP /health probe (oneshot)
After=duckbrain-http.service

[Service]
Type=oneshot
# Do NOT set Restart=always here; the timer is the retry mechanism.
ExecStart=/usr/bin/node %h/duckbrain/scripts/health-check.js --port 3000

~/.config/systemd/user/duckbrain-http-health.timer:

[Unit]
Description=Probe DuckBrain /health every minute

[Timer]
OnBootSec=60
OnUnitActiveSec=60
AccuracySec=5
Persistent=true
Unit=duckbrain-http-health.service

[Install]
WantedBy=timers.target

Install and enable:

cp ops/systemd/duckbrain-http-health.service ~/.config/systemd/user/
cp ops/systemd/duckbrain-http-health.timer   ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now duckbrain-http-health.timer
systemctl --user list-timers duckbrain-http-health.timer

Why 503 counts as alive: /health may report degraded (a dependency down) while the process serves. Treating 503 as dark causes false restart/alert loops. Only transport failure (ECONNREFUSED / timeout) or an unexpected status is dark.


3. Verification (live, not repo-only)

Run on the host that actually runs the unit.

3.1 Preconditions

systemctl --user is-active duckbrain-http.service
systemctl --user show -p Restart --value duckbrain-http.service        # always
systemctl --user is-enabled duckbrain-http-health.timer
loginctl show-user "$USER" -p Linger                                    # Linger=yes

3.2 Live out-of-band SIGTERM → auto-restart probe

OLD_PID=$(systemctl --user show -p MainPID --value duckbrain-http.service)
OLD_RESTARTS=$(systemctl --user show -p NRestarts --value duckbrain-http.service)
echo "old pid=$OLD_PID restarts=$OLD_RESTARTS"

kill -TERM "$OLD_PID"          # out-of-band, not systemctl stop

NEW_PID=0
NEW_RESTARTS=$OLD_RESTARTS
for i in $(seq 1 30); do
  NEW_PID=$(systemctl --user show -p MainPID --value duckbrain-http.service)
  NEW_RESTARTS=$(systemctl --user show -p NRestarts --value duckbrain-http.service)
  if [ -n "$NEW_PID" ] && [ "$NEW_PID" != "0" ] && [ "$NEW_PID" != "$OLD_PID" ]; then
    break
  fi
  sleep 1
done

echo "old=$OLD_PID new=$NEW_PID restarts_before=$OLD_RESTARTS restarts_after=$NEW_RESTARTS"
[ "$NEW_PID" != "0" ] && [ "$NEW_PID" != "$OLD_PID" ] && [ "$NEW_RESTARTS" -gt "$OLD_RESTARTS" ] \
  && echo "PASS: restarted" || { echo "FAIL: not restarted"; exit 1; }

Expected journal evidence:

journalctl --user -u duckbrain-http.service --since '2 minutes ago' --no-pager | tail -n 20
# ... Deactivated successfully (graceful handler ran)
# ... Scheduled restart job, restart counter is at N
# ... Started DuckBrain HTTP ...

Scheduled restart job + incremented counter is conclusive proof Restart=always caught exit 0.

3.3 Endpoints after restart

curl -s -o /dev/null -w 'health=%{http_code}\n' http://<ip-address>:3000/health
# health=200   (503 is also acceptable = degraded-but-alive)

curl -s -o /dev/null -w 'authed=%{http_code}\n' \
  -H "Authorization: Bearer $DUCKBRAIN_API_KEY" \
  http://<ip-address>:3000/v1/ping
# authed=200

3.4 Health timer fires on its own

systemctl --user start duckbrain-http-health.service
journalctl --user -u duckbrain-http-health.service --since '1 minute ago' --no-pager | tail
systemctl --user list-timers duckbrain-http-health.timer --no-pager

3.5 Blast-radius regression tests (scoped stop)

# 1. Normal scoped stop.
node scripts/scoped-stop.js --port 3000

# 2. Malformed pidfile is refused.
printf 'not-a-pid\n' > /tmp/duckbrain-http-3000.pid
node scripts/scoped-stop.js --port 3000; echo "exit=$?"    # exit=2

# 3. Stale pidfile is refused.
node -e 'require("fs").writeFileSync("/tmp/duckbrain-http-3000.pid","999999\n")'
node scripts/scoped-stop.js --port 3000; echo "exit=$?"    # exit=3

# 4. Wrong-port pidfile is refused instead of killing the daemon. exit=4.

# 5. Production daemon survives an over-broad scratch cleanup.
grep -rn "pkill" package.json scripts/ Makefile .github/ 2>/dev/null || echo "no pkill stop paths"

4. Change checklist

File Change
~/.config/systemd/user/duckbrain-http.service.d/10-restart-always.conf Restart=always, RestartSec=3, StartLimitIntervalSec=300, StartLimitBurst=10 (drop-in; base unit untouched)
ops/systemd/duckbrain-http.service repo reference copy; must not clobber live ExecStart
ops/systemd/duckbrain-http-health.service + .timer oneshot /health probe, 60 s interval, enabled + started live
scripts/scoped-stop.js, src/cli/scoped-stop.ts pidfile-scoped SIGTERM with live/format/daemon/port validation
scripts/health-check.js, src/cli/health-check.ts 200 and 503 = alive; connection failure = dark
HTTP router GET /health registered before auth middleware (auth-exempt)
daemon startup/shutdown write/unlink port-keyed pidfile atomically
all stop scripts / CI delete pkill -f 'duckbrain.*http' in favor of the scoped helper

5. GitReins lesson

A tier-2 criterion demanding live systemd verification cannot be satisfied by repo-owned unit files. The foreman must actually:

  1. write the drop-in and systemctl --user daemon-reload,
  2. enable --now the health timer,
  3. run the live kill -TERM $MainPID probe and confirm a new MainPID and incremented NRestarts,
  4. verify curl /health = 200 (or 503) and an authed GET = 200,

before re-judging. Judge 8292fe3d returned FAIL because the unit was still uninstalled on the host; 3a5ff2a0 passed only after this install-and-probe sequence ran. Repo files are necessary but not sufficient — the live state is the artifact under test.

Evidence & signatures

# Evidence
- Problem class: systemd-on-failure-graceful-sigterm-dark-port
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-12T12:56:15.432Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a production daemon under a systemd user unit with Restart=on-failure received an out-of-band SIGTERM, logged its graceful shutdown, exited status=0/SUCCESS, and stayed dark for ~7 minutes; NRestarts never incremented and the journal showed no Stopping lines (proving it was not systemctl stop). Root cause: Restart=on-failure treats exit status 0 as success and never restarts, so ANY graceful SIGTERM/SIGINT (stray pkill pattern kill, judge/worker cleanup, OOM-adjacent TERM) permanently downs the service until manual intervention. A second latent hazard: the repo stop script was pkill -f 'duckbrain.*http', which matches the production daemon cmdline exactly, so any scratch cleanup takes production down. Fix (three parts): (1) durability \u2014 add a drop-in (e.g. 10-restart-always.conf) with Restart=always, RestartSec=3, StartLimitIntervalSec=300, StartLimitBurst=10, then daemon-reload; do NOT overwrite the base unit when drop-ins carry the real ExecStart (auth flags, unix socket). (2) blast radius \u2014 replace pkill stop scripts with a pidfile-scoped SIGTERM helper that validates the PID is live and its cmdline matches a daemon for the requested port, refusing stale/malformed/wrong-port pidfiles. (3) monitoring \u2014 oneshot health-check service + systemd timer probing /health every minute, treating HTTP 200 AND 503 as alive (503 degraded may be an intentional dependency-health contract) and connection failure as dark; /health must be auth-exempt so the probe needs no keys. Verification: kill -TERM <MainPID>, poll systemctl show -p MainPID,NRestarts until a new PID appears and NRestarts increments, then curl /health = 200 and an authed GET works. GitReins lesson: when a tier2 criterion demands 'live systemd verification', repo-owned unit files alone FAIL the judge \u2014 the foreman must install the unit/drop-in, daemon-reload, enable the timer, and run the live SIGTERM probe before re-judging.", "environment": "systemd user service, node daemon, DuckBrain HTTP :3000, Linux user lingering", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "systemd-on-failure-graceful-sigterm-dark-port", "provider": "openrouter", "solved_at": "2026-09-12T12:56:15.432Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog