The crash loop (off-by-one.service, ~27k restarts / 37h) was invisible because no code path ever alerted on the restart counter:
Implemented and verified the fix. Artifacts written to /workspace/solution/:
- restart-watch.sh — drop-in detector
- test/test_restart_watch.sh — stubbed-systemctl harness (20/20 passing)
- SOLUTION.md — the write-up below
NRestarts crash loops (GAP-069 residual)The crash loop (off-by-one.service, ~27k restarts / 37h) was invisible because no code path ever alerted on the restart counter:
user-service-reset-monitor.sh snapshots NRestarts into /var/log/user-service-reset-monitor/periodic-snapshot.log, but a log line is not an alert. Nothing compares consecutive readings, so a counter climbing by thousands looks the same as one that never moves.NRestarts branch at all — no delta, no threshold, no event written to the JSONL board log.A crash loop is a rate problem, so the fix is delta-based (now vs. a persisted baseline), not absolute. Because the user manager resets NRestarts to 0 on restart, a counter decrease must be treated as a reset and silently re-baselined.
# .coding-hermes/board/.state/restart-baseline.json
{"off-by-one.service":{"nrestarts":27413,"updated_at":"2026-09-14T17:34:43-05:00"}}
# .gitignore
.coding-hermes/board/.state/
restart-watch.sh#!/usr/bin/env bash
#=============================================================================
# restart-watch.sh — NRestarts crash-loop detector for the staleness-checker
#
# Sourced by the staleness-checker (or its harness). Adds ONE side effect:
# when a watched systemd unit's restart counter jumps by >= threshold since
# the last recorded baseline, it appends an alert event to the JSONL board
# log through the alert writer. It ALWAYS returns 0 and never calls exit, so
# it cannot change the checker's exit status.
#
# Env knobs (all optional):
# RESTART_WATCH_UNIT unit to watch (default: off-by-one.service)
# RESTART_ALERT_THRESHOLD delta that triggers alert (default: 50)
# RESTART_BASELINE_FILE JSON baseline path (default: $BOARD_DIR/.state/restart-baseline.json)
# BOARD_DIR .coding-hermes/board dir (default: search from $PWD)
# RESTART_ALERT_TASK_ID board task the event links to (default: GAP-069)
# RESTART_ALERT_ACTOR event actor (default: staleness-checker)
# SYSTEMCTL systemctl binary override (default: systemctl)
# RESTART_WATCH_DISABLE=1 hard off switch
#=============================================================================
restart_watch_board_dir() {
if [[ -n "${BOARD_DIR:-}" ]]; then printf '%s\n' "$BOARD_DIR"; return 0; fi
local d
for d in "$PWD/.coding-hermes/board" "$PWD/.coding-hermes" "$PWD"; do
[[ -f "$d/events.jsonl" ]] && { printf '%s\n' "$d"; return 0; }
done
printf '%s\n' "$PWD/.coding-hermes/board"
}
# The alert-producing code path. If the repo already has `alert_writer`,
# the declare -F branch calls it; otherwise use boardctl, then a raw append.
restart_watch_alert() {
local unit="$1" before="$2" now="$3" delta="$4" threshold="$5"
local board detail
board="$(restart_watch_board_dir)"
detail=$(printf '{"kind":"service_restart_loop","unit":"%s","nrestarts_before":%s,"nrestarts_now":%s,"delta":%s,"threshold":%s,"detected_at":"%s"}' \
"$unit" "$before" "$now" "$delta" "$threshold" "$(date -Is 2>/dev/null || date)")
if declare -F alert_writer >/dev/null 2>&1; then
alert_writer "audit" "${RESTART_ALERT_TASK_ID:-GAP-069}" "$detail" || true
return 0
fi
if command -v boardctl >/dev/null 2>&1; then
if boardctl -C "$board" event \
--type audit \
--task-id "${RESTART_ALERT_TASK_ID:-GAP-069}" \
--actor "${RESTART_ALERT_ACTOR:-staleness-checker}" \
--detail-text "$detail" >/dev/null 2>&1; then
return 0
fi
fi
[[ -f "$board/events.jsonl" ]] || return 0
local next_id line
next_id=$( (grep -o '"id": *[0-9]\+' "$board/events.jsonl" 2>/dev/null | grep -o '[0-9]\+$' | sort -n | tail -1) )
next_id=$(( ${next_id:-0} + 1 ))
line=$(jq -cn \
--argjson id "$next_id" \
--arg ts "$(date -Is 2>/dev/null || date)" \
--arg tid "${RESTART_ALERT_TASK_ID:-GAP-069}" \
--arg actor "${RESTART_ALERT_ACTOR:-staleness-checker}" \
--arg detail "$detail" \
'{id:$id, timestamp:$ts, event_type:"audit", task_id:$tid, actor:$actor, detail:$detail, tick_number:null}' 2>/dev/null) || return 0
printf '%s\n' "$line" >>"$board/events.jsonl" 2>/dev/null || true
return 0
}
restart_watch_read_baseline() {
local file="$1" unit="$2"
[[ -f "$file" ]] || return 0
jq -r --arg u "$unit" '.[$u].nrestarts // empty' "$file" 2>/dev/null || true
}
restart_watch_write_baseline() {
local file="$1" unit="$2" value="$3"
local tmp
mkdir -p "$(dirname "$file")" 2>/dev/null || true
tmp=$(mktemp "${file}.XXXXXX" 2>/dev/null) || return 0
if [[ -f "$file" ]]; then
jq --arg u "$unit" --argjson n "$value" --arg ts "$(date -Is 2>/dev/null || date)" \
'.[$u] = {"nrestarts": $n, "updated_at": $ts}' \
"$file" >"$tmp" 2>/dev/null || { rm -f "$tmp"; return 0; }
else
jq -n --arg u "$unit" --argjson n "$value" --arg ts "$(date -Is 2>/dev/null || date)" \
'{($u): {"nrestarts": $n, "updated_at": $ts}}' \
>"$tmp" 2>/dev/null || { rm -f "$tmp"; return 0; }
fi
mv -f "$tmp" "$file" 2>/dev/null || rm -f "$tmp" 2>/dev/null
return 0
}
# Main entry point. Always returns 0.
restart_watch_check() {
[[ "${RESTART_WATCH_DISABLE:-0}" == "1" ]] && return 0
local unit="${RESTART_WATCH_UNIT:-off-by-one.service}"
local threshold="${RESTART_ALERT_THRESHOLD:-50}"
local baseline_file="${RESTART_BASELINE_FILE:-$(restart_watch_board_dir)/.state/restart-baseline.json}"
[[ "$threshold" =~ ^[0-9]+$ ]] || threshold=50
# 1) Query the live counter. Any failure/oddity is a silent no-op.
local now
now=$("${SYSTEMCTL:-systemctl}" show "$unit" -p NRestarts --value 2>/dev/null) || return 0
[[ "$now" =~ ^[0-9]+$ ]] || return 0
# 2) Compare against the baseline.
local before
before="$(restart_watch_read_baseline "$baseline_file" "$unit")"
# First run (or unreadable baseline): record it, do NOT alert.
if [[ -z "$before" ]]; then
restart_watch_write_baseline "$baseline_file" "$unit" "$now"
return 0
fi
# Counter went backwards => unit/manager restarted and NRestarts reset.
if (( now < before )); then
restart_watch_write_baseline "$baseline_file" "$unit" "$now"
return 0
fi
local delta=$(( now - before ))
if (( delta >= threshold )); then
# 3) ALERT: append to the JSONL board log, then re-arm the baseline.
restart_watch_alert "$unit" "$before" "$now" "$delta" "$threshold"
restart_watch_write_baseline "$baseline_file" "$unit" "$now"
fi
# Below threshold: leave baseline untouched so delta accumulates.
return 0
}
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
restart_watch_check
exit 0
fi
source "$(dirname "$0")/restart-watch.sh"
main() {
# ... existing staleness checks, which set $rc / return codes ...
rc=$?
# Existing alert-writer function name, if the checker already defines one.
# restart-watch.sh auto-detects `alert_writer`; boardctl is the fallback.
# Side-effect only: guaranteed return 0, feeds the JSONL board log.
restart_watch_check || true
return "$rc"
}
main "$@"
exit $?
Under set -euo pipefail this is still safe: every external call is guarded (|| return 0, || true) and the function ends return 0. Sub-threshold increments are not reset, so they accumulate toward the next alert.
RESTART_ALERT_THRESHOLD=200 systemctl --user restart staleness-checker.service
Harness with a stubbed systemctl that returns a scripted sequence of NRestarts values (FAIL → non-zero exit, empty line → no output). Actual result: 20 passed, 0 failed, covering first-run baseline, cumulative delta crossing, exactly-once alert + re-arm, silent reset, systemctl failure, exit-code preservation, and the disable switch.
== 1. first run records baseline, no alert ==
ok no event on first run
ok baseline recorded
== 2. delta >= threshold alerts exactly once and re-arms ==
ok run1 no alert
ok run1 baseline
ok run2 no alert
ok run2 baseline stays (cumulative)
ok run3 one alert
ok run3 re-arms baseline
ok run4 no duplicate alert
ok alert payload kind
ok alert delta
ok alert event_type
== 3. counter reset re-baselines silently ==
ok reset makes no event
ok reset re-baselines
== 4. systemctl failure / bad value is a silent no-op ==
ok FAIL leaves baseline
ok FAIL makes no event
ok empty leaves baseline
== 5. never alters exit codes ==
ok caller exit code preserved
ok function returns 0
== 6. disable switch ==
ok disabled makes no event
RESULT: 20 passed, 0 failed
End-to-end against a real day-one board, the alert is a genuine event (not a context field), and boardctl validate / doctor both return OK:
board: /tmp/bt2/.coding-hermes/board (topology A)
rows: 1 tasks, 1 events, 1 fixtures, header parsed
RESULT: OK (0 warning(s))
{"id": 1, "timestamp": "...", "event_type": "audit", "task_id": "GAP-069",
"actor": "staleness-checker",
"detail": "{\"kind\":\"service_restart_loop\",\"unit\":\"off-by-one.service\",\"nrestarts_before\":300,\"nrestarts_now\":400,\"delta\":100,\"threshold\":50,...}",
"tick_number": null}
| Criterion | Where it is satisfied |
|---|---|
Per-run NRestarts vs git-ignored JSON baseline |
restart_watch_check + .state/ in .gitignore |
delta >= threshold (env override, default 50) |
RESTART_ALERT_THRESHOLD, delta >= threshold branch |
| Appends event via alert writer | restart_watch_alert → alert_writer / boardctl event / raw JSONL append |
| First run records baseline, no alert | [[ -z "$before" ]] branch |
| Counter reset = silent re-baseline | (( now < before )) branch |
| Never alters exit codes | every path return 0; call site saves/restores rc |
Tested by stubbing systemctl |
test_restart_watch.sh (20/20) |
| Alert-shaped AC has an alert-producing path | restart_watch_alert emits an audit event; verified on the board |
# Evidence - Problem class: go-shell-crash-loop-restart-counter-alert - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-14T22:36:31.984Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "GAP-069 residual: systemd user service crash-loop (27k restarts/37h) invisible because nothing watched NRestarts. Fix: staleness-checker compares unit NRestarts against a git-ignored JSON baseline file per run; delta >= threshold (env-overridable, default 50) appends an event to the JSONL board log via the existing alert-writer function; first run records baseline, counter reset after restart = silent re-baseline; must never alter exit codes. Test by stubbing systemctl in the existing bash harness. Watch: tier-2 judge flagged the counter being RECORDED-not-ALERTED as a FAIL even when the other criteria passed \u2014 an alert-shaped AC needs an alert-producing code path, not context fields.", "environment": "linux-systemd-user", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-shell-crash-loop-restart-counter-alert", "provider": "openrouter", "solved_at": "2026-09-14T22:36:31.984Z", "version": ""}