Class: a scheduler tick is aborted by graceful shutdown with the message aborted by graceful shutdown - drain timed out with tick in flight.
Verified the recovery helper on sample data: it closes the target row, strips the 3 transient keys (34→31 canonical), and leaves peer rows byte-identical. Full solution written to ~/scheduler-drain-timeout-tick-crash.md.
scheduler-drain-timeout-tick-crashClass: a scheduler tick is aborted by graceful shutdown with the message aborted by graceful shutdown - drain timed out with tick in flight.
Distinct from: an agent iteration-cap death. Here the worker may have already finished: the fix commit is merged to the execution repo main and judged (tier1 + tier2 PASS) before the crash. Only the bookkeeping was lost: board row flip, event append, header bump, chore commit.
Golden rule: the work landed. Never respawn the worker. Adopt the landed artifact and re-derive the bookkeeping.
The scheduler drains in-flight ticks under a hard deadline. A tick slower than the deadline (typically the merge/judge tail) is killed with drain timed out with tick in flight. The tick's side effects are not atomic and are ordered wrong relative to the deadline:
tick: claim row -> run agent -> commit -> push branch -> merge to main -> judge (PASS)
^ work is durable here
-> flip board row -> append event -> bump header -> chore commit
^ death window: all of this can be lost
The crash leaves a split brain: execution repo ahead, board repo stale. A secondary contributor is non-idempotent/dirty bookkeeping: the crashed tick wrote a row with 34 keys while canonical rows have 31 (task_keys_uniform=drift), left uncommitted.
Three FAILs that all self-heal with the closeout:
| Canary signal | Meaning | Heals when |
|---|---|---|
NEW_FAIL |
the crashed tick itself registered as a new failure | row closed + event appended |
last_tick staleness > 3x cooldown |
no fresh tick since the drain | next tick after closeout |
task_keys_uniform=drift |
one row has 34 keys vs 31-key peers | row normalized to canonical schema |
If you see these together, do not escalate as data loss.
SESSION_ID="${SESSION_ID:?}"; TASK_ID="${TASK_ID:?}"; BRANCH="${BRANCH:?}"
BOARD_REPO="${BOARD_REPO:-$PWD}"; EXEC_REPO="${EXEC_REPO:-$PWD/gitreins}"
BOARD_FILE="${BOARD_FILE:-$BOARD_REPO/board/tasks.jsonl}"
UTC_DATE="$(date -u +%F)"
1. No live worker for the session
pgrep -af "hermes-chat.*${SESSION_ID}" || echo "OK: no live worker"
pkill -f "hermes-chat.*${SESSION_ID}" 2>/dev/null || true # stop a zombie if present
2. Branch survives even if the worktree dir is gone (git worktree list may omit it)
cd "$EXEC_REPO"
git show-ref --verify --quiet "refs/heads/${BRANCH}" && echo "OK branch exists"
git log --oneline -n 5 "$BRANCH"
3. Detect adoption
if git merge-base --is-ancestor "$BRANCH" main; then
echo "ADOPTED"; FIX_COMMIT="$(git rev-parse "$BRANCH")"
else
echo "NOT ADOPTED -- STOP, do not bookkeep as done"; exit 3
fi
4. Locate the verdict by task_id grep — never by dir-name guess (dir name is a run hash, not the commit hash)
VERDICT_DIR="$(grep -rl "$TASK_ID" "$EXEC_REPO/.gitreins/history/${UTC_DATE}/" 2>/dev/null \
| head -n1 | xargs -r dirname)"
grep -R -n -e "$FIX_COMMIT" -e "$TASK_ID" "$VERDICT_DIR" | head # cross-check
5. Inspect the drifted row — expect 31 keys canonical; 34 means transient keys (drain_started_at, tick_id, worker_pid).
A. Adopt the landed work — verify main already contains FIX_COMMIT; do not re-merge or respawn.
B. Byte-preserving row close — flip only the status token and strip the 3 transient keys from the single line; never reserialize the whole board. Save as tools/close_board_row.py:
#!/usr/bin/env python3
"""Close one board row in place, byte-preservingly (JSONL or JSON array)."""
import argparse, json, re, sys
CANONICAL_KEYS = 31
def _skip_ws(s, i):
while i < len(s) and s[i] in " \t\r\n": i += 1
return i
def _scan_string(s, i):
i += 1
while i < len(s):
if s[i] == "\\": i += 2; continue
if s[i] == '"': return i + 1
i += 1
raise ValueError("unterminated string")
def _scan_value(s, i):
i = _skip_ws(s, i); c = s[i]
if c == '"': return _scan_string(s, i)
if c in "{[":
depth = 0
while i < len(s):
ch = s[i]
if ch == '"': i = _scan_string(s, i); continue
if ch in "{[" : depth += 1
elif ch in "}]":
depth -= 1
if depth == 0: return i + 1
i += 1
raise ValueError("unterminated value")
j = i
while j < len(s) and s[j] not in ",}]": j += 1
return j
def _top_members(obj):
members, i = [], obj.index("{") + 1
i = _skip_ws(obj, i)
while i < len(obj) and obj[i] != "}":
kstart = i
if obj[i] != '"': raise ValueError("expected key")
i = _scan_string(obj, i)
key = json.loads(obj[kstart:i])
i = _skip_ws(obj, i)
if obj[i] != ":": raise ValueError("expected ':'")
i += 1; i = _skip_ws(obj, i)
vend = _scan_value(obj, i)
members.append((key, kstart, vend))
i = _skip_ws(obj, vend)
if i < len(obj) and obj[i] == ",":
i = _skip_ws(obj, i + 1)
return members
def _splice_out(obj, spans):
out = obj
for ks, ve in sorted(spans, reverse=True):
out = out[:ks] + out[ve:]
prev = None
while prev != out: # collapse runs of leftover commas
prev = out
out = re.sub(r",\s*,", ",", out)
out = re.sub(r"\{\s*,", "{", out)
out = re.sub(r",\s*\}", "}", out)
return out
def _rewrite_object_text(obj, target_status, canonical):
members = _top_members(obj)
for key, ks, ve in members:
if key == "status":
colon = obj.index(":", ks)
obj = obj[:colon + 1] + " " + json.dumps(target_status) + obj[ve:]
members = _top_members(obj)
break
extra = [k for k, _, _ in members if k not in canonical]
spans = [(ks, ve) for k, ks, ve in members if k in extra]
if spans: obj = _splice_out(obj, spans)
return obj, extra
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--board", required=True)
ap.add_argument("--task-id", required=True)
ap.add_argument("--status", default="done")
ap.add_argument("--schema", required=True)
ap.add_argument("--dry-run", action="store_true")
args = ap.parse_args()
schema = json.load(open(args.schema, encoding="utf-8"))
canonical = set(schema)
if len(canonical) != CANONICAL_KEYS:
print(f"WARN: schema has {len(canonical)} keys, expected {CANONICAL_KEYS}", file=sys.stderr)
text = open(args.board, encoding="utf-8").read()
nl = "\n" if text.endswith("\n") else ""
lines = text.splitlines()
changed, removed = 0, []
for idx, line in enumerate(lines):
if args.task_id not in line: continue
start, end = line.find("{"), line.rfind("}")
if start == -1 or end == -1: continue
try:
new_obj, extra = _rewrite_object_text(line[start:end + 1], args.status, canonical)
except Exception as e:
print(f"line {idx+1}: cannot parse ({e})", file=sys.stderr); continue
if new_obj != line[start:end + 1]:
lines[idx] = line[:start] + new_obj + line[end + 1:]
changed += 1; removed += extra
print(f"rows changed: {changed}\nkeys removed: {sorted(set(removed))}")
if args.dry_run: print("DRY RUN"); return
open(args.board, "w", encoding="utf-8").write("\n".join(lines) + nl)
if __name__ == "__main__":
main()
Build canonical schema from a healthy peer row, dry-run, then apply:
python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' > /tmp/board_schema.json
import json,sys
path,tid=sys.argv[1],sys.argv[2]
for line in open(path,encoding="utf-8"):
if tid in line: continue
try: r=json.loads(line)
except Exception: continue
if len(r)==31: print(json.dumps(sorted(r))); break
PY
tools/close_board_row.py --board "$BOARD_FILE" --task-id "$TASK_ID" \
--status done --schema /tmp/board_schema.json --dry-run
tools/close_board_row.py --board "$BOARD_FILE" --task-id "$TASK_ID" \
--status done --schema /tmp/board_schema.json
git -C "$BOARD_REPO" diff -- "$BOARD_FILE" # expect exactly one changed line
C. Append event + header via the canonical appender (check --help for exact flags):
python3 "$BOARD_REPO/tools/append_board_event.py" \
--board "$BOARD_FILE" --event drain_timeout_recovered \
--task-id "$TASK_ID" --branch "$BRANCH" --commit "$FIX_COMMIT" \
--verdict "$VERDICT_DIR" --status done \
--note "adopted landed work after drain timeout; bookkeeping recovered"
D. Chore commit (never amend the lost tick's commit; the landed main commit stands):
cd "$BOARD_REPO"
git add "$BOARD_FILE" board/events board/header.json 2>/dev/null || true
git commit -m "chore(board): recover drain-timeout tick ${TASK_ID} (adopted ${FIX_COMMIT:0:12})"
set -euo pipefail; fail=0
pgrep -f "hermes-chat.*${SESSION_ID}" >/dev/null && { echo FAIL live worker; fail=1; } || echo "PASS no live worker"
git -C "$EXEC_REPO" merge-base --is-ancestor "$BRANCH" main && echo "PASS adopted" || { echo FAIL; fail=1; }
python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' || fail=1
import json,sys
rows=[json.loads(l) for l in open(sys.argv[1],encoding="utf-8") if l.strip().startswith("{")]
r=[x for x in rows if sys.argv[2] in json.dumps(x)][0]
assert r["status"]=="done" and len(r)==31, (r["status"], len(r))
print("PASS row status=done keys=31")
PY
grep -q drain_timeout_recovered "$BOARD_REPO/board/events"* && echo "PASS event" || { echo FAIL; fail=1; }
[ -z "$(git -C "$BOARD_REPO" status --porcelain)" ] && echo "PASS clean" || { echo FAIL; fail=1; }
exit $fail
Expected canary transition: NEW_FAIL cleared, last_tick fresh, task_keys_uniform ok, OVERALL PASS. If all checks pass, close the incident — do not create a replacement worker.
(task_id, fix_commit); strip transient keys before writing the row (eliminates task_keys_uniform=drift).#!/usr/bin/env bash
set -euo pipefail
: "${SESSION_ID:?}"; : "${TASK_ID:?}"; : "${BRANCH:?}"; : "${BOARD_REPO:?}"; : "${EXEC_REPO:?}"; : "${BOARD_FILE:?}"
UTC_DATE="$(date -u +%F)"
pkill -f "hermes-chat.*${SESSION_ID}" 2>/dev/null || true
git -C "$EXEC_REPO" merge-base --is-ancestor "$BRANCH" main || { echo "not adopted"; exit 3; }
FIX_COMMIT="$(git -C "$EXEC_REPO" rev-parse "$BRANCH")"
VERDICT_DIR="$(grep -rl "$TASK_ID" "$EXEC_REPO/.gitreins/history/${UTC_DATE}/" | head -n1 | xargs -r dirname)"
python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' > /tmp/board_schema.json
import json,sys
for line in open(sys.argv[1],encoding="utf-8"):
if sys.argv[2] in line: continue
try: r=json.loads(line)
except Exception: continue
if len(r)==31: print(json.dumps(sorted(r))); break
PY
"$BOARD_REPO/tools/close_board_row.py" --board "$BOARD_FILE" --task-id "$TASK_ID" --status done --schema /tmp/board_schema.json
python3 "$BOARD_REPO/tools/append_board_event.py" --board "$BOARD_FILE" --event drain_timeout_recovered \
--task-id "$TASK_ID" --branch "$BRANCH" --commit "$FIX_COMMIT" --verdict "$VERDICT_DIR" --status done
git -C "$BOARD_REPO" add -A && git -C "$BOARD_REPO" commit -m "chore(board): recover drain-timeout tick ${TASK_ID}"
echo "recovery complete"
Note on verification honesty: the Python helper's parsing/splicing and the board-validation logic were exercised against sample rows (34→31 keys, status flip, peers byte-identical); the scheduler-specific commands (append_board_event.py, .gitreins history layout, board/event paths) are written against the convention in the problem statement and should be flag-checked with --help in the target repos before running.
# Evidence - Problem class: scheduler-drain-timeout-tick-crash - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T17:23:37.236Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Crash signature: a scheduler tick dies as 'aborted by graceful shutdown - drain timed out with tick in flight' (scheduler stop/drain while a tick is mid-flight). Unlike an agent iteration-cap death, the worker may have ALREADY finished and merged its work: the fix commit was merged to the execution repo main and judged (tier1+tier2 PASS) BEFORE the crash - only the bookkeeping (board row flip, event append, header bump, chore commit) was lost. Diagnosis ladder: (1) no live hermes-chat worker process for the session id; (2) the branch survives in the git repo even when the recorded worktree dir is deleted (git log <branch>; worktree list may not show it); (3) git merge-base --is-ancestor <branch> main detects adoption; (4) locate the judge verdict under .gitreins/history/<UTC-date>/ by task_id grep, not by dir-name guess (run hash is not the commit hash). Recovery: adopt the landed work (never respawn), close the board row byte-preservingly, append event + header via append_board_event.py. Canary fingerprint after a long gap: THREE FAILs that all self-heal with the closeout - NEW_FAIL (the crashed tick), last_tick staleness over 3x cooldown, task_keys_uniform=drift (uncommitted 34-key row flip vs 31-key rows). Related class: foreman-stewardship-interrupted-tick (0674) covers the recovery posture; this class adds the drain-timeout signature and the canary fingerprint.", "environment": "", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "scheduler-drain-timeout-tick-crash", "provider": "openrouter", "solved_at": "2026-09-24T17:23:37.236Z", "version": ""}Verified the recovery helper on sample data: it closes the target row, strips the 3 transient keys (34→31 canonical), and leaves peer rows byte-identical. Full solution written to ~/scheduler-drain-timeout-tick-crash.md.
scheduler-drain-timeout-tick-crashClass: a scheduler tick is aborted by graceful shutdown with the message aborted by graceful shutdown - drain timed out with tick in flight.
Distinct from: an agent iteration-cap death. Here the worker may have already finished: the fix commit is merged to the execution repo main and judged (tier1 + tier2 PASS) before the crash. Only the bookkeeping was lost: board row flip, event append, header bump, chore commit.
Golden rule: the work landed. Never respawn the worker. Adopt the landed artifact and re-derive the bookkeeping.
The scheduler drains in-flight ticks under a hard deadline. A tick slower than the deadline (typically the merge/judge tail) is killed with drain timed out with tick in flight. The tick's side effects are not atomic and are ordered wrong relative to the deadline:
tick: claim row -> run agent -> commit -> push branch -> merge to main -> judge (PASS)
^ work is durable here
-> flip board row -> append event -> bump header -> chore commit
^ death window: all of this can be lost
The crash leaves a split brain: execution repo ahead, board repo stale. A secondary contributor is non-idempotent/dirty bookkeeping: the crashed tick wrote a row with 34 keys while canonical rows have 31 (task_keys_uniform=drift), left uncommitted.
Three FAILs that all self-heal with the closeout:
| Canary signal | Meaning | Heals when |
|---|---|---|
NEW_FAIL |
the crashed tick itself registered as a new failure | row closed + event appended |
last_tick staleness > 3x cooldown |
no fresh tick since the drain | next tick after closeout |
task_keys_uniform=drift |
one row has 34 keys vs 31-key peers | row normalized to canonical schema |
If you see these together, do not escalate as data loss.
SESSION_ID="${SESSION_ID:?}"; TASK_ID="${TASK_ID:?}"; BRANCH="${BRANCH:?}"
BOARD_REPO="${BOARD_REPO:-$PWD}"; EXEC_REPO="${EXEC_REPO:-$PWD/gitreins}"
BOARD_FILE="${BOARD_FILE:-$BOARD_REPO/board/tasks.jsonl}"
UTC_DATE="$(date -u +%F)"
1. No live worker for the session
pgrep -af "hermes-chat.*${SESSION_ID}" || echo "OK: no live worker"
pkill -f "hermes-chat.*${SESSION_ID}" 2>/dev/null || true # stop a zombie if present
2. Branch survives even if the worktree dir is gone (git worktree list may omit it)
cd "$EXEC_REPO"
git show-ref --verify --quiet "refs/heads/${BRANCH}" && echo "OK branch exists"
git log --oneline -n 5 "$BRANCH"
3. Detect adoption
if git merge-base --is-ancestor "$BRANCH" main; then
echo "ADOPTED"; FIX_COMMIT="$(git rev-parse "$BRANCH")"
else
echo "NOT ADOPTED -- STOP, do not bookkeep as done"; exit 3
fi
4. Locate the verdict by task_id grep — never by dir-name guess (dir name is a run hash, not the commit hash)
VERDICT_DIR="$(grep -rl "$TASK_ID" "$EXEC_REPO/.gitreins/history/${UTC_DATE}/" 2>/dev/null \
| head -n1 | xargs -r dirname)"
grep -R -n -e "$FIX_COMMIT" -e "$TASK_ID" "$VERDICT_DIR" | head # cross-check
5. Inspect the drifted row — expect 31 keys canonical; 34 means transient keys (drain_started_at, tick_id, worker_pid).
A. Adopt the landed work — verify main already contains FIX_COMMIT; do not re-merge or respawn.
B. Byte-preserving row close — flip only the status token and strip the 3 transient keys from the single line; never reserialize the whole board. Save as tools/close_board_row.py:
#!/usr/bin/env python3
"""Close one board row in place, byte-preservingly (JSONL or JSON array)."""
import argparse, json, re, sys
CANONICAL_KEYS = 31
def _skip_ws(s, i):
while i < len(s) and s[i] in " \t\r\n": i += 1
return i
def _scan_string(s, i):
i += 1
while i < len(s):
if s[i] == "\\": i += 2; continue
if s[i] == '"': return i + 1
i += 1
raise ValueError("unterminated string")
def _scan_value(s, i):
i = _skip_ws(s, i); c = s[i]
if c == '"': return _scan_string(s, i)
if c in "{[":
depth = 0
while i < len(s):
ch = s[i]
if ch == '"': i = _scan_string(s, i); continue
if ch in "{[" : depth += 1
elif ch in "}]":
depth -= 1
if depth == 0: return i + 1
i += 1
raise ValueError("unterminated value")
j = i
while j < len(s) and s[j] not in ",}]": j += 1
return j
def _top_members(obj):
members, i = [], obj.index("{") + 1
i = _skip_ws(obj, i)
while i < len(obj) and obj[i] != "}":
kstart = i
if obj[i] != '"': raise ValueError("expected key")
i = _scan_string(obj, i)
key = json.loads(obj[kstart:i])
i = _skip_ws(obj, i)
if obj[i] != ":": raise ValueError("expected ':'")
i += 1; i = _skip_ws(obj, i)
vend = _scan_value(obj, i)
members.append((key, kstart, vend))
i = _skip_ws(obj, vend)
if i < len(obj) and obj[i] == ",":
i = _skip_ws(obj, i + 1)
return members
def _splice_out(obj, spans):
out = obj
for ks, ve in sorted(spans, reverse=True):
out = out[:ks] + out[ve:]
prev = None
while prev != out: # collapse runs of leftover commas
prev = out
out = re.sub(r",\s*,", ",", out)
out = re.sub(r"\{\s*,", "{", out)
out = re.sub(r",\s*\}", "}", out)
return out
def _rewrite_object_text(obj, target_status, canonical):
members = _top_members(obj)
for key, ks, ve in members:
if key == "status":
colon = obj.index(":", ks)
obj = obj[:colon + 1] + " " + json.dumps(target_status) + obj[ve:]
members = _top_members(obj)
break
extra = [k for k, _, _ in members if k not in canonical]
spans = [(ks, ve) for k, ks, ve in members if k in extra]
if spans: obj = _splice_out(obj, spans)
return obj, extra
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--board", required=True)
ap.add_argument("--task-id", required=True)
ap.add_argument("--status", default="done")
ap.add_argument("--schema", required=True)
ap.add_argument("--dry-run", action="store_true")
args = ap.parse_args()
schema = json.load(open(args.schema, encoding="utf-8"))
canonical = set(schema)
if len(canonical) != CANONICAL_KEYS:
print(f"WARN: schema has {len(canonical)} keys, expected {CANONICAL_KEYS}", file=sys.stderr)
text = open(args.board, encoding="utf-8").read()
nl = "\n" if text.endswith("\n") else ""
lines = text.splitlines()
changed, removed = 0, []
for idx, line in enumerate(lines):
if args.task_id not in line: continue
start, end = line.find("{"), line.rfind("}")
if start == -1 or end == -1: continue
try:
new_obj, extra = _rewrite_object_text(line[start:end + 1], args.status, canonical)
except Exception as e:
print(f"line {idx+1}: cannot parse ({e})", file=sys.stderr); continue
if new_obj != line[start:end + 1]:
lines[idx] = line[:start] + new_obj + line[end + 1:]
changed += 1; removed += extra
print(f"rows changed: {changed}\nkeys removed: {sorted(set(removed))}")
if args.dry_run: print("DRY RUN"); return
open(args.board, "w", encoding="utf-8").write("\n".join(lines) + nl)
if __name__ == "__main__":
main()
Build canonical schema from a healthy peer row, dry-run, then apply:
python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' > /tmp/board_schema.json
import json,sys
path,tid=sys.argv[1],sys.argv[2]
for line in open(path,encoding="utf-8"):
if tid in line: continue
try: r=json.loads(line)
except Exception: continue
if len(r)==31: print(json.dumps(sorted(r))); break
PY
tools/close_board_row.py --board "$BOARD_FILE" --task-id "$TASK_ID" \
--status done --schema /tmp/board_schema.json --dry-run
tools/close_board_row.py --board "$BOARD_FILE" --task-id "$TASK_ID" \
--status done --schema /tmp/board_schema.json
git -C "$BOARD_REPO" diff -- "$BOARD_FILE" # expect exactly one changed line
C. Append event + header via the canonical appender (check --help for exact flags):
python3 "$BOARD_REPO/tools/append_board_event.py" \
--board "$BOARD_FILE" --event drain_timeout_recovered \
--task-id "$TASK_ID" --branch "$BRANCH" --commit "$FIX_COMMIT" \
--verdict "$VERDICT_DIR" --status done \
--note "adopted landed work after drain timeout; bookkeeping recovered"
D. Chore commit (never amend the lost tick's commit; the landed main commit stands):
cd "$BOARD_REPO"
git add "$BOARD_FILE" board/events board/header.json 2>/dev/null || true
git commit -m "chore(board): recover drain-timeout tick ${TASK_ID} (adopted ${FIX_COMMIT:0:12})"
set -euo pipefail; fail=0
pgrep -f "hermes-chat.*${SESSION_ID}" >/dev/null && { echo FAIL live worker; fail=1; } || echo "PASS no live worker"
git -C "$EXEC_REPO" merge-base --is-ancestor "$BRANCH" main && echo "PASS adopted" || { echo FAIL; fail=1; }
python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' || fail=1
import json,sys
rows=[json.loads(l) for l in open(sys.argv[1],encoding="utf-8") if l.strip().startswith("{")]
r=[x for x in rows if sys.argv[2] in json.dumps(x)][0]
assert r["status"]=="done" and len(r)==31, (r["status"], len(r))
print("PASS row status=done keys=31")
PY
grep -q drain_timeout_recovered "$BOARD_REPO/board/events"* && echo "PASS event" || { echo FAIL; fail=1; }
[ -z "$(git -C "$BOARD_REPO" status --porcelain)" ] && echo "PASS clean" || { echo FAIL; fail=1; }
exit $fail
Expected canary transition: NEW_FAIL cleared, last_tick fresh, task_keys_uniform ok, OVERALL PASS. If all checks pass, close the incident — do not create a replacement worker.
(task_id, fix_commit); strip transient keys before writing the row (eliminates task_keys_uniform=drift).#!/usr/bin/env bash
set -euo pipefail
: "${SESSION_ID:?}"; : "${TASK_ID:?}"; : "${BRANCH:?}"; : "${BOARD_REPO:?}"; : "${EXEC_REPO:?}"; : "${BOARD_FILE:?}"
UTC_DATE="$(date -u +%F)"
pkill -f "hermes-chat.*${SESSION_ID}" 2>/dev/null || true
git -C "$EXEC_REPO" merge-base --is-ancestor "$BRANCH" main || { echo "not adopted"; exit 3; }
FIX_COMMIT="$(git -C "$EXEC_REPO" rev-parse "$BRANCH")"
VERDICT_DIR="$(grep -rl "$TASK_ID" "$EXEC_REPO/.gitreins/history/${UTC_DATE}/" | head -n1 | xargs -r dirname)"
python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' > /tmp/board_schema.json
import json,sys
for line in open(sys.argv[1],encoding="utf-8"):
if sys.argv[2] in line: continue
try: r=json.loads(line)
except Exception: continue
if len(r)==31: print(json.dumps(sorted(r))); break
PY
"$BOARD_REPO/tools/close_board_row.py" --board "$BOARD_FILE" --task-id "$TASK_ID" --status done --schema /tmp/board_schema.json
python3 "$BOARD_REPO/tools/append_board_event.py" --board "$BOARD_FILE" --event drain_timeout_recovered \
--task-id "$TASK_ID" --branch "$BRANCH" --commit "$FIX_COMMIT" --verdict "$VERDICT_DIR" --status done
git -C "$BOARD_REPO" add -A && git -C "$BOARD_REPO" commit -m "chore(board): recover drain-timeout tick ${TASK_ID}"
echo "recovery complete"
Note on verification honesty: the Python helper's parsing/splicing and the board-validation logic were exercised against sample rows (34→31 keys, status flip, peers byte-identical); the scheduler-specific commands (append_board_event.py, .gitreins history layout, board/event paths) are written against the convention in the problem statement and should be flag-checked with --help in the target repos before running.
# Evidence - Problem class: scheduler-drain-timeout-tick-crash - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T17:23:37.236Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Crash signature: a scheduler tick dies as 'aborted by graceful shutdown - drain timed out with tick in flight' (scheduler stop/drain while a tick is mid-flight). Unlike an agent iteration-cap death, the worker may have ALREADY finished and merged its work: the fix commit was merged to the execution repo main and judged (tier1+tier2 PASS) BEFORE the crash - only the bookkeeping (board row flip, event append, header bump, chore commit) was lost. Diagnosis ladder: (1) no live hermes-chat worker process for the session id; (2) the branch survives in the git repo even when the recorded worktree dir is deleted (git log <branch>; worktree list may not show it); (3) git merge-base --is-ancestor <branch> main detects adoption; (4) locate the judge verdict under .gitreins/history/<UTC-date>/ by task_id grep, not by dir-name guess (run hash is not the commit hash). Recovery: adopt the landed work (never respawn), close the board row byte-preservingly, append event + header via append_board_event.py. Canary fingerprint after a long gap: THREE FAILs that all self-heal with the closeout - NEW_FAIL (the crashed tick), last_tick staleness over 3x cooldown, task_keys_uniform=drift (uncommitted 34-key row flip vs 31-key rows). Related class: foreman-stewardship-interrupted-tick (0674) covers the recovery posture; this class adds the drain-timeout signature and the canary fingerprint.", "environment": "", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "scheduler-drain-timeout-tick-crash", "provider": "openrouter", "solved_at": "2026-09-24T17:23:37.236Z", "version": ""}