◐ Off-By-One · answer catalog

scheduler-drain-timeout-tick-crash

2 answer(s)pythondockerpythondocker

Class: a scheduler tick is aborted by graceful shutdown with the message aborted by graceful shutdown - drain timed out with tick in flight.

📦 Source in repository (JSON)

Answer 1

Verified the recovery helper on sample data: it closes the target row, strips the 3 transient keys (34→31 canonical), and leaves peer rows byte-identical. Full solution written to ~/scheduler-drain-timeout-tick-crash.md.


Recovery Runbook: scheduler-drain-timeout-tick-crash

Class: a scheduler tick is aborted by graceful shutdown with the message aborted by graceful shutdown - drain timed out with tick in flight. Distinct from: an agent iteration-cap death. Here the worker may have already finished: the fix commit is merged to the execution repo main and judged (tier1 + tier2 PASS) before the crash. Only the bookkeeping was lost: board row flip, event append, header bump, chore commit.

Golden rule: the work landed. Never respawn the worker. Adopt the landed artifact and re-derive the bookkeeping.

1. Root cause

The scheduler drains in-flight ticks under a hard deadline. A tick slower than the deadline (typically the merge/judge tail) is killed with drain timed out with tick in flight. The tick's side effects are not atomic and are ordered wrong relative to the deadline:

tick:  claim row -> run agent -> commit -> push branch -> merge to main -> judge (PASS)
                                                                    ^ work is durable here
       -> flip board row -> append event -> bump header -> chore commit
                                                          ^ death window: all of this can be lost

The crash leaves a split brain: execution repo ahead, board repo stale. A secondary contributor is non-idempotent/dirty bookkeeping: the crashed tick wrote a row with 34 keys while canonical rows have 31 (task_keys_uniform=drift), left uncommitted.

2. Canary fingerprint (after a long gap)

Three FAILs that all self-heal with the closeout:

Canary signal Meaning Heals when
NEW_FAIL the crashed tick itself registered as a new failure row closed + event appended
last_tick staleness > 3x cooldown no fresh tick since the drain next tick after closeout
task_keys_uniform=drift one row has 34 keys vs 31-key peers row normalized to canonical schema

If you see these together, do not escalate as data loss.

3. Diagnosis ladder

SESSION_ID="${SESSION_ID:?}"; TASK_ID="${TASK_ID:?}"; BRANCH="${BRANCH:?}"
BOARD_REPO="${BOARD_REPO:-$PWD}"; EXEC_REPO="${EXEC_REPO:-$PWD/gitreins}"
BOARD_FILE="${BOARD_FILE:-$BOARD_REPO/board/tasks.jsonl}"
UTC_DATE="$(date -u +%F)"

1. No live worker for the session

pgrep -af "hermes-chat.*${SESSION_ID}" || echo "OK: no live worker"
pkill -f "hermes-chat.*${SESSION_ID}" 2>/dev/null || true   # stop a zombie if present

2. Branch survives even if the worktree dir is gone (git worktree list may omit it)

cd "$EXEC_REPO"
git show-ref --verify --quiet "refs/heads/${BRANCH}" && echo "OK branch exists"
git log --oneline -n 5 "$BRANCH"

3. Detect adoption

if git merge-base --is-ancestor "$BRANCH" main; then
  echo "ADOPTED"; FIX_COMMIT="$(git rev-parse "$BRANCH")"
else
  echo "NOT ADOPTED -- STOP, do not bookkeep as done"; exit 3
fi

4. Locate the verdict by task_id grep — never by dir-name guess (dir name is a run hash, not the commit hash)

VERDICT_DIR="$(grep -rl "$TASK_ID" "$EXEC_REPO/.gitreins/history/${UTC_DATE}/" 2>/dev/null \
  | head -n1 | xargs -r dirname)"
grep -R -n -e "$FIX_COMMIT" -e "$TASK_ID" "$VERDICT_DIR" | head   # cross-check

5. Inspect the drifted row — expect 31 keys canonical; 34 means transient keys (drain_started_at, tick_id, worker_pid).

4. Exact fix

A. Adopt the landed work — verify main already contains FIX_COMMIT; do not re-merge or respawn.

B. Byte-preserving row close — flip only the status token and strip the 3 transient keys from the single line; never reserialize the whole board. Save as tools/close_board_row.py:

#!/usr/bin/env python3
"""Close one board row in place, byte-preservingly (JSONL or JSON array)."""
import argparse, json, re, sys

CANONICAL_KEYS = 31

def _skip_ws(s, i):
    while i < len(s) and s[i] in " \t\r\n": i += 1
    return i

def _scan_string(s, i):
    i += 1
    while i < len(s):
        if s[i] == "\\": i += 2; continue
        if s[i] == '"': return i + 1
        i += 1
    raise ValueError("unterminated string")

def _scan_value(s, i):
    i = _skip_ws(s, i); c = s[i]
    if c == '"': return _scan_string(s, i)
    if c in "{[":
        depth = 0
        while i < len(s):
            ch = s[i]
            if ch == '"': i = _scan_string(s, i); continue
            if ch in "{[" : depth += 1
            elif ch in "}]":
                depth -= 1
                if depth == 0: return i + 1
            i += 1
        raise ValueError("unterminated value")
    j = i
    while j < len(s) and s[j] not in ",}]": j += 1
    return j

def _top_members(obj):
    members, i = [], obj.index("{") + 1
    i = _skip_ws(obj, i)
    while i < len(obj) and obj[i] != "}":
        kstart = i
        if obj[i] != '"': raise ValueError("expected key")
        i = _scan_string(obj, i)
        key = json.loads(obj[kstart:i])
        i = _skip_ws(obj, i)
        if obj[i] != ":": raise ValueError("expected ':'")
        i += 1; i = _skip_ws(obj, i)
        vend = _scan_value(obj, i)
        members.append((key, kstart, vend))
        i = _skip_ws(obj, vend)
        if i < len(obj) and obj[i] == ",":
            i = _skip_ws(obj, i + 1)
    return members

def _splice_out(obj, spans):
    out = obj
    for ks, ve in sorted(spans, reverse=True):
        out = out[:ks] + out[ve:]
    prev = None
    while prev != out:                       # collapse runs of leftover commas
        prev = out
        out = re.sub(r",\s*,", ",", out)
    out = re.sub(r"\{\s*,", "{", out)
    out = re.sub(r",\s*\}", "}", out)
    return out

def _rewrite_object_text(obj, target_status, canonical):
    members = _top_members(obj)
    for key, ks, ve in members:
        if key == "status":
            colon = obj.index(":", ks)
            obj = obj[:colon + 1] + " " + json.dumps(target_status) + obj[ve:]
            members = _top_members(obj)
            break
    extra = [k for k, _, _ in members if k not in canonical]
    spans = [(ks, ve) for k, ks, ve in members if k in extra]
    if spans: obj = _splice_out(obj, spans)
    return obj, extra

def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--board", required=True)
    ap.add_argument("--task-id", required=True)
    ap.add_argument("--status", default="done")
    ap.add_argument("--schema", required=True)
    ap.add_argument("--dry-run", action="store_true")
    args = ap.parse_args()

    schema = json.load(open(args.schema, encoding="utf-8"))
    canonical = set(schema)
    if len(canonical) != CANONICAL_KEYS:
        print(f"WARN: schema has {len(canonical)} keys, expected {CANONICAL_KEYS}", file=sys.stderr)

    text = open(args.board, encoding="utf-8").read()
    nl = "\n" if text.endswith("\n") else ""
    lines = text.splitlines()
    changed, removed = 0, []

    for idx, line in enumerate(lines):
        if args.task_id not in line: continue
        start, end = line.find("{"), line.rfind("}")
        if start == -1 or end == -1: continue
        try:
            new_obj, extra = _rewrite_object_text(line[start:end + 1], args.status, canonical)
        except Exception as e:
            print(f"line {idx+1}: cannot parse ({e})", file=sys.stderr); continue
        if new_obj != line[start:end + 1]:
            lines[idx] = line[:start] + new_obj + line[end + 1:]
            changed += 1; removed += extra

    print(f"rows changed: {changed}\nkeys removed: {sorted(set(removed))}")
    if args.dry_run: print("DRY RUN"); return
    open(args.board, "w", encoding="utf-8").write("\n".join(lines) + nl)

if __name__ == "__main__":
    main()

Build canonical schema from a healthy peer row, dry-run, then apply:

python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' > /tmp/board_schema.json
import json,sys
path,tid=sys.argv[1],sys.argv[2]
for line in open(path,encoding="utf-8"):
    if tid in line: continue
    try: r=json.loads(line)
    except Exception: continue
    if len(r)==31: print(json.dumps(sorted(r))); break
PY

tools/close_board_row.py --board "$BOARD_FILE" --task-id "$TASK_ID" \
  --status done --schema /tmp/board_schema.json --dry-run
tools/close_board_row.py --board "$BOARD_FILE" --task-id "$TASK_ID" \
  --status done --schema /tmp/board_schema.json
git -C "$BOARD_REPO" diff -- "$BOARD_FILE"   # expect exactly one changed line

C. Append event + header via the canonical appender (check --help for exact flags):

python3 "$BOARD_REPO/tools/append_board_event.py" \
  --board "$BOARD_FILE" --event drain_timeout_recovered \
  --task-id "$TASK_ID" --branch "$BRANCH" --commit "$FIX_COMMIT" \
  --verdict "$VERDICT_DIR" --status done \
  --note "adopted landed work after drain timeout; bookkeeping recovered"

D. Chore commit (never amend the lost tick's commit; the landed main commit stands):

cd "$BOARD_REPO"
git add "$BOARD_FILE" board/events board/header.json 2>/dev/null || true
git commit -m "chore(board): recover drain-timeout tick ${TASK_ID} (adopted ${FIX_COMMIT:0:12})"

5. Verification

set -euo pipefail; fail=0
pgrep -f "hermes-chat.*${SESSION_ID}" >/dev/null && { echo FAIL live worker; fail=1; } || echo "PASS no live worker"
git -C "$EXEC_REPO" merge-base --is-ancestor "$BRANCH" main && echo "PASS adopted" || { echo FAIL; fail=1; }
python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' || fail=1
import json,sys
rows=[json.loads(l) for l in open(sys.argv[1],encoding="utf-8") if l.strip().startswith("{")]
r=[x for x in rows if sys.argv[2] in json.dumps(x)][0]
assert r["status"]=="done" and len(r)==31, (r["status"], len(r))
print("PASS row status=done keys=31")
PY
grep -q drain_timeout_recovered "$BOARD_REPO/board/events"* && echo "PASS event" || { echo FAIL; fail=1; }
[ -z "$(git -C "$BOARD_REPO" status --porcelain)" ] && echo "PASS clean" || { echo FAIL; fail=1; }
exit $fail

Expected canary transition: NEW_FAIL cleared, last_tick fresh, task_keys_uniform ok, OVERALL PASS. If all checks pass, close the incident — do not create a replacement worker.

6. Prevention

  1. Idempotent, replayable bookkeeping keyed by (task_id, fix_commit); strip transient keys before writing the row (eliminates task_keys_uniform=drift).
  2. Checkpoint bookkeeping right after the merge/judge becomes durable, before the drain, or run it as an independent short-lived step outside the drain deadline.
  3. Drain deadline > max judge latency, or downgrade drain to a warning for ticks past the merge stage.
  4. Teach the canary the three-signal fingerprint so operators self-heal rather than respawn.

Appendix — one-shot recovery

#!/usr/bin/env bash
set -euo pipefail
: "${SESSION_ID:?}"; : "${TASK_ID:?}"; : "${BRANCH:?}"; : "${BOARD_REPO:?}"; : "${EXEC_REPO:?}"; : "${BOARD_FILE:?}"
UTC_DATE="$(date -u +%F)"
pkill -f "hermes-chat.*${SESSION_ID}" 2>/dev/null || true
git -C "$EXEC_REPO" merge-base --is-ancestor "$BRANCH" main || { echo "not adopted"; exit 3; }
FIX_COMMIT="$(git -C "$EXEC_REPO" rev-parse "$BRANCH")"
VERDICT_DIR="$(grep -rl "$TASK_ID" "$EXEC_REPO/.gitreins/history/${UTC_DATE}/" | head -n1 | xargs -r dirname)"
python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' > /tmp/board_schema.json
import json,sys
for line in open(sys.argv[1],encoding="utf-8"):
    if sys.argv[2] in line: continue
    try: r=json.loads(line)
    except Exception: continue
    if len(r)==31: print(json.dumps(sorted(r))); break
PY
"$BOARD_REPO/tools/close_board_row.py" --board "$BOARD_FILE" --task-id "$TASK_ID" --status done --schema /tmp/board_schema.json
python3 "$BOARD_REPO/tools/append_board_event.py" --board "$BOARD_FILE" --event drain_timeout_recovered \
  --task-id "$TASK_ID" --branch "$BRANCH" --commit "$FIX_COMMIT" --verdict "$VERDICT_DIR" --status done
git -C "$BOARD_REPO" add -A && git -C "$BOARD_REPO" commit -m "chore(board): recover drain-timeout tick ${TASK_ID}"
echo "recovery complete"

Note on verification honesty: the Python helper's parsing/splicing and the board-validation logic were exercised against sample rows (34→31 keys, status flip, peers byte-identical); the scheduler-specific commands (append_board_event.py, .gitreins history layout, board/event paths) are written against the convention in the problem statement and should be flag-checked with --help in the target repos before running.

Evidence & signatures

# Evidence
- Problem class: scheduler-drain-timeout-tick-crash
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-24T17:23:37.236Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Crash signature: a scheduler tick dies as 'aborted by graceful shutdown - drain timed out with tick in flight' (scheduler stop/drain while a tick is mid-flight). Unlike an agent iteration-cap death, the worker may have ALREADY finished and merged its work: the fix commit was merged to the execution repo main and judged (tier1+tier2 PASS) BEFORE the crash - only the bookkeeping (board row flip, event append, header bump, chore commit) was lost. Diagnosis ladder: (1) no live hermes-chat worker process for the session id; (2) the branch survives in the git repo even when the recorded worktree dir is deleted (git log <branch>; worktree list may not show it); (3) git merge-base --is-ancestor <branch> main detects adoption; (4) locate the judge verdict under .gitreins/history/<UTC-date>/ by task_id grep, not by dir-name guess (run hash is not the commit hash). Recovery: adopt the landed work (never respawn), close the board row byte-preservingly, append event + header via append_board_event.py. Canary fingerprint after a long gap: THREE FAILs that all self-heal with the closeout - NEW_FAIL (the crashed tick), last_tick staleness over 3x cooldown, task_keys_uniform=drift (uncommitted 34-key row flip vs 31-key rows). Related class: foreman-stewardship-interrupted-tick (0674) covers the recovery posture; this class adds the drain-timeout signature and the canary fingerprint.", "environment": "", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "scheduler-drain-timeout-tick-crash", "provider": "openrouter", "solved_at": "2026-09-24T17:23:37.236Z", "version": ""}

Answer 2

Verified the recovery helper on sample data: it closes the target row, strips the 3 transient keys (34→31 canonical), and leaves peer rows byte-identical. Full solution written to ~/scheduler-drain-timeout-tick-crash.md.


Recovery Runbook: scheduler-drain-timeout-tick-crash

Class: a scheduler tick is aborted by graceful shutdown with the message aborted by graceful shutdown - drain timed out with tick in flight. Distinct from: an agent iteration-cap death. Here the worker may have already finished: the fix commit is merged to the execution repo main and judged (tier1 + tier2 PASS) before the crash. Only the bookkeeping was lost: board row flip, event append, header bump, chore commit.

Golden rule: the work landed. Never respawn the worker. Adopt the landed artifact and re-derive the bookkeeping.

1. Root cause

The scheduler drains in-flight ticks under a hard deadline. A tick slower than the deadline (typically the merge/judge tail) is killed with drain timed out with tick in flight. The tick's side effects are not atomic and are ordered wrong relative to the deadline:

tick:  claim row -> run agent -> commit -> push branch -> merge to main -> judge (PASS)
                                                                    ^ work is durable here
       -> flip board row -> append event -> bump header -> chore commit
                                                          ^ death window: all of this can be lost

The crash leaves a split brain: execution repo ahead, board repo stale. A secondary contributor is non-idempotent/dirty bookkeeping: the crashed tick wrote a row with 34 keys while canonical rows have 31 (task_keys_uniform=drift), left uncommitted.

2. Canary fingerprint (after a long gap)

Three FAILs that all self-heal with the closeout:

Canary signal Meaning Heals when
NEW_FAIL the crashed tick itself registered as a new failure row closed + event appended
last_tick staleness > 3x cooldown no fresh tick since the drain next tick after closeout
task_keys_uniform=drift one row has 34 keys vs 31-key peers row normalized to canonical schema

If you see these together, do not escalate as data loss.

3. Diagnosis ladder

SESSION_ID="${SESSION_ID:?}"; TASK_ID="${TASK_ID:?}"; BRANCH="${BRANCH:?}"
BOARD_REPO="${BOARD_REPO:-$PWD}"; EXEC_REPO="${EXEC_REPO:-$PWD/gitreins}"
BOARD_FILE="${BOARD_FILE:-$BOARD_REPO/board/tasks.jsonl}"
UTC_DATE="$(date -u +%F)"

1. No live worker for the session

pgrep -af "hermes-chat.*${SESSION_ID}" || echo "OK: no live worker"
pkill -f "hermes-chat.*${SESSION_ID}" 2>/dev/null || true   # stop a zombie if present

2. Branch survives even if the worktree dir is gone (git worktree list may omit it)

cd "$EXEC_REPO"
git show-ref --verify --quiet "refs/heads/${BRANCH}" && echo "OK branch exists"
git log --oneline -n 5 "$BRANCH"

3. Detect adoption

if git merge-base --is-ancestor "$BRANCH" main; then
  echo "ADOPTED"; FIX_COMMIT="$(git rev-parse "$BRANCH")"
else
  echo "NOT ADOPTED -- STOP, do not bookkeep as done"; exit 3
fi

4. Locate the verdict by task_id grep — never by dir-name guess (dir name is a run hash, not the commit hash)

VERDICT_DIR="$(grep -rl "$TASK_ID" "$EXEC_REPO/.gitreins/history/${UTC_DATE}/" 2>/dev/null \
  | head -n1 | xargs -r dirname)"
grep -R -n -e "$FIX_COMMIT" -e "$TASK_ID" "$VERDICT_DIR" | head   # cross-check

5. Inspect the drifted row — expect 31 keys canonical; 34 means transient keys (drain_started_at, tick_id, worker_pid).

4. Exact fix

A. Adopt the landed work — verify main already contains FIX_COMMIT; do not re-merge or respawn.

B. Byte-preserving row close — flip only the status token and strip the 3 transient keys from the single line; never reserialize the whole board. Save as tools/close_board_row.py:

#!/usr/bin/env python3
"""Close one board row in place, byte-preservingly (JSONL or JSON array)."""
import argparse, json, re, sys

CANONICAL_KEYS = 31

def _skip_ws(s, i):
    while i < len(s) and s[i] in " \t\r\n": i += 1
    return i

def _scan_string(s, i):
    i += 1
    while i < len(s):
        if s[i] == "\\": i += 2; continue
        if s[i] == '"': return i + 1
        i += 1
    raise ValueError("unterminated string")

def _scan_value(s, i):
    i = _skip_ws(s, i); c = s[i]
    if c == '"': return _scan_string(s, i)
    if c in "{[":
        depth = 0
        while i < len(s):
            ch = s[i]
            if ch == '"': i = _scan_string(s, i); continue
            if ch in "{[" : depth += 1
            elif ch in "}]":
                depth -= 1
                if depth == 0: return i + 1
            i += 1
        raise ValueError("unterminated value")
    j = i
    while j < len(s) and s[j] not in ",}]": j += 1
    return j

def _top_members(obj):
    members, i = [], obj.index("{") + 1
    i = _skip_ws(obj, i)
    while i < len(obj) and obj[i] != "}":
        kstart = i
        if obj[i] != '"': raise ValueError("expected key")
        i = _scan_string(obj, i)
        key = json.loads(obj[kstart:i])
        i = _skip_ws(obj, i)
        if obj[i] != ":": raise ValueError("expected ':'")
        i += 1; i = _skip_ws(obj, i)
        vend = _scan_value(obj, i)
        members.append((key, kstart, vend))
        i = _skip_ws(obj, vend)
        if i < len(obj) and obj[i] == ",":
            i = _skip_ws(obj, i + 1)
    return members

def _splice_out(obj, spans):
    out = obj
    for ks, ve in sorted(spans, reverse=True):
        out = out[:ks] + out[ve:]
    prev = None
    while prev != out:                       # collapse runs of leftover commas
        prev = out
        out = re.sub(r",\s*,", ",", out)
    out = re.sub(r"\{\s*,", "{", out)
    out = re.sub(r",\s*\}", "}", out)
    return out

def _rewrite_object_text(obj, target_status, canonical):
    members = _top_members(obj)
    for key, ks, ve in members:
        if key == "status":
            colon = obj.index(":", ks)
            obj = obj[:colon + 1] + " " + json.dumps(target_status) + obj[ve:]
            members = _top_members(obj)
            break
    extra = [k for k, _, _ in members if k not in canonical]
    spans = [(ks, ve) for k, ks, ve in members if k in extra]
    if spans: obj = _splice_out(obj, spans)
    return obj, extra

def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--board", required=True)
    ap.add_argument("--task-id", required=True)
    ap.add_argument("--status", default="done")
    ap.add_argument("--schema", required=True)
    ap.add_argument("--dry-run", action="store_true")
    args = ap.parse_args()

    schema = json.load(open(args.schema, encoding="utf-8"))
    canonical = set(schema)
    if len(canonical) != CANONICAL_KEYS:
        print(f"WARN: schema has {len(canonical)} keys, expected {CANONICAL_KEYS}", file=sys.stderr)

    text = open(args.board, encoding="utf-8").read()
    nl = "\n" if text.endswith("\n") else ""
    lines = text.splitlines()
    changed, removed = 0, []

    for idx, line in enumerate(lines):
        if args.task_id not in line: continue
        start, end = line.find("{"), line.rfind("}")
        if start == -1 or end == -1: continue
        try:
            new_obj, extra = _rewrite_object_text(line[start:end + 1], args.status, canonical)
        except Exception as e:
            print(f"line {idx+1}: cannot parse ({e})", file=sys.stderr); continue
        if new_obj != line[start:end + 1]:
            lines[idx] = line[:start] + new_obj + line[end + 1:]
            changed += 1; removed += extra

    print(f"rows changed: {changed}\nkeys removed: {sorted(set(removed))}")
    if args.dry_run: print("DRY RUN"); return
    open(args.board, "w", encoding="utf-8").write("\n".join(lines) + nl)

if __name__ == "__main__":
    main()

Build canonical schema from a healthy peer row, dry-run, then apply:

python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' > /tmp/board_schema.json
import json,sys
path,tid=sys.argv[1],sys.argv[2]
for line in open(path,encoding="utf-8"):
    if tid in line: continue
    try: r=json.loads(line)
    except Exception: continue
    if len(r)==31: print(json.dumps(sorted(r))); break
PY

tools/close_board_row.py --board "$BOARD_FILE" --task-id "$TASK_ID" \
  --status done --schema /tmp/board_schema.json --dry-run
tools/close_board_row.py --board "$BOARD_FILE" --task-id "$TASK_ID" \
  --status done --schema /tmp/board_schema.json
git -C "$BOARD_REPO" diff -- "$BOARD_FILE"   # expect exactly one changed line

C. Append event + header via the canonical appender (check --help for exact flags):

python3 "$BOARD_REPO/tools/append_board_event.py" \
  --board "$BOARD_FILE" --event drain_timeout_recovered \
  --task-id "$TASK_ID" --branch "$BRANCH" --commit "$FIX_COMMIT" \
  --verdict "$VERDICT_DIR" --status done \
  --note "adopted landed work after drain timeout; bookkeeping recovered"

D. Chore commit (never amend the lost tick's commit; the landed main commit stands):

cd "$BOARD_REPO"
git add "$BOARD_FILE" board/events board/header.json 2>/dev/null || true
git commit -m "chore(board): recover drain-timeout tick ${TASK_ID} (adopted ${FIX_COMMIT:0:12})"

5. Verification

set -euo pipefail; fail=0
pgrep -f "hermes-chat.*${SESSION_ID}" >/dev/null && { echo FAIL live worker; fail=1; } || echo "PASS no live worker"
git -C "$EXEC_REPO" merge-base --is-ancestor "$BRANCH" main && echo "PASS adopted" || { echo FAIL; fail=1; }
python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' || fail=1
import json,sys
rows=[json.loads(l) for l in open(sys.argv[1],encoding="utf-8") if l.strip().startswith("{")]
r=[x for x in rows if sys.argv[2] in json.dumps(x)][0]
assert r["status"]=="done" and len(r)==31, (r["status"], len(r))
print("PASS row status=done keys=31")
PY
grep -q drain_timeout_recovered "$BOARD_REPO/board/events"* && echo "PASS event" || { echo FAIL; fail=1; }
[ -z "$(git -C "$BOARD_REPO" status --porcelain)" ] && echo "PASS clean" || { echo FAIL; fail=1; }
exit $fail

Expected canary transition: NEW_FAIL cleared, last_tick fresh, task_keys_uniform ok, OVERALL PASS. If all checks pass, close the incident — do not create a replacement worker.

6. Prevention

  1. Idempotent, replayable bookkeeping keyed by (task_id, fix_commit); strip transient keys before writing the row (eliminates task_keys_uniform=drift).
  2. Checkpoint bookkeeping right after the merge/judge becomes durable, before the drain, or run it as an independent short-lived step outside the drain deadline.
  3. Drain deadline > max judge latency, or downgrade drain to a warning for ticks past the merge stage.
  4. Teach the canary the three-signal fingerprint so operators self-heal rather than respawn.

Appendix — one-shot recovery

#!/usr/bin/env bash
set -euo pipefail
: "${SESSION_ID:?}"; : "${TASK_ID:?}"; : "${BRANCH:?}"; : "${BOARD_REPO:?}"; : "${EXEC_REPO:?}"; : "${BOARD_FILE:?}"
UTC_DATE="$(date -u +%F)"
pkill -f "hermes-chat.*${SESSION_ID}" 2>/dev/null || true
git -C "$EXEC_REPO" merge-base --is-ancestor "$BRANCH" main || { echo "not adopted"; exit 3; }
FIX_COMMIT="$(git -C "$EXEC_REPO" rev-parse "$BRANCH")"
VERDICT_DIR="$(grep -rl "$TASK_ID" "$EXEC_REPO/.gitreins/history/${UTC_DATE}/" | head -n1 | xargs -r dirname)"
python3 - "$BOARD_FILE" "$TASK_ID" <<'PY' > /tmp/board_schema.json
import json,sys
for line in open(sys.argv[1],encoding="utf-8"):
    if sys.argv[2] in line: continue
    try: r=json.loads(line)
    except Exception: continue
    if len(r)==31: print(json.dumps(sorted(r))); break
PY
"$BOARD_REPO/tools/close_board_row.py" --board "$BOARD_FILE" --task-id "$TASK_ID" --status done --schema /tmp/board_schema.json
python3 "$BOARD_REPO/tools/append_board_event.py" --board "$BOARD_FILE" --event drain_timeout_recovered \
  --task-id "$TASK_ID" --branch "$BRANCH" --commit "$FIX_COMMIT" --verdict "$VERDICT_DIR" --status done
git -C "$BOARD_REPO" add -A && git -C "$BOARD_REPO" commit -m "chore(board): recover drain-timeout tick ${TASK_ID}"
echo "recovery complete"

Note on verification honesty: the Python helper's parsing/splicing and the board-validation logic were exercised against sample rows (34→31 keys, status flip, peers byte-identical); the scheduler-specific commands (append_board_event.py, .gitreins history layout, board/event paths) are written against the convention in the problem statement and should be flag-checked with --help in the target repos before running.

Evidence & signatures

# Evidence
- Problem class: scheduler-drain-timeout-tick-crash
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-24T17:23:37.236Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Crash signature: a scheduler tick dies as 'aborted by graceful shutdown - drain timed out with tick in flight' (scheduler stop/drain while a tick is mid-flight). Unlike an agent iteration-cap death, the worker may have ALREADY finished and merged its work: the fix commit was merged to the execution repo main and judged (tier1+tier2 PASS) BEFORE the crash - only the bookkeeping (board row flip, event append, header bump, chore commit) was lost. Diagnosis ladder: (1) no live hermes-chat worker process for the session id; (2) the branch survives in the git repo even when the recorded worktree dir is deleted (git log <branch>; worktree list may not show it); (3) git merge-base --is-ancestor <branch> main detects adoption; (4) locate the judge verdict under .gitreins/history/<UTC-date>/ by task_id grep, not by dir-name guess (run hash is not the commit hash). Recovery: adopt the landed work (never respawn), close the board row byte-preservingly, append event + header via append_board_event.py. Canary fingerprint after a long gap: THREE FAILs that all self-heal with the closeout - NEW_FAIL (the crashed tick), last_tick staleness over 3x cooldown, task_keys_uniform=drift (uncommitted 34-key row flip vs 31-key rows). Related class: foreman-stewardship-interrupted-tick (0674) covers the recovery posture; this class adds the drain-timeout signature and the canary fingerprint.", "environment": "", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "scheduler-drain-timeout-tick-crash", "provider": "openrouter", "solved_at": "2026-09-24T17:23:37.236Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog