foreman-refired-tick-wave-harvest
Solution written to ~/SOLUTION.md.
The failure is a non-idempotent tick handler on a reused tick id:
--min-interval bucket re-addresses the same manifest path.dispatched was treated as terminal enough to skip but never harvested — so re-entry either duplicates legs or leaves landed rows open forever.-Q log stays 0 bytes, and pgrep self-matches the reconciler, so live workers were misclassified.old-dangerous unit had ExecStartPre=fuser -k 9090/tcp, Restart=always, RestartSec=3, which re-fires the same bucket.tick.sh — flock single-owner lease per tick id + manifest gate; an existing manifest always routes to harvest, never dispatch.harvest.sh — the resumable 7-step protocol: enumerate legs, git merge-base --is-ancestor to prove merges, read existing verdict.json, force-complete hung merged legs (--force --skip-tier2) then judge --async, close judged rows with their own evidence, leave live sessions untouched, atomically mark the manifest done.pending.sh — select new work only from tasks not seen in any wave manifest.fuser -k, use Restart=on-failure + backoff, EnvironmentFile, fix the 0600 nobody unit.I built a throwaway repo and exercised every branch with a stub gitreins: fresh admit, re-fire denial, concurrent lease denial, landed+judged close, landed+unjudged force-complete/judge, live-leg preservation, and second-run idempotency — all passed. The merge-base probes returned 0/1 correctly.
Notably, verification caught a real bug: sqlite3 defaults to |-separated columns, so read -r n age swallowed both values and silently marked every session dead. Fixed with IFS='|' read -r n age < <(sqlite3 ...), documented in the doc's gotcha section.
# Evidence - Problem class: foreman-refired-tick-wave-harvest - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-22T21:55:40.919Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A scheduler tick re-fired with the SAME tick id as an earlier incarnation that had already dispatched a wave (manifest .coding-hermes/waves/<tick-id>.json, status dispatched). Freshness protocol that worked: (1) verify via git log + wave manifest + gitreins tasks.yaml which wave legs landed; (2) merge-base --is-ancestor each wt branch tip against main to prove merges; (3) read .gitreins/history/<date>/<id>/verdict.json for tasks already judged but rows left open; (4) close the landed rows with their OWN verdict/commit evidence and mark the manifest done; (5) leave live legs alive - the earlier incarnation may still have workers running (verify via state.db session message count + age, NOT the -Q log which stays 0 bytes; not pgrep which self-matches); (6) for merged-but-unjudged legs: gitreins task complete --force --skip-tier2 (kill the hang after the state flips), then gitreins judge --async; (7) pick NEW work only from what remains pending.", "environment": "coding-hermes scheduler tick, gitreins harness, JSONL board", "language": "process", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "foreman-refired-tick-wave-harvest", "provider": "openrouter", "solved_at": "2026-09-22T21:55:40.919Z", "version": ""}Solution written to ~/SOLUTION.md.
The failure is a non-idempotent tick handler on a reused tick id:
--min-interval bucket re-addresses the same manifest path.dispatched was treated as terminal enough to skip but never harvested — so re-entry either duplicates legs or leaves landed rows open forever.-Q log stays 0 bytes, and pgrep self-matches the reconciler, so live workers were misclassified.old-dangerous unit had ExecStartPre=fuser -k 9090/tcp, Restart=always, RestartSec=3, which re-fires the same bucket.tick.sh — flock single-owner lease per tick id + manifest gate; an existing manifest always routes to harvest, never dispatch.harvest.sh — the resumable 7-step protocol: enumerate legs, git merge-base --is-ancestor to prove merges, read existing verdict.json, force-complete hung merged legs (--force --skip-tier2) then judge --async, close judged rows with their own evidence, leave live sessions untouched, atomically mark the manifest done.pending.sh — select new work only from tasks not seen in any wave manifest.fuser -k, use Restart=on-failure + backoff, EnvironmentFile, fix the 0600 nobody unit.I built a throwaway repo and exercised every branch with a stub gitreins: fresh admit, re-fire denial, concurrent lease denial, landed+judged close, landed+unjudged force-complete/judge, live-leg preservation, and second-run idempotency — all passed. The merge-base probes returned 0/1 correctly.
Notably, verification caught a real bug: sqlite3 defaults to |-separated columns, so read -r n age swallowed both values and silently marked every session dead. Fixed with IFS='|' read -r n age < <(sqlite3 ...), documented in the doc's gotcha section.
# Evidence - Problem class: foreman-refired-tick-wave-harvest - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-22T21:55:40.919Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A scheduler tick re-fired with the SAME tick id as an earlier incarnation that had already dispatched a wave (manifest .coding-hermes/waves/<tick-id>.json, status dispatched). Freshness protocol that worked: (1) verify via git log + wave manifest + gitreins tasks.yaml which wave legs landed; (2) merge-base --is-ancestor each wt branch tip against main to prove merges; (3) read .gitreins/history/<date>/<id>/verdict.json for tasks already judged but rows left open; (4) close the landed rows with their OWN verdict/commit evidence and mark the manifest done; (5) leave live legs alive - the earlier incarnation may still have workers running (verify via state.db session message count + age, NOT the -Q log which stays 0 bytes; not pgrep which self-matches); (6) for merged-but-unjudged legs: gitreins task complete --force --skip-tier2 (kill the hang after the state flips), then gitreins judge --async; (7) pick NEW work only from what remains pending.", "environment": "coding-hermes scheduler tick, gitreins harness, JSONL board", "language": "process", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "foreman-refired-tick-wave-harvest", "provider": "openrouter", "solved_at": "2026-09-22T21:55:40.919Z", "version": ""}