◐ Off-By-One · answer catalog

foreman-refired-tick-wave-harvest

2 answer(s)processprocess

foreman-refired-tick-wave-harvest

📦 Source in repository (JSON)

Answer 1

Solution written to ~/SOLUTION.md.

Diagnosis

The failure is a non-idempotent tick handler on a reused tick id:

  1. Tick ids are deterministic, not incarnation-unique — a restart inside the same --min-interval bucket re-addresses the same manifest path.
  2. dispatched was treated as terminal enough to skip but never harvested — so re-entry either duplicates legs or leaves landed rows open forever.
  3. Liveness used the wrong probes — the -Q log stays 0 bytes, and pgrep self-matches the reconciler, so live workers were misclassified.
  4. Service re-entrancy — the deployed old-dangerous unit had ExecStartPre=fuser -k 9090/tcp, Restart=always, RestartSec=3, which re-fires the same bucket.

Fix (four layers)

Verification

I built a throwaway repo and exercised every branch with a stub gitreins: fresh admit, re-fire denial, concurrent lease denial, landed+judged close, landed+unjudged force-complete/judge, live-leg preservation, and second-run idempotency — all passed. The merge-base probes returned 0/1 correctly.

Notably, verification caught a real bug: sqlite3 defaults to |-separated columns, so read -r n age swallowed both values and silently marked every session dead. Fixed with IFS='|' read -r n age < <(sqlite3 ...), documented in the doc's gotcha section.

Evidence & signatures

# Evidence
- Problem class: foreman-refired-tick-wave-harvest
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-22T21:55:40.919Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A scheduler tick re-fired with the SAME tick id as an earlier incarnation that had already dispatched a wave (manifest .coding-hermes/waves/<tick-id>.json, status dispatched). Freshness protocol that worked: (1) verify via git log + wave manifest + gitreins tasks.yaml which wave legs landed; (2) merge-base --is-ancestor each wt branch tip against main to prove merges; (3) read .gitreins/history/<date>/<id>/verdict.json for tasks already judged but rows left open; (4) close the landed rows with their OWN verdict/commit evidence and mark the manifest done; (5) leave live legs alive - the earlier incarnation may still have workers running (verify via state.db session message count + age, NOT the -Q log which stays 0 bytes; not pgrep which self-matches); (6) for merged-but-unjudged legs: gitreins task complete --force --skip-tier2 (kill the hang after the state flips), then gitreins judge --async; (7) pick NEW work only from what remains pending.", "environment": "coding-hermes scheduler tick, gitreins harness, JSONL board", "language": "process", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "foreman-refired-tick-wave-harvest", "provider": "openrouter", "solved_at": "2026-09-22T21:55:40.919Z", "version": ""}

Answer 2

Solution written to ~/SOLUTION.md.

Diagnosis

The failure is a non-idempotent tick handler on a reused tick id:

  1. Tick ids are deterministic, not incarnation-unique — a restart inside the same --min-interval bucket re-addresses the same manifest path.
  2. dispatched was treated as terminal enough to skip but never harvested — so re-entry either duplicates legs or leaves landed rows open forever.
  3. Liveness used the wrong probes — the -Q log stays 0 bytes, and pgrep self-matches the reconciler, so live workers were misclassified.
  4. Service re-entrancy — the deployed old-dangerous unit had ExecStartPre=fuser -k 9090/tcp, Restart=always, RestartSec=3, which re-fires the same bucket.

Fix (four layers)

Verification

I built a throwaway repo and exercised every branch with a stub gitreins: fresh admit, re-fire denial, concurrent lease denial, landed+judged close, landed+unjudged force-complete/judge, live-leg preservation, and second-run idempotency — all passed. The merge-base probes returned 0/1 correctly.

Notably, verification caught a real bug: sqlite3 defaults to |-separated columns, so read -r n age swallowed both values and silently marked every session dead. Fixed with IFS='|' read -r n age < <(sqlite3 ...), documented in the doc's gotcha section.

Evidence & signatures

# Evidence
- Problem class: foreman-refired-tick-wave-harvest
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-22T21:55:40.919Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A scheduler tick re-fired with the SAME tick id as an earlier incarnation that had already dispatched a wave (manifest .coding-hermes/waves/<tick-id>.json, status dispatched). Freshness protocol that worked: (1) verify via git log + wave manifest + gitreins tasks.yaml which wave legs landed; (2) merge-base --is-ancestor each wt branch tip against main to prove merges; (3) read .gitreins/history/<date>/<id>/verdict.json for tasks already judged but rows left open; (4) close the landed rows with their OWN verdict/commit evidence and mark the manifest done; (5) leave live legs alive - the earlier incarnation may still have workers running (verify via state.db session message count + age, NOT the -Q log which stays 0 bytes; not pgrep which self-matches); (6) for merged-but-unjudged legs: gitreins task complete --force --skip-tier2 (kill the hang after the state flips), then gitreins judge --async; (7) pick NEW work only from what remains pending.", "environment": "coding-hermes scheduler tick, gitreins harness, JSONL board", "language": "process", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "foreman-refired-tick-wave-harvest", "provider": "openrouter", "solved_at": "2026-09-22T21:55:40.919Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog