Problem class: scheduler-foreman-stand-down-live-kanban-worker
All tests pass: the preflight gate correctly stands down on a live owner (and clears a free worktree), and the stand-down is verified byte-identical to the worktree/index while creating the remote ref.
Problem class: scheduler-foreman-stand-down-live-kanban-worker
One-line rule: A pending board row on a worktree that already contains a live worker is not a re-dispatch invitation. The foreman's only correct action is a read-only, zero-write stand-down, plus a remote-ref-only push of the existing branch. Never take a second writer on a shared worktree.
A foreman tick fires onto wt/<task> and finds:
| Observation | Wrong conclusion it invites |
|---|---|
Board row still reads pending |
"Nobody claimed it; dispatch." |
Branch is N commits ahead of origin/main, unpushed |
"Stranded/abandoned; save it by re-dispatching." |
| Worktree has uncommitted edits | "Dead tick left a mess; clean up and restart." |
pgrep -af 'hermes chat' shows workers |
"I can see the workers, ownership is clear." (It isn't.) |
Every one of these is consistent with a healthy, in-flight worker. The pending status is a stale snapshot, not ground truth.
The scheduler decides to dispatch by reading tasks.status = 'pending' and then starts a worker. The transition pending -> in_progress (and the worker_pid write) happens in a separate step. There is a window — and, under retries, a permanent split — where a live worker exists while the row still reads pending. A foreman that keys off status alone will re-dispatch into a live writer.
pgrep -af 'hermes chat' enumerates workers from every lane. It cannot say which worktree a given PID owns. The only reliable attribution key is the process's working directory: /proc/<pid>/cwd. A task is foreign-owned iff some live coding-hermes-worker ... chat -q 'work kanban task <id>' process has cwd == <worktree>.
tasks.worker_pid + tasks.last_heartbeat_at and task_events(kind='heartbeat') are the liveness record, but only if the foreman actually reads them — and only if heartbeats are fresh. A stale DB must never be trusted over a live /proc match.
task_runs for the same card shows five prior runs ending status=blocked with summaries naming "a live duplicate worker (PID …) still owns the staged edits in this shared worktree." Adding a sixth writer is the highest-risk action available: two git add/git commit streams race the same index and two full test suites contend for the same CPU, flaking the legitimate run.
origin/mainWriting a board row from a worktree whose board files are 3 commits behind origin/main manufactures a board merge conflict on the eventual feature merge. Stewardship therefore belongs in operational telemetry (DuckBrain) and the foreman report — not in the git-tracked board.
Sibling rows on the same board may already be marked complete for the same failing CI test. Verify a failure is not already filed before filing another row.
Two layers: (A) a mandatory read-only preflight gate every foreman tick runs before dispatch, and (B) an atomic claim in the scheduler so the pending+live-worker split can no longer occur. A stand-down action handles ticks already in flight.
foreman-preflight.shExit codes: 0 CLEAR (dispatch may proceed), 42 STAND DOWN (foreign live owner), 3 UNKNOWN (do not dispatch; escalate). Reads only.
#!/usr/bin/env bash
# foreman-preflight.sh -- READ-ONLY ownership gate for a scheduler foreman tick.
# Usage: foreman-preflight.sh <worktree-path> [task-id]
set -euo pipefail
WT_RAW="${1:?usage: foreman-preflight.sh <worktree-path> [task-id]}"
TASK="${2:-}"
HEARTBEAT_TTL_SECONDS="${HEARTBEAT_TTL_SECONDS:-180}"
BOARDS_DIR="${HERMES_KANBAN_BOARDS:-$HOME/.hermes/kanban/boards}"
WT="$(cd "$WT_RAW" && pwd -P)"
printf 'worktree=%s\ntask=%s\nheartbeat_ttl=%ss\n\n' "$WT" "$TASK" "$HEARTBEAT_TTL_SECONDS"
# 1. /proc scan: the ONLY reliable ownership attribution.
owner_pids=()
while IFS= read -r -d '' p; do
pid="${p#/proc/}"; pid="${pid%/cwd}"
[[ "$pid" =~ ^[0-9]+$ ]] || continue
cwd="$(readlink "$p" 2>/dev/null || true)"
[[ "$cwd" == "$WT" ]] || continue
cmd="$(tr '\0' ' ' < "/proc/$pid/cmdline" 2>/dev/null || true)"
case "$cmd" in
*coding-hermes-worker*|*"kanban task"*)
owner_pids+=("$pid")
printf 'OWNER pid=%s cwd=%s\n argv=%s\n' "$pid" "$cwd" "$cmd" ;;
*)
printf 'cwd-match (not a worker) pid=%s argv=%s\n' "$pid" "$cmd" ;;
esac
done < <(find /proc -maxdepth 2 -name cwd -print0 2>/dev/null)
# 2. kanban.db via READ-ONLY URI: worker_pid, heartbeat, task_runs history.
db_live_pids=()
if [[ -n "$TASK" ]]; then
shopt -s nullglob
for db in "$BOARDS_DIR"/*/kanban.db; do
uri="file:${db}?mode=ro"
row="$(sqlite3 "$uri" \
"SELECT id,status,worker_pid,last_heartbeat_at,
CAST(strftime('%s','now') - strftime('%s',last_heartbeat_at) AS INTEGER)
FROM tasks WHERE id='${TASK//\'/\'\'}';" 2>/dev/null || true)"
[[ -n "$row" ]] && printf '\nboard=%s\n task=%s\n' "$db" "$row"
hb="$(sqlite3 "$uri" \
"SELECT COALESCE(MAX(created_at),'') FROM task_events
WHERE task_id='${TASK//\'/\'\'}' AND kind='heartbeat';" 2>/dev/null || true)"
[[ -n "$hb" ]] && printf ' last_heartbeat_event=%s\n' "$hb"
rpid="$(sqlite3 "$uri" \
"SELECT COALESCE(worker_pid,'') FROM tasks WHERE id='${TASK//\'/\'\'}';" 2>/dev/null || true)"
[[ -n "$rpid" ]] && db_live_pids+=("$rpid")
runs="$(sqlite3 -separator ' | ' "$uri" \
"SELECT id,status,substr(COALESCE(summary,''),1,140) FROM task_runs
WHERE task_id='${TASK//\'/\'\'}' ORDER BY id DESC LIMIT 5;" 2>/dev/null || true)"
if [[ -n "$runs" ]]; then
printf ' task_runs (most recent 5):\n'
while IFS= read -r line; do printf ' %s\n' "$line"; done <<<"$runs"
fi
done
shopt -u nullglob
fi
# 3. Children of the owner prove real in-flight progress (e.g. `go test`).
for pid in "${owner_pids[@]:-}"; do
[[ -n "$pid" ]] || continue
printf '\nchildren of owner pid=%s:\n' "$pid"
pgrep -P "$pid" -a 2>/dev/null | sed 's/^/ /' || printf ' (none)\n'
done
# 4. Decision.
owner_alive=0
for pid in "${owner_pids[@]:-}"; do
[[ -n "$pid" ]] || continue
if kill -0 "$pid" 2>/dev/null; then owner_alive=1; fi
done
if [[ "$owner_alive" -eq 1 ]]; then
printf '\nDECISION: STAND DOWN (exit 42) -- live worker cwd owns %s\n' "$WT"; exit 42
fi
for rpid in "${db_live_pids[@]:-}"; do
[[ -n "$rpid" ]] || continue
if kill -0 "$rpid" 2>/dev/null; then
printf '\nDECISION: STAND DOWN (exit 42) -- board worker_pid=%s is alive\n' "$rpid"; exit 42
fi
done
if [[ ${#owner_pids[@]} -eq 0 && ${#db_live_pids[@]} -eq 0 ]]; then
printf '\nDECISION: CLEAR (exit 0) -- no live owner evidence\n'; exit 0
fi
printf '\nDECISION: UNKNOWN (exit 3) -- do not dispatch; escalate\n'; exit 3
foreman-standdown.shThe only write permitted on a STAND DOWN is an existing branch ref to the remote. It touches neither the worktree nor the index, so it cannot collide with the live worker's in-flight git add/commit, and it protects the N stranded commits from worktree reaping.
#!/usr/bin/env bash
# foreman-standdown.sh -- zero-write stand-down. Usage: foreman-standdown.sh <worktree-path>
set -euo pipefail
WT="$(cd "${1:?usage: foreman-standdown.sh <worktree-path>}" && pwd -P)"
RPT_DIR="${HERMES_FOREMAN_REPORT_DIR:-$HOME/.hermes/foreman/reports}"
mkdir -p "$RPT_DIR"
RPT="$RPT_DIR/standdown-$(date -u +%Y%m%dT%H%M%SZ).md"
branch="$(git -C "$WT" rev-parse --abbrev-ref HEAD)"
# Verify push scope BEFORE pushing: `push: branches: [main]` => zero CI load.
ci_scope="$(grep -RhoE 'branches:[[:space:]]*\[[^]]*\]' "$WT"/.github/workflows 2>/dev/null | sort -u || true)"
{
printf '# Foreman stand-down %s\n\n' "$(date -u +%FT%TZ)"
printf -- '- worktree: `%s`\n' "$WT"
printf -- '- branch: `%s`\n' "$branch"
printf -- '- HEAD: `%s`\n' "$(git -C "$WT" rev-parse HEAD)"
printf -- '- unpushed commits: `%s`\n' "$(git -C "$WT" rev-list --count @{u}..HEAD 2>/dev/null || echo '?')"
printf -- '- CI push scope: `%s`\n' "${ci_scope:-<none found>}"
printf -- '- action: push existing branch ref only (no index/worktree/board writes)\n'
} >"$RPT"
git -C "$WT" push origin "HEAD:refs/heads/$branch" # sole write: remote ref
printf '\nstand-down recorded: %s\n' "$RPT"
var ErrTaskOwned = errors.New("task already has a live owner")
// ClaimTask atomically moves a task pending -> in_progress ONLY if no live
// worker owns it. Closes the pending+live-worker window in §2.1.
func (s *Scheduler) ClaimTask(ctx context.Context, taskID string) error {
return s.db.InTx(ctx, func(tx *sql.Tx) error {
var status string
var workerPID *int
var lastHB *time.Time
if err := tx.QueryRowContext(ctx,
`SELECT status, worker_pid, last_heartbeat_at
FROM tasks WHERE id = ?`, taskID,
).Scan(&status, &workerPID, &lastHB); err != nil {
return err
}
if status == "in_progress" && workerPID != nil &&
pidAlive(*workerPID) && heartbeatFresh(lastHB, heartbeatTTL) {
return ErrTaskOwned
}
res, err := tx.ExecContext(ctx,
`UPDATE tasks
SET status = 'in_progress', worker_pid = ?, last_heartbeat_at = now()
WHERE id = ? AND status = 'pending'`,
os.Getpid(), taskID)
if err != nil {
return err
}
if n, _ := res.RowsAffected(); n == 0 {
return ErrTaskOwned // someone else won the race; do not dispatch
}
return nil
})
}
Defence-in-depth gate on the tick:
func (s *Scheduler) onForemanTick(ctx context.Context, wt, taskID string) error {
switch owned, err := s.probeOwnership(ctx, wt, taskID); {
case err != nil:
return fmt.Errorf("ownership probe failed for %s: %w", wt, err)
case owned:
// READ-ONLY stand-down. No dispatch, no commit, no board edit, no judge.
return s.recordStandDown(ctx, wt, taskID) // telemetry only
}
return s.ClaimTask(ctx, taskID) // may still return ErrTaskOwned
}
wt=/path/to/wt/<task>; task=<id>
if ! bash foreman-preflight.sh "$wt" "$task"; then
bash foreman-standdown.sh "$wt" # exit 42 or 3: remote-ref push only
exit 0
fi
# CLEAR: safe to dispatch; the atomic claim still guards the race.
Prohibited on a STAND DOWN: dispatch, git add, git commit, board edit, judge run, and any second test suite.
Both suites pass. Files are in ~/solution/.
test-preflight.sh=== CASE 1: live owner in worktree ===
OWNER pid=147 cwd=/tmp/.../wt-live
argv=hermes --skills coding-hermes-worker chat -q 'work kanban task K-1234' ...
board=.../kanban.db
task=K-1234|in_progress|147|<fresh heartbeat>|0
last_heartbeat_event=<fresh>
task_runs (most recent 5):
5 | blocked | a live duplicate worker (PID 999) still owns the staged edits ...
children of owner pid=147:
149 sleep 1
DECISION: STAND DOWN (exit 42) -- live worker cwd owns /tmp/.../wt-live
PASS: live owner -> STAND DOWN (42)
=== CASE 2: free worktree, no owner ===
DECISION: CLEAR (exit 0) -- no live owner evidence
PASS: free worktree -> CLEAR (0)
RESULT: pass=2 fail=0
This exercises all three independent signals: the /proc cwd+argv match, the fresh DB heartbeat/worker_pid, and the live child process.
test-standdown.shPASS: worktree status unchanged
PASS: index hash unchanged
PASS: local HEAD unchanged
PASS: remote branch created at HEAD
RESULT: pass=4 fail=0
The remote branch is created at current HEAD while the dirty worktree and the index remain byte-identical — safe to run while a worker is mid-commit.
# Attribution, not lane enumeration:
for p in /proc/[0-9]*/cwd; do
[ "$(readlink "$p" 2>/dev/null)" = "$(cd wt/<task> && pwd -P)" ] && \
tr '\0' ' ' < "${p%/cwd}/cmdline" && echo
done
# Liveness + history, read-only:
sqlite3 "file:$HOME/.hermes/kanban/boards/<board>/kanban.db?mode=ro" \
"SELECT id,status,worker_pid,last_heartbeat_at FROM tasks WHERE id='<id>';" \
"SELECT id,status,substr(summary,1,160) FROM task_runs WHERE task_id='<id>' ORDER BY id DESC LIMIT 5;"
# Push scope before any push:
grep -R 'branches:' .github/workflows
Expected: an OWNER argv line, a fresh last_heartbeat_at, and task_runs showing prior blocked duplicates. Verdict: STAND DOWN — do not create a sixth writer.
pgrep -af 'hermes chat' for ownership — it is lane-agnostic. Use /proc/<pid>/cwd.pending row plus unpushed commits is not proof of abandonment.pwd -P (resolved) on both sides of the cwd match; symlinked worktrees otherwise yield false negatives.worker_pid must not override a live /proc match, and vice versa: either signal being live means stand down.3 and escalate — never guess by dispatching.Artifacts written: ~/solution/SOLUTION.md, foreman-preflight.sh, foreman-standdown.sh, test-preflight.sh, test-standdown.sh.
# Evidence - Problem class: scheduler-foreman-stand-down-live-kanban-worker - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-10-05T05:34:59.375Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A coding-hermes SCHEDULER foreman tick fired onto a git worktree (wt/<task>) whose board row was still `pending`, while a LIVE kanban implementation worker owned the same worktree. Symptoms that mislead: the branch was N commits ahead of origin/main and unpushed, the worktree had uncommitted edits, and the board row read pending -- all of which look like an abandoned/stranded tick begging for a re-dispatch. The correct foreman response is a READ-ONLY ZERO-WRITE STAND-DOWN, not a dispatch. Evidence chain that established foreign ownership: (1) scan /proc/<pid>/cwd for every pid and match the worktree path -- found ONE live `hermes ... --skills coding-hermes-worker ... chat -q 'work kanban task <id>'` process whose cwd was the worktree; (2) read the kanban card from ~/.hermes/kanban/boards/<board>/kanban.db (sqlite3 read-only URI `file:<path>?mode=ro`): tasks.worker_pid + last_heartbeat_at + task_events(kind=heartbeat) proved liveness with a fresh heartbeat and worker_pid matching the /proc scan; (3) process CHILDREN proved real progress -- a live `go test -short -p 1 ./internal/scheduler -run <TestName> -count=1` under the worker pid, plus a growing worker log; (4) task_runs history was the decisive artifact: FIVE prior runs of the SAME card ended status=blocked with summaries naming `a live duplicate worker (PID ...) still owns the staged edits in this shared worktree` -- i.e. the duplicate-dispatch defect had already happened repeatedly, so adding a sixth writer was the highest-risk action available. What was done instead: (a) NO dispatch, NO git add/commit, NO board edit, NO gitreins judge run -- each would have either collided with the live writer's index or contended with its in-flight full test suite and flaked its run; (b) push the EXISTING branch ref only (`git push origin HEAD:refs/heads/<wt-branch>`) -- a remote-ref write that touches neither the worktree nor the index, protecting N stranded commits from worktree reaping; (c) verify branch protection/trigger scope first: .github/workflows/* only trigger on `push: branches: [main]`, so pushing the wt branch costs zero CI load; (d) record the stewardship in DuckBrain + the report rather than the git-tracked board, because the worktree's board checkout was BEHIND origin/main (3 commits) and a board write there would have created a board merge conflict on the eventual feature merge. Key pitfalls: `pgrep -af 'hermes chat'` enumerates OTHER lanes' workers and cannot attribute ownership -- use /proc/<pid>/cwd; a paused/stale-looking row plus unpushed commits is NOT proof of abandonment; completed sibling rows on the same board (marked complete for the same failing CI test) show the board can carry duplicates, so verify a CI failure is not already filed before filing another row.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "scheduler-foreman-stand-down-live-kanban-worker", "provider": "openrouter", "solved_at": "2026-10-05T05:34:59.376Z", "version": ""}