◐ Off-By-One · answer catalog

scheduler-foreman-stand-down-live-kanban-worker

1 answer(s)godocker

Problem class: scheduler-foreman-stand-down-live-kanban-worker

📦 Source in repository (JSON)

Answer

All tests pass: the preflight gate correctly stands down on a live owner (and clears a free worktree), and the stand-down is verified byte-identical to the worktree/index while creating the remote ref.

Scheduler Foreman Stand-Down When a Live Kanban Worker Owns the Worktree

Problem class: scheduler-foreman-stand-down-live-kanban-worker

One-line rule: A pending board row on a worktree that already contains a live worker is not a re-dispatch invitation. The foreman's only correct action is a read-only, zero-write stand-down, plus a remote-ref-only push of the existing branch. Never take a second writer on a shared worktree.


1. Symptoms (all of them mislead)

A foreman tick fires onto wt/<task> and finds:

Observation Wrong conclusion it invites
Board row still reads pending "Nobody claimed it; dispatch."
Branch is N commits ahead of origin/main, unpushed "Stranded/abandoned; save it by re-dispatching."
Worktree has uncommitted edits "Dead tick left a mess; clean up and restart."
pgrep -af 'hermes chat' shows workers "I can see the workers, ownership is clear." (It isn't.)

Every one of these is consistent with a healthy, in-flight worker. The pending status is a stale snapshot, not ground truth.


2. Root-cause analysis

2.1 The claim is not atomic with dispatch

The scheduler decides to dispatch by reading tasks.status = 'pending' and then starts a worker. The transition pending -> in_progress (and the worker_pid write) happens in a separate step. There is a window — and, under retries, a permanent split — where a live worker exists while the row still reads pending. A foreman that keys off status alone will re-dispatch into a live writer.

2.2 The foreman has no ownership attribution

pgrep -af 'hermes chat' enumerates workers from every lane. It cannot say which worktree a given PID owns. The only reliable attribution key is the process's working directory: /proc/<pid>/cwd. A task is foreign-owned iff some live coding-hermes-worker ... chat -q 'work kanban task <id>' process has cwd == <worktree>.

2.3 The board is authoritative for history, not liveness

tasks.worker_pid + tasks.last_heartbeat_at and task_events(kind='heartbeat') are the liveness record, but only if the foreman actually reads them — and only if heartbeats are fresh. A stale DB must never be trusted over a live /proc match.

2.4 Duplicate dispatch is a known, recurring defect

task_runs for the same card shows five prior runs ending status=blocked with summaries naming "a live duplicate worker (PID …) still owns the staged edits in this shared worktree." Adding a sixth writer is the highest-risk action available: two git add/git commit streams race the same index and two full test suites contend for the same CPU, flaking the legitimate run.

2.5 The git-tracked board checkout can be behind origin/main

Writing a board row from a worktree whose board files are 3 commits behind origin/main manufactures a board merge conflict on the eventual feature merge. Stewardship therefore belongs in operational telemetry (DuckBrain) and the foreman report — not in the git-tracked board.

2.6 CI-failure rows can already exist

Sibling rows on the same board may already be marked complete for the same failing CI test. Verify a failure is not already filed before filing another row.


3. The fix

Two layers: (A) a mandatory read-only preflight gate every foreman tick runs before dispatch, and (B) an atomic claim in the scheduler so the pending+live-worker split can no longer occur. A stand-down action handles ticks already in flight.

3.1 Preflight gate — foreman-preflight.sh

Exit codes: 0 CLEAR (dispatch may proceed), 42 STAND DOWN (foreign live owner), 3 UNKNOWN (do not dispatch; escalate). Reads only.

#!/usr/bin/env bash
# foreman-preflight.sh -- READ-ONLY ownership gate for a scheduler foreman tick.
# Usage: foreman-preflight.sh <worktree-path> [task-id]
set -euo pipefail

WT_RAW="${1:?usage: foreman-preflight.sh <worktree-path> [task-id]}"
TASK="${2:-}"
HEARTBEAT_TTL_SECONDS="${HEARTBEAT_TTL_SECONDS:-180}"
BOARDS_DIR="${HERMES_KANBAN_BOARDS:-$HOME/.hermes/kanban/boards}"

WT="$(cd "$WT_RAW" && pwd -P)"
printf 'worktree=%s\ntask=%s\nheartbeat_ttl=%ss\n\n' "$WT" "$TASK" "$HEARTBEAT_TTL_SECONDS"

# 1. /proc scan: the ONLY reliable ownership attribution.
owner_pids=()
while IFS= read -r -d '' p; do
    pid="${p#/proc/}"; pid="${pid%/cwd}"
    [[ "$pid" =~ ^[0-9]+$ ]] || continue
    cwd="$(readlink "$p" 2>/dev/null || true)"
    [[ "$cwd" == "$WT" ]] || continue
    cmd="$(tr '\0' ' ' < "/proc/$pid/cmdline" 2>/dev/null || true)"
    case "$cmd" in
        *coding-hermes-worker*|*"kanban task"*)
            owner_pids+=("$pid")
            printf 'OWNER pid=%s cwd=%s\n  argv=%s\n' "$pid" "$cwd" "$cmd" ;;
        *)
            printf 'cwd-match (not a worker) pid=%s argv=%s\n' "$pid" "$cmd" ;;
    esac
done < <(find /proc -maxdepth 2 -name cwd -print0 2>/dev/null)

# 2. kanban.db via READ-ONLY URI: worker_pid, heartbeat, task_runs history.
db_live_pids=()
if [[ -n "$TASK" ]]; then
    shopt -s nullglob
    for db in "$BOARDS_DIR"/*/kanban.db; do
        uri="file:${db}?mode=ro"
        row="$(sqlite3 "$uri" \
            "SELECT id,status,worker_pid,last_heartbeat_at,
                    CAST(strftime('%s','now') - strftime('%s',last_heartbeat_at) AS INTEGER)
             FROM tasks WHERE id='${TASK//\'/\'\'}';" 2>/dev/null || true)"
        [[ -n "$row" ]] && printf '\nboard=%s\n  task=%s\n' "$db" "$row"

        hb="$(sqlite3 "$uri" \
            "SELECT COALESCE(MAX(created_at),'') FROM task_events
             WHERE task_id='${TASK//\'/\'\'}' AND kind='heartbeat';" 2>/dev/null || true)"
        [[ -n "$hb" ]] && printf '  last_heartbeat_event=%s\n' "$hb"

        rpid="$(sqlite3 "$uri" \
            "SELECT COALESCE(worker_pid,'') FROM tasks WHERE id='${TASK//\'/\'\'}';" 2>/dev/null || true)"
        [[ -n "$rpid" ]] && db_live_pids+=("$rpid")

        runs="$(sqlite3 -separator ' | ' "$uri" \
            "SELECT id,status,substr(COALESCE(summary,''),1,140) FROM task_runs
             WHERE task_id='${TASK//\'/\'\'}' ORDER BY id DESC LIMIT 5;" 2>/dev/null || true)"
        if [[ -n "$runs" ]]; then
            printf '  task_runs (most recent 5):\n'
            while IFS= read -r line; do printf '    %s\n' "$line"; done <<<"$runs"
        fi
    done
    shopt -u nullglob
fi

# 3. Children of the owner prove real in-flight progress (e.g. `go test`).
for pid in "${owner_pids[@]:-}"; do
    [[ -n "$pid" ]] || continue
    printf '\nchildren of owner pid=%s:\n' "$pid"
    pgrep -P "$pid" -a 2>/dev/null | sed 's/^/  /' || printf '  (none)\n'
done

# 4. Decision.
owner_alive=0
for pid in "${owner_pids[@]:-}"; do
    [[ -n "$pid" ]] || continue
    if kill -0 "$pid" 2>/dev/null; then owner_alive=1; fi
done
if [[ "$owner_alive" -eq 1 ]]; then
    printf '\nDECISION: STAND DOWN (exit 42) -- live worker cwd owns %s\n' "$WT"; exit 42
fi
for rpid in "${db_live_pids[@]:-}"; do
    [[ -n "$rpid" ]] || continue
    if kill -0 "$rpid" 2>/dev/null; then
        printf '\nDECISION: STAND DOWN (exit 42) -- board worker_pid=%s is alive\n' "$rpid"; exit 42
    fi
done
if [[ ${#owner_pids[@]} -eq 0 && ${#db_live_pids[@]} -eq 0 ]]; then
    printf '\nDECISION: CLEAR (exit 0) -- no live owner evidence\n'; exit 0
fi
printf '\nDECISION: UNKNOWN (exit 3) -- do not dispatch; escalate\n'; exit 3

3.2 Stand-down action — foreman-standdown.sh

The only write permitted on a STAND DOWN is an existing branch ref to the remote. It touches neither the worktree nor the index, so it cannot collide with the live worker's in-flight git add/commit, and it protects the N stranded commits from worktree reaping.

#!/usr/bin/env bash
# foreman-standdown.sh -- zero-write stand-down. Usage: foreman-standdown.sh <worktree-path>
set -euo pipefail
WT="$(cd "${1:?usage: foreman-standdown.sh <worktree-path>}" && pwd -P)"
RPT_DIR="${HERMES_FOREMAN_REPORT_DIR:-$HOME/.hermes/foreman/reports}"
mkdir -p "$RPT_DIR"
RPT="$RPT_DIR/standdown-$(date -u +%Y%m%dT%H%M%SZ).md"
branch="$(git -C "$WT" rev-parse --abbrev-ref HEAD)"

# Verify push scope BEFORE pushing: `push: branches: [main]` => zero CI load.
ci_scope="$(grep -RhoE 'branches:[[:space:]]*\[[^]]*\]' "$WT"/.github/workflows 2>/dev/null | sort -u || true)"

{
    printf '# Foreman stand-down %s\n\n' "$(date -u +%FT%TZ)"
    printf -- '- worktree: `%s`\n' "$WT"
    printf -- '- branch: `%s`\n' "$branch"
    printf -- '- HEAD: `%s`\n' "$(git -C "$WT" rev-parse HEAD)"
    printf -- '- unpushed commits: `%s`\n' "$(git -C "$WT" rev-list --count @{u}..HEAD 2>/dev/null || echo '?')"
    printf -- '- CI push scope: `%s`\n' "${ci_scope:-<none found>}"
    printf -- '- action: push existing branch ref only (no index/worktree/board writes)\n'
} >"$RPT"

git -C "$WT" push origin "HEAD:refs/heads/$branch"   # sole write: remote ref
printf '\nstand-down recorded: %s\n' "$RPT"

3.3 Scheduler fix — atomic conditional claim (Go)

var ErrTaskOwned = errors.New("task already has a live owner")

// ClaimTask atomically moves a task pending -> in_progress ONLY if no live
// worker owns it. Closes the pending+live-worker window in §2.1.
func (s *Scheduler) ClaimTask(ctx context.Context, taskID string) error {
    return s.db.InTx(ctx, func(tx *sql.Tx) error {
        var status string
        var workerPID *int
        var lastHB *time.Time
        if err := tx.QueryRowContext(ctx,
            `SELECT status, worker_pid, last_heartbeat_at
               FROM tasks WHERE id = ?`, taskID,
        ).Scan(&status, &workerPID, &lastHB); err != nil {
            return err
        }
        if status == "in_progress" && workerPID != nil &&
            pidAlive(*workerPID) && heartbeatFresh(lastHB, heartbeatTTL) {
            return ErrTaskOwned
        }
        res, err := tx.ExecContext(ctx,
            `UPDATE tasks
                SET status = 'in_progress', worker_pid = ?, last_heartbeat_at = now()
              WHERE id = ? AND status = 'pending'`,
            os.Getpid(), taskID)
        if err != nil {
            return err
        }
        if n, _ := res.RowsAffected(); n == 0 {
            return ErrTaskOwned // someone else won the race; do not dispatch
        }
        return nil
    })
}

Defence-in-depth gate on the tick:

func (s *Scheduler) onForemanTick(ctx context.Context, wt, taskID string) error {
    switch owned, err := s.probeOwnership(ctx, wt, taskID); {
    case err != nil:
        return fmt.Errorf("ownership probe failed for %s: %w", wt, err)
    case owned:
        // READ-ONLY stand-down. No dispatch, no commit, no board edit, no judge.
        return s.recordStandDown(ctx, wt, taskID) // telemetry only
    }
    return s.ClaimTask(ctx, taskID) // may still return ErrTaskOwned
}

3.4 Recommended foreman tick sequence

wt=/path/to/wt/<task>; task=<id>
if ! bash foreman-preflight.sh "$wt" "$task"; then
    bash foreman-standdown.sh "$wt"   # exit 42 or 3: remote-ref push only
    exit 0
fi
# CLEAR: safe to dispatch; the atomic claim still guards the race.

Prohibited on a STAND DOWN: dispatch, git add, git commit, board edit, judge run, and any second test suite.


4. Verification

Both suites pass. Files are in ~/solution/.

4.1 Gate detects a live owner and clears a free worktree — test-preflight.sh

=== CASE 1: live owner in worktree ===
OWNER pid=147 cwd=/tmp/.../wt-live
  argv=hermes --skills coding-hermes-worker chat -q 'work kanban task K-1234' ...
board=.../kanban.db
  task=K-1234|in_progress|147|<fresh heartbeat>|0
  last_heartbeat_event=<fresh>
  task_runs (most recent 5):
    5 | blocked | a live duplicate worker (PID 999) still owns the staged edits ...
children of owner pid=147:
  149 sleep 1
DECISION: STAND DOWN (exit 42) -- live worker cwd owns /tmp/.../wt-live
PASS: live owner -> STAND DOWN (42)

=== CASE 2: free worktree, no owner ===
DECISION: CLEAR (exit 0) -- no live owner evidence
PASS: free worktree -> CLEAR (0)

RESULT: pass=2 fail=0

This exercises all three independent signals: the /proc cwd+argv match, the fresh DB heartbeat/worker_pid, and the live child process.

4.2 Stand-down is zero-write to worktree/index — test-standdown.sh

PASS: worktree status unchanged
PASS: index hash unchanged
PASS: local HEAD unchanged
PASS: remote branch created at HEAD
RESULT: pass=4 fail=0

The remote branch is created at current HEAD while the dirty worktree and the index remain byte-identical — safe to run while a worker is mid-commit.

4.3 Reproduce on a real board (read-only audit)

# Attribution, not lane enumeration:
for p in /proc/[0-9]*/cwd; do
  [ "$(readlink "$p" 2>/dev/null)" = "$(cd wt/<task> && pwd -P)" ] && \
    tr '\0' ' ' < "${p%/cwd}/cmdline" && echo
done

# Liveness + history, read-only:
sqlite3 "file:$HOME/.hermes/kanban/boards/<board>/kanban.db?mode=ro" \
  "SELECT id,status,worker_pid,last_heartbeat_at FROM tasks WHERE id='<id>';" \
  "SELECT id,status,substr(summary,1,160) FROM task_runs WHERE task_id='<id>' ORDER BY id DESC LIMIT 5;"

# Push scope before any push:
grep -R 'branches:' .github/workflows

Expected: an OWNER argv line, a fresh last_heartbeat_at, and task_runs showing prior blocked duplicates. Verdict: STAND DOWN — do not create a sixth writer.


5. Pitfalls checklist


Artifacts written: ~/solution/SOLUTION.md, foreman-preflight.sh, foreman-standdown.sh, test-preflight.sh, test-standdown.sh.

Evidence & signatures

# Evidence
- Problem class: scheduler-foreman-stand-down-live-kanban-worker
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-10-05T05:34:59.375Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A coding-hermes SCHEDULER foreman tick fired onto a git worktree (wt/<task>) whose board row was still `pending`, while a LIVE kanban implementation worker owned the same worktree. Symptoms that mislead: the branch was N commits ahead of origin/main and unpushed, the worktree had uncommitted edits, and the board row read pending -- all of which look like an abandoned/stranded tick begging for a re-dispatch. The correct foreman response is a READ-ONLY ZERO-WRITE STAND-DOWN, not a dispatch. Evidence chain that established foreign ownership: (1) scan /proc/<pid>/cwd for every pid and match the worktree path -- found ONE live `hermes ... --skills coding-hermes-worker ... chat -q 'work kanban task <id>'` process whose cwd was the worktree; (2) read the kanban card from ~/.hermes/kanban/boards/<board>/kanban.db (sqlite3 read-only URI `file:<path>?mode=ro`): tasks.worker_pid + last_heartbeat_at + task_events(kind=heartbeat) proved liveness with a fresh heartbeat and worker_pid matching the /proc scan; (3) process CHILDREN proved real progress -- a live `go test -short -p 1 ./internal/scheduler -run <TestName> -count=1` under the worker pid, plus a growing worker log; (4) task_runs history was the decisive artifact: FIVE prior runs of the SAME card ended status=blocked with summaries naming `a live duplicate worker (PID ...) still owns the staged edits in this shared worktree` -- i.e. the duplicate-dispatch defect had already happened repeatedly, so adding a sixth writer was the highest-risk action available. What was done instead: (a) NO dispatch, NO git add/commit, NO board edit, NO gitreins judge run -- each would have either collided with the live writer's index or contended with its in-flight full test suite and flaked its run; (b) push the EXISTING branch ref only (`git push origin HEAD:refs/heads/<wt-branch>`) -- a remote-ref write that touches neither the worktree nor the index, protecting N stranded commits from worktree reaping; (c) verify branch protection/trigger scope first: .github/workflows/* only trigger on `push: branches: [main]`, so pushing the wt branch costs zero CI load; (d) record the stewardship in DuckBrain + the report rather than the git-tracked board, because the worktree's board checkout was BEHIND origin/main (3 commits) and a board write there would have created a board merge conflict on the eventual feature merge. Key pitfalls: `pgrep -af 'hermes chat'` enumerates OTHER lanes' workers and cannot attribute ownership -- use /proc/<pid>/cwd; a paused/stale-looking row plus unpushed commits is NOT proof of abandonment; completed sibling rows on the same board (marked complete for the same failing CI test) show the board can carry duplicates, so verify a CI failure is not already filed before filing another row.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "scheduler-foreman-stand-down-live-kanban-worker", "provider": "openrouter", "solved_at": "2026-10-05T05:34:59.376Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog