Verdict: flake, not a code regression. Two independent defects compound it:
I have a fully reproduced root cause, an applied patch, and passing tests. Here is the solution.
regression suite (masked by continue-on-error) that leaks a half-created agentVerdict: flake, not a code regression. Two independent defects compound it:
regression and root-suite carry continue-on-error: true, so GitHub reports the run as success while the job is red — a run-level green proves nothing about them.userdel/isolation teardown are no-ops and the host is left with a passwd entry, no SSH key, and no registry row.The stall itself is an intermittent live-host issue (rootless-docker install / systemd user-manager wait on bunker-mvp), not a commit regression: commit 358a10c is board-only on top of 0bd45e4, and the same spawn code was green one run earlier.
gh run list reports conclusion=success for run 35084607865 (sha 358a10c). The truth is job-level:
gh run list --limit 5 --json databaseId,headSha,conclusion,status,displayTitle
gh run view 35084607865 --json jobs \
--jq '.jobs[] | {name, conclusion,
failed: [.steps[] | select(.conclusion=="failure") | .name]}'
# run 35084607865 (sha 358a10c) -> job "regression", step "Regression suite": FAILURE
Because .github/workflows/ci.yml sets continue-on-error: true on both self-hosted jobs
(lines ~178 and ~252), the job failure is swallowed at the workflow level:
regression:
continue-on-error: true
root-suite:
continue-on-error: true
Rule for every tick: report job+step truth, never run-level success.
The step's verbatim failure (from gh api repos/<org>/<repo>/actions/jobs/<job_id>/logs;
gh run view --log-failed and --job --log can both return empty on this runner):
##[error]Process completed with exit code 1.
(Regression suite: PASS: 24 / FAIL: 9;
spawn with explicit ID unresponsive for 300.0s)
✗ spawn with explicit ID (... "Agent created: regr-alpha")
✗ returns Docker SSH URL
✗ returns port range
✗ SSH key persisted for regr-alpha ([ -f /etc/bunkerd/ssh/regr-alpha ])
✗ list shows regr-alpha
✓ Linux user created for regr-alpha
⚠ agent metrics: not_found: agent "regr-alpha" not found
Timings are the tell: ── 4. Spawn ── at 10:28:23, explicit-ID check resolved only at 10:33:23 = 300.0 s (the client's full spawn deadline, internal/cli/spawn.go:119), while the very next spawn passed 12 s later.
| RED | CONTROL | |
|---|---|---|
| run | 35084607865 |
35083640714 |
| sha | 358a10c (board-only commit) |
0bd45e4 (parent) |
| job | regression (job 104756644313) |
regression (job 104753557205) |
| explicit-ID spawn | 300.0 s, timeout | 11 s, pass |
| suite | PASS 24 / FAIL 9 | PASS 33 / FAIL 0 |
| battery | SKIPPED | CERTIFIED BINARY … verdict: MATCH |
358a10c differs from 0bd45e4 only in board/doc files (.coding-hermes/board/*, .gitreins/tasks.yaml, internal/agent/SKILL.md). Same code, green then red ⇒ intermittent live-host stall.
internal/agent/manager_spawn.go defines the rollback closure and — before the fix — fed it the live request context:
cleanup := func() {
...
removeUserSliceLimits(ctx, agentID, m.logger) // ctx cancelled => no-op
cmd := exec.CommandContext(ctx, "userdel", "-r", "bunker-"+agentID) // no-op
...
m.removeIsolation(ctx, agentID) // no-op
}
The server request timeout is 300s (/etc/bunkerd/config.yaml: server.request_timeout). When the client deadline fires (or the connection drops), ctx is cancelled. exec.CommandContext with a cancelled context returns immediately without running the command, so the passwd entry created by useradd survives, while the SSH key and registry row never get written. internal/server/service.go then collapsed the error into a bare CodeInternal, so the harness only saw "timeout".
installRootlessDocker (internal/agent/rootless.go:105) runs waitForUserManager (bounded only by the request context — effectively 300 s) and then curl + the rootless installer under the same request context. Concurrent spawns / a stuck user manager / install contention can park there until the 300 s deadline. Startup reconciliation (Reconcile) destroys orphans, but it runs only at daemon start, so a leak created between runs persists until the next restart.
While diagnosing, the exact fingerprint was reproduced live. An authenticated SpawnAgent whose client disconnected after 8 s produced:
audit: /bunker.v1.Bunkerd/SpawnAgent durationMs=10654 outcome=internal agentId=""
/etc/passwd : bunker-2eae301d:x:1001:1001::~:/bin/bash (mtime 12:35:56)
/etc/subuid : bunker-2eae301d:165536:65536
/etc/apparmor.d/home.bunker-2eae301d.bin.rootlesskit
~ : DOES NOT EXIST
/run/user/1001 : DOES NOT EXIST
ListAgents : {} <-- no registry row
Passwd entry present, no SSH key/home, registry empty — the half-created spawn.
internal/agent/manager_spawn.go:
// cleanupTimeout bounds the detached rollback context. Teardown of a single
// agent (userdel -r, systemctl reset-failed, isolation dirs) is fast; the bound
// only exists so a wedged rollback cannot pin the spawn goroutine forever.
const cleanupTimeout = 30 * time.Second
// detachedCleanupContext returns a context that survives cancellation of the
// request context it derives from. Rollback/compensation for a failed or
// timed-out spawn must still run after the client disconnects or its deadline
// fires; using the request ctx directly made userdel a no-op and left
// half-created agents behind (INT-CI-005).
func detachedCleanupContext(ctx context.Context) (context.Context, context.CancelFunc) {
return context.WithTimeout(context.WithoutCancel(ctx), cleanupTimeout)
}
Inside the cleanup closure, derive and use it (cctx), so every compensating command runs even after cancellation:
cleanup := func() {
cctx, cancel := detachedCleanupContext(ctx)
defer cancel()
if portRangeAllocated { m.portAlloc.Free(agentID) }
if createdUserSlice { removeUserSliceLimits(cctx, agentID, m.logger) }
if createdUser {
cmd := exec.CommandContext(cctx, "userdel", "-r", "bunker-"+agentID)
...
}
...
m.removeIsolation(cctx, agentID)
}
internal/server/service.go maps cancellation instead of hiding it:
resp, err := s.agentMgr.Spawn(ctx, req.Msg)
if err != nil {
s.logger.Error("spawn agent failed", "error", err)
switch {
case errors.Is(err, context.DeadlineExceeded):
return nil, connect.NewError(connect.CodeDeadlineExceeded, err)
case errors.Is(err, context.Canceled):
return nil, connect.NewError(connect.CodeCanceled, err)
default:
return nil, connect.NewError(connect.CodeInternal, err)
}
}
Append a gate job to .github/workflows/ci.yml (keeps continue-on-error for an offline runner, but turns a real job failure into a run failure):
ci-job-gate:
name: Job-level CI gate (INT-CI-005)
runs-on: ubuntu-latest
needs: [regression, root-suite]
if: always()
steps:
- name: Fail when a continue-on-error job actually failed
env:
NEEDS: ${{ toJSON(needs) }}
run: |
set -euo pipefail
echo "$NEEDS"
failed=0
for job in regression root-suite; do
result=$(printf '%s' "$NEEDS" | jq -r --arg job "$job" '.[$job].result')
echo "$job -> $result"
case "$result" in
failure|cancelled)
echo "::error::job '$job' concluded '$result' and was masked by continue-on-error=true (INT-CI-005)"
failed=1 ;;
skipped)
echo "::warning::job '$job' was skipped (no self-hosted runner or PR build) — not masking a failure" ;;
esac
done
exit "$failed"
In regression-tests.sh section 4: time the spawn, dump the journal window and leaked host state on failure, and add a killed-spawn self-clean check.
# 4a. Spawn with explicit ID
SPAWN_T0=$(date +%s)
set +e
OUT=$(bunker spawn --agent-id regr-alpha 2>&1)
SPAWN_RC=$?
set -e
SPAWN_ELAPSED=$(( $(date +%s) - SPAWN_T0 ))
if [ "$SPAWN_RC" -ne 0 ]; then
printf 'spawn with explicit ID exited rc=%s after %ss\n' "$SPAWN_RC" "$SPAWN_ELAPSED"
printf '%s\n' "$OUT"
command -v journalctl >/dev/null 2>&1 && {
echo "── daemon journal (last $((SPAWN_ELAPSED + 5))s) ──"
journalctl -u bunkerd --since "-$((SPAWN_ELAPSED + 5))s" --no-pager 2>/dev/null | tail -n 80 || true
}
echo "── leaked host state ──"
getent passwd "bunker-regr-alpha" || echo "no passwd entry"
ls -l /etc/bunkerd/ssh/regr-alpha 2>&1 || true
fi
assert 'echo "$OUT" | grep -q "Agent created: regr-alpha"' "spawn with explicit ID"
# 4f. a spawn whose client is killed must leave no passwd entry / key
if command -v timeout >/dev/null 2>&1; then
timeout 5 bunker spawn --agent-id regr-leak >/dev/null 2>&1 || true
for _ in 1 2 3 4 5 6; do grep -q '^bunker-regr-leak:' /etc/passwd || break; sleep 1; done
assert '! grep -q "^bunker-regr-leak:" /etc/passwd' "timed-out spawn leaves no Linux user"
assert '[ ! -e /etc/bunkerd/ssh/regr-leak ]' "timed-out spawn leaves no SSH key"
fi
Bound the slow, network/user-manager steps so a stall is attributed before the client deadline, and sweep orphans periodically (not only at startup):
waitForUserManager and the rootless install a dedicated deadline, e.g. context.WithTimeout(ctx, cfg.Agent.RootlessInstallTimeout) (default ~180 s), and wrap errors as spawn step "installRootlessDocker": ….AgentManager.Reconcile under destroy-only semantics) so bunker-* users absent from the registry are removed without a daemon restart./tmp/bunker)$ go test ./internal/agent/ -run 'TestDetachedCleanupContext' -v
=== RUN TestDetachedCleanupContext_SurvivesRequestCancellation
--- PASS: TestDetachedCleanupContext_SurvivesRequestCancellation (0.00s)
=== RUN TestDetachedCleanupContext_HasBoundedTimeout
--- PASS: TestDetachedCleanupContext_HasBoundedTimeout (0.00s)
ok github.com/deployBunker/bunker/internal/agent 0.005s
$ go build ./... # OK
$ go vet ./internal/agent/ ./internal/server/ # OK
$ gofmt -l internal/agent/manager_spawn.go internal/server/service.go # (empty)
$ bash -n regression-tests.sh # OK
$ python3 -c 'import yaml; yaml.safe_load(open(".github/workflows/ci.yml"))' # OK (jobs: build-and-test, gitreins-guard, regression, root-suite, ci-job-gate)
The one failing test in the package, TestApplyUserSliceLimits_NotRoot_Coverage, fails identically on a stashed clean tree (sandbox /etc is read-only) — pre-existing and unrelated.
The new test proves the mechanism directly: exec.CommandContext on a cancelled context fails; the context returned by detachedCleanupContext still executes.
# Start a spawn, kill the client at 5s (cancels the request ctx like a 300s
# deadline or a dropped connection does), then wait for daemon rollback.
timeout 5 bunker spawn --agent-id regr-leak || true
sleep 5
getent passwd bunker-regr-leak && echo "LEAK" || echo "clean"
[ -e /etc/bunkerd/ssh/regr-leak ] && echo "LEAK" || echo "clean"
bunker list | grep -q regr-leak && echo "LEAK" || echo "clean"
Expected after the fix: clean / clean / clean. Before the fix, getent returns the passwd entry (as reproduced with bunker-2eae301d).
regression step spawn with explicit ID is green for 10 consecutive pushes to main./etc/bunkerd/ssh/<id> key, and no partial registry row.CodeDeadlineExceeded: spawn step "installRootlessDocker": context deadline exceeded), not a bare 300.0s timeout, and the run-level conclusion is red whenever the job is red.Runbook for the next recurrence if it still stalls:
journalctl -u bunkerd --since '2026-09-16 10:28' --until '2026-09-16 10:34' --no-pager
grep -c '^bunker-' /etc/passwd # orphan users
bunker list # registry view
ls -l /etc/bunkerd/ssh/ # key vs passwd mismatch
Apply the following (produced and built at /tmp/bunker; git diff):
diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml
--- a/.github/workflows/ci.yml
+++ b/.github/workflows/ci.yml
@@ -276,3 +276,40 @@ jobs:
run: bash scripts/root-suite.sh
timeout-minutes: 15
+
+ # ── Tier 2: job-level truth gate (INT-CI-005) ────────────────────
+ # `regression` and `root-suite` carry continue-on-error: true so a missing
+ # self-hosted runner cannot block the workflow. The cost is that GitHub
+ # reports the RUN as success even when the job is RED: run 35084607865
+ # showed conclusion=success while job `regression` / step `Regression suite`
+ # was FAILURE (24 PASS / 9 FAIL). This gate converts a real job-level
+ # failure back into a run-level failure. `skipped` stays non-fatal so PR
+ # builds and an offline runner do not break the build; `cancelled` is fatal.
+ ci-job-gate:
+ name: Job-level CI gate (INT-CI-005)
+ runs-on: ubuntu-latest
+ needs: [regression, root-suite]
+ if: always()
+ steps:
+ - name: Fail when a continue-on-error job actually failed
+ env:
+ NEEDS: ${{ toJSON(needs) }}
+ run: |
+ set -euo pipefail
+ echo "$NEEDS"
+ failed=0
+ for job in regression root-suite; do
+ result=$(printf '%s' "$NEEDS" | jq -r --arg job "$job" '.[$job].result')
+ echo "$job -> $result"
+ case "$result" in
+ failure|cancelled)
+ echo "::error::job '$job' concluded '$result' and was masked by continue-on-error=true (INT-CI-005)"
+ failed=1
+ ;;
+ skipped)
+ echo "::warning::job '$job' was skipped (no self-hosted runner or PR build) — not masking a failure"
+ ;;
+ esac
+ done
+ exit "$failed"
diff --git a/internal/agent/manager_spawn.go b/internal/agent/manager_spawn.go
--- a/internal/agent/manager_spawn.go
+++ b/internal/agent/manager_spawn.go
@@ -20,6 +20,20 @@ import (
"github.com/deployBunker/bunker/internal/resource"
)
+// cleanupTimeout bounds the detached rollback context. Teardown of a single
+// agent (userdel -r, systemctl reset-failed, isolation dirs) is fast; the bound
+// only exists so a wedged rollback cannot pin the spawn goroutine forever.
+const cleanupTimeout = 30 * time.Second
+
+// detachedCleanupContext returns a context that survives cancellation of the
+// request context it derives from. Rollback/compensation for a failed or
+// timed-out spawn must still run after the client disconnects or its deadline
+// fires; using the request ctx directly made userdel a no-op and left
+// half-created agents behind (INT-CI-005).
+func detachedCleanupContext(ctx context.Context) (context.Context, context.CancelFunc) {
+ return context.WithTimeout(context.WithoutCancel(ctx), cleanupTimeout)
+}
+
func (m *AgentManager) Spawn(ctx context.Context, req *v1.SpawnAgentRequest) (*v1.SpawnAgentResponse, error) {
@@ -102,17 +116,26 @@ func (m *AgentManager) Spawn(ctx context.Context, req *v1.SpawnAgentRequest) (*v
portRangeAllocated := m.portAlloc != nil
cleanup := func() {
+ // INT-CI-005: rollback MUST NOT inherit the request context. A client
+ // timeout / disconnect cancels ctx, which makes every compensating
+ // command (userdel, systemctl reset-failed, isolation teardown) return
+ // immediately without running — leaving a half-created agent: a passwd
+ // entry with no SSH key and no registry row. Derive a short-lived
+ // detached context so cleanup always executes.
+ cctx, cancel := detachedCleanupContext(ctx)
+ defer cancel()
+
// Free the port range first — the in-memory allocator leaks
// permanently if a failed spawn never releases it.
if portRangeAllocated {
m.portAlloc.Free(agentID)
}
if createdUserSlice {
- removeUserSliceLimits(ctx, agentID, m.logger)
+ removeUserSliceLimits(cctx, agentID, m.logger)
}
if createdUser {
m.logger.Warn("rolling back: removing user", "username", "bunker-"+agentID)
- cmd := exec.CommandContext(ctx, "userdel", "-r", "bunker-"+agentID)
+ cmd := exec.CommandContext(cctx, "userdel", "-r", "bunker-"+agentID)
if out, err := cmd.CombinedOutput(); err != nil {
@@ -131,7 +154,7 @@ func (m *AgentManager) Spawn(ctx context.Context, req *v1.SpawnAgentRequest) (*v
// GAP-075: drop the scratch and private-/tmp instance directories so a
// failed spawn cannot leave a provisioned exchange point or tmp
// instance behind for an agent that does not exist.
- m.removeIsolation(ctx, agentID)
+ m.removeIsolation(cctx, agentID)
}
// ── Step 2: Create Linux user ──────────────────────────────────
diff --git a/internal/server/service.go b/internal/server/service.go
--- a/internal/server/service.go
+++ b/internal/server/service.go
@@ -3,6 +3,7 @@ package server
import (
"context"
+ "errors"
"fmt"
"log/slog"
"os"
@@ -174,7 +175,19 @@ func (s *bunkerdService) SpawnAgent(ctx context.Context, req *connect.Request[v1
resp, err := s.agentMgr.Spawn(ctx, req.Msg)
if err != nil {
s.logger.Error("spawn agent failed", "error", err)
- return nil, connect.NewError(connect.CodeInternal, err)
+ // INT-CI-005: do not collapse a client timeout / disconnect into a bare
+ // CodeInternal. The CLI (and the regression harness) needs to tell a
+ // real daemon fault from "the spawn outlived its 300s deadline".
+ switch {
+ case errors.Is(err, context.DeadlineExceeded):
+ return nil, connect.NewError(connect.CodeDeadlineExceeded, err)
+ case errors.Is(err, context.Canceled):
+ return nil, connect.NewError(connect.CodeCanceled, err)
+ default:
+ return nil, connect.NewError(connect.CodeInternal, err)
+ }
}
// Generate an agent-scoped opaque API sub-key when JWT auth is enabled.
diff --git a/internal/agent/manager_spawn_rollback_test.go b/internal/agent/manager_spawn_rollback_test.go
new file mode 100644
--- /dev/null
+++ b/internal/agent/manager_spawn_rollback_test.go
@@ -0,0 +1,43 @@
+package agent
+
+import (
+ "context"
+ "os/exec"
+ "testing"
+ "time"
+)
+
+// TestDetachedCleanupContext_SurvivesRequestCancellation pins INT-CI-005:
+// rollback commands must run even when the spawn's request context is
+// cancelled or deadline-exceeded. Before the fix, the cleanup closure fed the
+// live request ctx to exec.CommandContext, so a client timeout turned userdel
+// into a no-op and left a passwd entry with no SSH key and no registry row.
+func TestDetachedCleanupContext_SurvivesRequestCancellation(t *testing.T) {
+ parent, cancel := context.WithCancel(context.Background())
+ cancel() // simulate client disconnect / the 300s deadline firing
+
+ // Baseline: the old behaviour — the request ctx is already dead, so the
+ // compensating command never runs.
+ if err := exec.CommandContext(parent, "true").Run(); err == nil {
+ t.Fatal("sanity: a command on a cancelled context unexpectedly succeeded")
+ }
+
+ // Fixed behaviour: the detached rollback ctx still executes commands.
+ cctx, ccancel := detachedCleanupContext(parent)
+ defer ccancel()
+ if err := exec.CommandContext(cctx, "true").Run(); err != nil {
+ t.Fatalf("detached cleanup context must run rollback commands, got: %v", err)
+ }
+}
+
+func TestDetachedCleanupContext_HasBoundedTimeout(t *testing.T) {
+ cctx, cancel := detachedCleanupContext(context.Background())
+ defer cancel()
+ dl, ok := cctx.Deadline()
+ if !ok {
+ t.Fatal("detached cleanup context must carry a deadline so a wedged rollback cannot hang spawn")
+ }
+ if d := time.Until(dl); d <= 0 || d > cleanupTimeout {
+ t.Fatalf("detached cleanup deadline out of range: %v (cleanupTimeout=%v)", d, cleanupTimeout)
+ }
+}
diff --git a/regression-tests.sh b/regression-tests.sh
--- a/regression-tests.sh
+++ b/regression-tests.sh
@@ -167,7 +167,26 @@ echo ""
# 4a. Spawn with explicit ID
+# INT-CI-005: time the spawn and, on failure, print the daemon-side reason
+# (journal window) plus the leaked host state — a bare 300s timeout is not
+# attributable on its own.
+SPAWN_T0=$(date +%s)
+set +e
OUT=$(bunker spawn --agent-id regr-alpha 2>&1)
+SPAWN_RC=$?
+set -e
+SPAWN_ELAPSED=$(( $(date +%s) - SPAWN_T0 ))
+if [ "$SPAWN_RC" -ne 0 ]; then
+ printf 'spawn with explicit ID exited rc=%s after %ss\n' "$SPAWN_RC" "$SPAWN_ELAPSED"
+ printf '%s\n' "$OUT"
+ if command -v journalctl >/dev/null 2>&1; then
+ echo "── daemon journal (last $((SPAWN_ELAPSED + 5))s) ──"
+ journalctl -u bunkerd --since "-$((SPAWN_ELAPSED + 5))s" --no-pager 2>/dev/null | tail -n 80 || true
+ fi
+ echo "── leaked host state ──"
+ getent passwd "bunker-regr-alpha" || echo "no passwd entry"
+ ls -l /etc/bunkerd/ssh/regr-alpha 2>&1 || true
+fi
assert 'echo "$OUT" | grep -q "Agent created: regr-alpha"' "spawn with explicit ID"
@@ -202,6 +221,22 @@ assert '[ -f "$AUTH_KEYS" ]' "authorized_keys exists"
assert 'grep -q "ssh-ed25519" '"$AUTH_KEYS" "authorized_keys has ssh-ed25519 key"
assert 'grep -q "DOCKER_HOST=" '"$AUTH_KEYS" "authorized_keys has DOCKER_HOST environment"
+# 4f. INT-CI-005 self-clean check: a spawn whose client is killed mid-flight
+# (here: `timeout 5`, which cancels the request context the same way a 300s
+# deadline or a dropped connection does) must leave the host fully clean —
+# no passwd entry, no SSH key, no registry row.
+if command -v timeout >/dev/null 2>&1; then
+ timeout 5 bunker spawn --agent-id regr-leak >/dev/null 2>&1 || true
+ for _ in 1 2 3 4 5 6; do
+ grep -q '^bunker-regr-leak:' /etc/passwd || break
+ sleep 1
+ done
+ assert '! grep -q "^bunker-regr-leak:" /etc/passwd' "timed-out spawn leaves no Linux user"
+ assert '[ ! -e /etc/bunkerd/ssh/regr-leak ]' "timed-out spawn leaves no SSH key"
+fi
+
echo ""
One-line summary: job-level red is real and hidden by continue-on-error; the red is an intermittent spawn stall whose cancellation makes rollback a no-op; run rollback under context.WithoutCancel, gate the workflow on job results, and make the harness report the daemon reason and assert no leaked agent.
# Evidence - Problem class: ci-job-level-red-masked-by-continue-on-error - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-16T12:41:07.272Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: `gh run list` shows conclusion=success for every recent run, yet the newest completed run's `regression` job is RED. On this repo (deployBunker/bunker) jobs `root-suite` and `regression` carry continue-on-error: true, so a run-level green proves NOTHING about them.\n\nATTRIBUTION PROCEDURE (what actually settled it): (1) `gh run list --limit 5 --json databaseId,headSha,conclusion,status,displayTitle`; (2) `gh run view <id> --json jobs --jq '.jobs[] | {name, conclusion, failed: [.steps[]|select(.conclusion==\"failure\")|.name]}'` -> run 35084607865 (sha 358a10c) job `regression`, step `Regression suite`, FAILURE; (3) get the step's real log with `gh api repos/<org>/<repo>/actions/jobs/<job_id>/logs` because `gh run view --log-failed` and `gh run view --job <id> --log` can both return EMPTY on this runner; (4) read TIMINGS, not just PASS/FAIL: the harness printed '\u2500\u2500 4. Spawn \u2500\u2500' at 10:28:23 and resolved 'spawn with explicit ID' only at 10:33:23 = 300.0s, i.e. the client's full spawn deadline, while the next check (auto-generated ID) passed 12s later; (5) find the CONTROL run: 358a10c is a board-only commit on top of 0bd45e4, and 0bd45e4's run 35083640714 executed the SAME code with the SAME harness green (explicit-ID spawn 11s, 'PASS: 33 / FAIL: 0', battery step 'CERTIFIED BINARY ... verdict: MATCH'). Same code, green one run and 9 red checks the next => intermittent live-host stall, NOT a code regression introduced by a commit. VERIFY the control by job id, not by run order: the run that contains a step named the same as the failing one, on a commit whose only diff is (board/docs) files.\n\nHALF-CREATED RESOURCE FINGERPRINT: a spawn that stalls past its deadline leaves the host in a partial state \u2014 '\u2713 Linux user created for regr-alpha' passed while '\u2717 SSH key persisted for regr-alpha ([ -f /etc/bunkerd/ssh/regr-alpha ])' failed and every downstream read said not_found ('agent metrics: not_found: agent \"regr-alpha\" not found'). So the leak is a passwd entry with no SSH key and no registry row \u2014 the fix must make a timed-out spawn either complete or self-clean, and the harness must print the daemon-side reason instead of a bare timeout.\n\nFIX/OUTCOME: filed as a P1 board row (INT-CI-005) rather than hand-patched, because the failure is a live-host flake on the CI runner (bunker-mvp itself) and needs the daemon journal at the stalled window (`journalctl -u bunkerd --since ... --until ...`) plus 10 consecutive green spawn-with-explicit-ID runs as the exit criterion. Report the job-level truth (job+step+verbatim failing checks+timings) in the tick output, never the run-level 'success'.", "environment": "GitHub Actions self-hosted runner 'bunker-mvp' (root), VERSIONLESS go 1.26.5, BUNKERD_COEXIST=1 on :29090/:28080", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-job-level-red-masked-by-continue-on-error", "provider": "openrouter", "solved_at": "2026-09-16T12:41:07.272Z", "version": ""}I have a fully reproduced root cause, an applied patch, and passing tests. Here is the solution.
regression suite (masked by continue-on-error) that leaks a half-created agentVerdict: flake, not a code regression. Two independent defects compound it:
regression and root-suite carry continue-on-error: true, so GitHub reports the run as success while the job is red — a run-level green proves nothing about them.userdel/isolation teardown are no-ops and the host is left with a passwd entry, no SSH key, and no registry row.The stall itself is an intermittent live-host issue (rootless-docker install / systemd user-manager wait on bunker-mvp), not a commit regression: commit 358a10c is board-only on top of 0bd45e4, and the same spawn code was green one run earlier.
gh run list reports conclusion=success for run 35084607865 (sha 358a10c). The truth is job-level:
gh run list --limit 5 --json databaseId,headSha,conclusion,status,displayTitle
gh run view 35084607865 --json jobs \
--jq '.jobs[] | {name, conclusion,
failed: [.steps[] | select(.conclusion=="failure") | .name]}'
# run 35084607865 (sha 358a10c) -> job "regression", step "Regression suite": FAILURE
Because .github/workflows/ci.yml sets continue-on-error: true on both self-hosted jobs
(lines ~178 and ~252), the job failure is swallowed at the workflow level:
regression:
continue-on-error: true
root-suite:
continue-on-error: true
Rule for every tick: report job+step truth, never run-level success.
The step's verbatim failure (from gh api repos/<org>/<repo>/actions/jobs/<job_id>/logs;
gh run view --log-failed and --job --log can both return empty on this runner):
##[error]Process completed with exit code 1.
(Regression suite: PASS: 24 / FAIL: 9;
spawn with explicit ID unresponsive for 300.0s)
✗ spawn with explicit ID (... "Agent created: regr-alpha")
✗ returns Docker SSH URL
✗ returns port range
✗ SSH key persisted for regr-alpha ([ -f /etc/bunkerd/ssh/regr-alpha ])
✗ list shows regr-alpha
✓ Linux user created for regr-alpha
⚠ agent metrics: not_found: agent "regr-alpha" not found
Timings are the tell: ── 4. Spawn ── at 10:28:23, explicit-ID check resolved only at 10:33:23 = 300.0 s (the client's full spawn deadline, internal/cli/spawn.go:119), while the very next spawn passed 12 s later.
| RED | CONTROL | |
|---|---|---|
| run | 35084607865 |
35083640714 |
| sha | 358a10c (board-only commit) |
0bd45e4 (parent) |
| job | regression (job 104756644313) |
regression (job 104753557205) |
| explicit-ID spawn | 300.0 s, timeout | 11 s, pass |
| suite | PASS 24 / FAIL 9 | PASS 33 / FAIL 0 |
| battery | SKIPPED | CERTIFIED BINARY … verdict: MATCH |
358a10c differs from 0bd45e4 only in board/doc files (.coding-hermes/board/*, .gitreins/tasks.yaml, internal/agent/SKILL.md). Same code, green then red ⇒ intermittent live-host stall.
internal/agent/manager_spawn.go defines the rollback closure and — before the fix — fed it the live request context:
cleanup := func() {
...
removeUserSliceLimits(ctx, agentID, m.logger) // ctx cancelled => no-op
cmd := exec.CommandContext(ctx, "userdel", "-r", "bunker-"+agentID) // no-op
...
m.removeIsolation(ctx, agentID) // no-op
}
The server request timeout is 300s (/etc/bunkerd/config.yaml: server.request_timeout). When the client deadline fires (or the connection drops), ctx is cancelled. exec.CommandContext with a cancelled context returns immediately without running the command, so the passwd entry created by useradd survives, while the SSH key and registry row never get written. internal/server/service.go then collapsed the error into a bare CodeInternal, so the harness only saw "timeout".
installRootlessDocker (internal/agent/rootless.go:105) runs waitForUserManager (bounded only by the request context — effectively 300 s) and then curl + the rootless installer under the same request context. Concurrent spawns / a stuck user manager / install contention can park there until the 300 s deadline. Startup reconciliation (Reconcile) destroys orphans, but it runs only at daemon start, so a leak created between runs persists until the next restart.
While diagnosing, the exact fingerprint was reproduced live. An authenticated SpawnAgent whose client disconnected after 8 s produced:
audit: /bunker.v1.Bunkerd/SpawnAgent durationMs=10654 outcome=internal agentId=""
/etc/passwd : bunker-2eae301d:x:1001:1001::~:/bin/bash (mtime 12:35:56)
/etc/subuid : bunker-2eae301d:165536:65536
/etc/apparmor.d/home.bunker-2eae301d.bin.rootlesskit
~ : DOES NOT EXIST
/run/user/1001 : DOES NOT EXIST
ListAgents : {} <-- no registry row
Passwd entry present, no SSH key/home, registry empty — the half-created spawn.
internal/agent/manager_spawn.go:
// cleanupTimeout bounds the detached rollback context. Teardown of a single
// agent (userdel -r, systemctl reset-failed, isolation dirs) is fast; the bound
// only exists so a wedged rollback cannot pin the spawn goroutine forever.
const cleanupTimeout = 30 * time.Second
// detachedCleanupContext returns a context that survives cancellation of the
// request context it derives from. Rollback/compensation for a failed or
// timed-out spawn must still run after the client disconnects or its deadline
// fires; using the request ctx directly made userdel a no-op and left
// half-created agents behind (INT-CI-005).
func detachedCleanupContext(ctx context.Context) (context.Context, context.CancelFunc) {
return context.WithTimeout(context.WithoutCancel(ctx), cleanupTimeout)
}
Inside the cleanup closure, derive and use it (cctx), so every compensating command runs even after cancellation:
cleanup := func() {
cctx, cancel := detachedCleanupContext(ctx)
defer cancel()
if portRangeAllocated { m.portAlloc.Free(agentID) }
if createdUserSlice { removeUserSliceLimits(cctx, agentID, m.logger) }
if createdUser {
cmd := exec.CommandContext(cctx, "userdel", "-r", "bunker-"+agentID)
...
}
...
m.removeIsolation(cctx, agentID)
}
internal/server/service.go maps cancellation instead of hiding it:
resp, err := s.agentMgr.Spawn(ctx, req.Msg)
if err != nil {
s.logger.Error("spawn agent failed", "error", err)
switch {
case errors.Is(err, context.DeadlineExceeded):
return nil, connect.NewError(connect.CodeDeadlineExceeded, err)
case errors.Is(err, context.Canceled):
return nil, connect.NewError(connect.CodeCanceled, err)
default:
return nil, connect.NewError(connect.CodeInternal, err)
}
}
Append a gate job to .github/workflows/ci.yml (keeps continue-on-error for an offline runner, but turns a real job failure into a run failure):
ci-job-gate:
name: Job-level CI gate (INT-CI-005)
runs-on: ubuntu-latest
needs: [regression, root-suite]
if: always()
steps:
- name: Fail when a continue-on-error job actually failed
env:
NEEDS: ${{ toJSON(needs) }}
run: |
set -euo pipefail
echo "$NEEDS"
failed=0
for job in regression root-suite; do
result=$(printf '%s' "$NEEDS" | jq -r --arg job "$job" '.[$job].result')
echo "$job -> $result"
case "$result" in
failure|cancelled)
echo "::error::job '$job' concluded '$result' and was masked by continue-on-error=true (INT-CI-005)"
failed=1 ;;
skipped)
echo "::warning::job '$job' was skipped (no self-hosted runner or PR build) — not masking a failure" ;;
esac
done
exit "$failed"
In regression-tests.sh section 4: time the spawn, dump the journal window and leaked host state on failure, and add a killed-spawn self-clean check.
# 4a. Spawn with explicit ID
SPAWN_T0=$(date +%s)
set +e
OUT=$(bunker spawn --agent-id regr-alpha 2>&1)
SPAWN_RC=$?
set -e
SPAWN_ELAPSED=$(( $(date +%s) - SPAWN_T0 ))
if [ "$SPAWN_RC" -ne 0 ]; then
printf 'spawn with explicit ID exited rc=%s after %ss\n' "$SPAWN_RC" "$SPAWN_ELAPSED"
printf '%s\n' "$OUT"
command -v journalctl >/dev/null 2>&1 && {
echo "── daemon journal (last $((SPAWN_ELAPSED + 5))s) ──"
journalctl -u bunkerd --since "-$((SPAWN_ELAPSED + 5))s" --no-pager 2>/dev/null | tail -n 80 || true
}
echo "── leaked host state ──"
getent passwd "bunker-regr-alpha" || echo "no passwd entry"
ls -l /etc/bunkerd/ssh/regr-alpha 2>&1 || true
fi
assert 'echo "$OUT" | grep -q "Agent created: regr-alpha"' "spawn with explicit ID"
# 4f. a spawn whose client is killed must leave no passwd entry / key
if command -v timeout >/dev/null 2>&1; then
timeout 5 bunker spawn --agent-id regr-leak >/dev/null 2>&1 || true
for _ in 1 2 3 4 5 6; do grep -q '^bunker-regr-leak:' /etc/passwd || break; sleep 1; done
assert '! grep -q "^bunker-regr-leak:" /etc/passwd' "timed-out spawn leaves no Linux user"
assert '[ ! -e /etc/bunkerd/ssh/regr-leak ]' "timed-out spawn leaves no SSH key"
fi
Bound the slow, network/user-manager steps so a stall is attributed before the client deadline, and sweep orphans periodically (not only at startup):
waitForUserManager and the rootless install a dedicated deadline, e.g. context.WithTimeout(ctx, cfg.Agent.RootlessInstallTimeout) (default ~180 s), and wrap errors as spawn step "installRootlessDocker": ….AgentManager.Reconcile under destroy-only semantics) so bunker-* users absent from the registry are removed without a daemon restart./tmp/bunker)$ go test ./internal/agent/ -run 'TestDetachedCleanupContext' -v
=== RUN TestDetachedCleanupContext_SurvivesRequestCancellation
--- PASS: TestDetachedCleanupContext_SurvivesRequestCancellation (0.00s)
=== RUN TestDetachedCleanupContext_HasBoundedTimeout
--- PASS: TestDetachedCleanupContext_HasBoundedTimeout (0.00s)
ok github.com/deployBunker/bunker/internal/agent 0.005s
$ go build ./... # OK
$ go vet ./internal/agent/ ./internal/server/ # OK
$ gofmt -l internal/agent/manager_spawn.go internal/server/service.go # (empty)
$ bash -n regression-tests.sh # OK
$ python3 -c 'import yaml; yaml.safe_load(open(".github/workflows/ci.yml"))' # OK (jobs: build-and-test, gitreins-guard, regression, root-suite, ci-job-gate)
The one failing test in the package, TestApplyUserSliceLimits_NotRoot_Coverage, fails identically on a stashed clean tree (sandbox /etc is read-only) — pre-existing and unrelated.
The new test proves the mechanism directly: exec.CommandContext on a cancelled context fails; the context returned by detachedCleanupContext still executes.
# Start a spawn, kill the client at 5s (cancels the request ctx like a 300s
# deadline or a dropped connection does), then wait for daemon rollback.
timeout 5 bunker spawn --agent-id regr-leak || true
sleep 5
getent passwd bunker-regr-leak && echo "LEAK" || echo "clean"
[ -e /etc/bunkerd/ssh/regr-leak ] && echo "LEAK" || echo "clean"
bunker list | grep -q regr-leak && echo "LEAK" || echo "clean"
Expected after the fix: clean / clean / clean. Before the fix, getent returns the passwd entry (as reproduced with bunker-2eae301d).
regression step spawn with explicit ID is green for 10 consecutive pushes to main./etc/bunkerd/ssh/<id> key, and no partial registry row.CodeDeadlineExceeded: spawn step "installRootlessDocker": context deadline exceeded), not a bare 300.0s timeout, and the run-level conclusion is red whenever the job is red.Runbook for the next recurrence if it still stalls:
journalctl -u bunkerd --since '2026-09-16 10:28' --until '2026-09-16 10:34' --no-pager
grep -c '^bunker-' /etc/passwd # orphan users
bunker list # registry view
ls -l /etc/bunkerd/ssh/ # key vs passwd mismatch
Apply the following (produced and built at /tmp/bunker; git diff):
diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml
--- a/.github/workflows/ci.yml
+++ b/.github/workflows/ci.yml
@@ -276,3 +276,40 @@ jobs:
run: bash scripts/root-suite.sh
timeout-minutes: 15
+
+ # ── Tier 2: job-level truth gate (INT-CI-005) ────────────────────
+ # `regression` and `root-suite` carry continue-on-error: true so a missing
+ # self-hosted runner cannot block the workflow. The cost is that GitHub
+ # reports the RUN as success even when the job is RED: run 35084607865
+ # showed conclusion=success while job `regression` / step `Regression suite`
+ # was FAILURE (24 PASS / 9 FAIL). This gate converts a real job-level
+ # failure back into a run-level failure. `skipped` stays non-fatal so PR
+ # builds and an offline runner do not break the build; `cancelled` is fatal.
+ ci-job-gate:
+ name: Job-level CI gate (INT-CI-005)
+ runs-on: ubuntu-latest
+ needs: [regression, root-suite]
+ if: always()
+ steps:
+ - name: Fail when a continue-on-error job actually failed
+ env:
+ NEEDS: ${{ toJSON(needs) }}
+ run: |
+ set -euo pipefail
+ echo "$NEEDS"
+ failed=0
+ for job in regression root-suite; do
+ result=$(printf '%s' "$NEEDS" | jq -r --arg job "$job" '.[$job].result')
+ echo "$job -> $result"
+ case "$result" in
+ failure|cancelled)
+ echo "::error::job '$job' concluded '$result' and was masked by continue-on-error=true (INT-CI-005)"
+ failed=1
+ ;;
+ skipped)
+ echo "::warning::job '$job' was skipped (no self-hosted runner or PR build) — not masking a failure"
+ ;;
+ esac
+ done
+ exit "$failed"
diff --git a/internal/agent/manager_spawn.go b/internal/agent/manager_spawn.go
--- a/internal/agent/manager_spawn.go
+++ b/internal/agent/manager_spawn.go
@@ -20,6 +20,20 @@ import (
"github.com/deployBunker/bunker/internal/resource"
)
+// cleanupTimeout bounds the detached rollback context. Teardown of a single
+// agent (userdel -r, systemctl reset-failed, isolation dirs) is fast; the bound
+// only exists so a wedged rollback cannot pin the spawn goroutine forever.
+const cleanupTimeout = 30 * time.Second
+
+// detachedCleanupContext returns a context that survives cancellation of the
+// request context it derives from. Rollback/compensation for a failed or
+// timed-out spawn must still run after the client disconnects or its deadline
+// fires; using the request ctx directly made userdel a no-op and left
+// half-created agents behind (INT-CI-005).
+func detachedCleanupContext(ctx context.Context) (context.Context, context.CancelFunc) {
+ return context.WithTimeout(context.WithoutCancel(ctx), cleanupTimeout)
+}
+
func (m *AgentManager) Spawn(ctx context.Context, req *v1.SpawnAgentRequest) (*v1.SpawnAgentResponse, error) {
@@ -102,17 +116,26 @@ func (m *AgentManager) Spawn(ctx context.Context, req *v1.SpawnAgentRequest) (*v
portRangeAllocated := m.portAlloc != nil
cleanup := func() {
+ // INT-CI-005: rollback MUST NOT inherit the request context. A client
+ // timeout / disconnect cancels ctx, which makes every compensating
+ // command (userdel, systemctl reset-failed, isolation teardown) return
+ // immediately without running — leaving a half-created agent: a passwd
+ // entry with no SSH key and no registry row. Derive a short-lived
+ // detached context so cleanup always executes.
+ cctx, cancel := detachedCleanupContext(ctx)
+ defer cancel()
+
// Free the port range first — the in-memory allocator leaks
// permanently if a failed spawn never releases it.
if portRangeAllocated {
m.portAlloc.Free(agentID)
}
if createdUserSlice {
- removeUserSliceLimits(ctx, agentID, m.logger)
+ removeUserSliceLimits(cctx, agentID, m.logger)
}
if createdUser {
m.logger.Warn("rolling back: removing user", "username", "bunker-"+agentID)
- cmd := exec.CommandContext(ctx, "userdel", "-r", "bunker-"+agentID)
+ cmd := exec.CommandContext(cctx, "userdel", "-r", "bunker-"+agentID)
if out, err := cmd.CombinedOutput(); err != nil {
@@ -131,7 +154,7 @@ func (m *AgentManager) Spawn(ctx context.Context, req *v1.SpawnAgentRequest) (*v
// GAP-075: drop the scratch and private-/tmp instance directories so a
// failed spawn cannot leave a provisioned exchange point or tmp
// instance behind for an agent that does not exist.
- m.removeIsolation(ctx, agentID)
+ m.removeIsolation(cctx, agentID)
}
// ── Step 2: Create Linux user ──────────────────────────────────
diff --git a/internal/server/service.go b/internal/server/service.go
--- a/internal/server/service.go
+++ b/internal/server/service.go
@@ -3,6 +3,7 @@ package server
import (
"context"
+ "errors"
"fmt"
"log/slog"
"os"
@@ -174,7 +175,19 @@ func (s *bunkerdService) SpawnAgent(ctx context.Context, req *connect.Request[v1
resp, err := s.agentMgr.Spawn(ctx, req.Msg)
if err != nil {
s.logger.Error("spawn agent failed", "error", err)
- return nil, connect.NewError(connect.CodeInternal, err)
+ // INT-CI-005: do not collapse a client timeout / disconnect into a bare
+ // CodeInternal. The CLI (and the regression harness) needs to tell a
+ // real daemon fault from "the spawn outlived its 300s deadline".
+ switch {
+ case errors.Is(err, context.DeadlineExceeded):
+ return nil, connect.NewError(connect.CodeDeadlineExceeded, err)
+ case errors.Is(err, context.Canceled):
+ return nil, connect.NewError(connect.CodeCanceled, err)
+ default:
+ return nil, connect.NewError(connect.CodeInternal, err)
+ }
}
// Generate an agent-scoped opaque API sub-key when JWT auth is enabled.
diff --git a/internal/agent/manager_spawn_rollback_test.go b/internal/agent/manager_spawn_rollback_test.go
new file mode 100644
--- /dev/null
+++ b/internal/agent/manager_spawn_rollback_test.go
@@ -0,0 +1,43 @@
+package agent
+
+import (
+ "context"
+ "os/exec"
+ "testing"
+ "time"
+)
+
+// TestDetachedCleanupContext_SurvivesRequestCancellation pins INT-CI-005:
+// rollback commands must run even when the spawn's request context is
+// cancelled or deadline-exceeded. Before the fix, the cleanup closure fed the
+// live request ctx to exec.CommandContext, so a client timeout turned userdel
+// into a no-op and left a passwd entry with no SSH key and no registry row.
+func TestDetachedCleanupContext_SurvivesRequestCancellation(t *testing.T) {
+ parent, cancel := context.WithCancel(context.Background())
+ cancel() // simulate client disconnect / the 300s deadline firing
+
+ // Baseline: the old behaviour — the request ctx is already dead, so the
+ // compensating command never runs.
+ if err := exec.CommandContext(parent, "true").Run(); err == nil {
+ t.Fatal("sanity: a command on a cancelled context unexpectedly succeeded")
+ }
+
+ // Fixed behaviour: the detached rollback ctx still executes commands.
+ cctx, ccancel := detachedCleanupContext(parent)
+ defer ccancel()
+ if err := exec.CommandContext(cctx, "true").Run(); err != nil {
+ t.Fatalf("detached cleanup context must run rollback commands, got: %v", err)
+ }
+}
+
+func TestDetachedCleanupContext_HasBoundedTimeout(t *testing.T) {
+ cctx, cancel := detachedCleanupContext(context.Background())
+ defer cancel()
+ dl, ok := cctx.Deadline()
+ if !ok {
+ t.Fatal("detached cleanup context must carry a deadline so a wedged rollback cannot hang spawn")
+ }
+ if d := time.Until(dl); d <= 0 || d > cleanupTimeout {
+ t.Fatalf("detached cleanup deadline out of range: %v (cleanupTimeout=%v)", d, cleanupTimeout)
+ }
+}
diff --git a/regression-tests.sh b/regression-tests.sh
--- a/regression-tests.sh
+++ b/regression-tests.sh
@@ -167,7 +167,26 @@ echo ""
# 4a. Spawn with explicit ID
+# INT-CI-005: time the spawn and, on failure, print the daemon-side reason
+# (journal window) plus the leaked host state — a bare 300s timeout is not
+# attributable on its own.
+SPAWN_T0=$(date +%s)
+set +e
OUT=$(bunker spawn --agent-id regr-alpha 2>&1)
+SPAWN_RC=$?
+set -e
+SPAWN_ELAPSED=$(( $(date +%s) - SPAWN_T0 ))
+if [ "$SPAWN_RC" -ne 0 ]; then
+ printf 'spawn with explicit ID exited rc=%s after %ss\n' "$SPAWN_RC" "$SPAWN_ELAPSED"
+ printf '%s\n' "$OUT"
+ if command -v journalctl >/dev/null 2>&1; then
+ echo "── daemon journal (last $((SPAWN_ELAPSED + 5))s) ──"
+ journalctl -u bunkerd --since "-$((SPAWN_ELAPSED + 5))s" --no-pager 2>/dev/null | tail -n 80 || true
+ fi
+ echo "── leaked host state ──"
+ getent passwd "bunker-regr-alpha" || echo "no passwd entry"
+ ls -l /etc/bunkerd/ssh/regr-alpha 2>&1 || true
+fi
assert 'echo "$OUT" | grep -q "Agent created: regr-alpha"' "spawn with explicit ID"
@@ -202,6 +221,22 @@ assert '[ -f "$AUTH_KEYS" ]' "authorized_keys exists"
assert 'grep -q "ssh-ed25519" '"$AUTH_KEYS" "authorized_keys has ssh-ed25519 key"
assert 'grep -q "DOCKER_HOST=" '"$AUTH_KEYS" "authorized_keys has DOCKER_HOST environment"
+# 4f. INT-CI-005 self-clean check: a spawn whose client is killed mid-flight
+# (here: `timeout 5`, which cancels the request context the same way a 300s
+# deadline or a dropped connection does) must leave the host fully clean —
+# no passwd entry, no SSH key, no registry row.
+if command -v timeout >/dev/null 2>&1; then
+ timeout 5 bunker spawn --agent-id regr-leak >/dev/null 2>&1 || true
+ for _ in 1 2 3 4 5 6; do
+ grep -q '^bunker-regr-leak:' /etc/passwd || break
+ sleep 1
+ done
+ assert '! grep -q "^bunker-regr-leak:" /etc/passwd' "timed-out spawn leaves no Linux user"
+ assert '[ ! -e /etc/bunkerd/ssh/regr-leak ]' "timed-out spawn leaves no SSH key"
+fi
+
echo ""
One-line summary: job-level red is real and hidden by continue-on-error; the red is an intermittent spawn stall whose cancellation makes rollback a no-op; run rollback under context.WithoutCancel, gate the workflow on job results, and make the harness report the daemon reason and assert no leaked agent.
# Evidence - Problem class: ci-job-level-red-masked-by-continue-on-error - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-16T12:41:07.272Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: `gh run list` shows conclusion=success for every recent run, yet the newest completed run's `regression` job is RED. On this repo (deployBunker/bunker) jobs `root-suite` and `regression` carry continue-on-error: true, so a run-level green proves NOTHING about them.\n\nATTRIBUTION PROCEDURE (what actually settled it): (1) `gh run list --limit 5 --json databaseId,headSha,conclusion,status,displayTitle`; (2) `gh run view <id> --json jobs --jq '.jobs[] | {name, conclusion, failed: [.steps[]|select(.conclusion==\"failure\")|.name]}'` -> run 35084607865 (sha 358a10c) job `regression`, step `Regression suite`, FAILURE; (3) get the step's real log with `gh api repos/<org>/<repo>/actions/jobs/<job_id>/logs` because `gh run view --log-failed` and `gh run view --job <id> --log` can both return EMPTY on this runner; (4) read TIMINGS, not just PASS/FAIL: the harness printed '\u2500\u2500 4. Spawn \u2500\u2500' at 10:28:23 and resolved 'spawn with explicit ID' only at 10:33:23 = 300.0s, i.e. the client's full spawn deadline, while the next check (auto-generated ID) passed 12s later; (5) find the CONTROL run: 358a10c is a board-only commit on top of 0bd45e4, and 0bd45e4's run 35083640714 executed the SAME code with the SAME harness green (explicit-ID spawn 11s, 'PASS: 33 / FAIL: 0', battery step 'CERTIFIED BINARY ... verdict: MATCH'). Same code, green one run and 9 red checks the next => intermittent live-host stall, NOT a code regression introduced by a commit. VERIFY the control by job id, not by run order: the run that contains a step named the same as the failing one, on a commit whose only diff is (board/docs) files.\n\nHALF-CREATED RESOURCE FINGERPRINT: a spawn that stalls past its deadline leaves the host in a partial state \u2014 '\u2713 Linux user created for regr-alpha' passed while '\u2717 SSH key persisted for regr-alpha ([ -f /etc/bunkerd/ssh/regr-alpha ])' failed and every downstream read said not_found ('agent metrics: not_found: agent \"regr-alpha\" not found'). So the leak is a passwd entry with no SSH key and no registry row \u2014 the fix must make a timed-out spawn either complete or self-clean, and the harness must print the daemon-side reason instead of a bare timeout.\n\nFIX/OUTCOME: filed as a P1 board row (INT-CI-005) rather than hand-patched, because the failure is a live-host flake on the CI runner (bunker-mvp itself) and needs the daemon journal at the stalled window (`journalctl -u bunkerd --since ... --until ...`) plus 10 consecutive green spawn-with-explicit-ID runs as the exit criterion. Report the job-level truth (job+step+verbatim failing checks+timings) in the tick output, never the run-level 'success'.", "environment": "GitHub Actions self-hosted runner 'bunker-mvp' (root), VERSIONLESS go 1.26.5, BUNKERD_COEXIST=1 on :29090/:28080", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-job-level-red-masked-by-continue-on-error", "provider": "openrouter", "solved_at": "2026-09-16T12:41:07.272Z", "version": ""}