Repo: deployBunker/bunker · File fixed: scripts/root-suite.sh (commit 60793b2, +14/−1)
All four cases pass — including the critical one: with FAKE_GO_RC=42 the EXIT trap sees rc=42 and the wrapper exits 42. Without the explicit exit "$rc", the trailing echo would have masked the failure as 0.
I've written the solution to ~/INT-CI-006-solution.md. Here it is:
root-suite panics test timed out after 5m0s: a hardcoded-timeout budget cliff, not a hangRepo: deployBunker/bunker · File fixed: scripts/root-suite.sh (commit 60793b2, +14/−1)
Task: INT-CI-006 · Verdict: Tier 1 PASS, Tier 2 PASS (0ae2a263)
Failing run: 35106585699 (sha 7249291), job root-suite = FAILURE
Control run: 35104867990 (sha 17318af), job root-suite = success, job wall 5m24s
Root-gated CI suite goes red intermittently with panic: test timed out after 5m0s, same code green on adjacent runs. The failing commit 7249291 was board-only — no code/test change. The timeout is a hardcoded go test -timeout 300s sitting seconds below the suite's real wall clock. A second bug masked it: the job is continue-on-error: true, so the run reported success while the job failed (sibling class ci-job-level-red-masked-by-continue-on-error).
The suite is 17 root-gated tests doing 22 real agent spawn/destroy cycles, each starting a rootless dockerd. Real wall clock ≈ 280–300s against a hardcoded 300s timeout — a cliff. A good run clears its own alarm by seconds; a slightly slower one reds with zero code change. Rule of thumb: if the last green run's test phase finished within ~10% of the timeout, the timeout is a cliff.
How to tell budget exhaustion from a hang (do this first):
1. Expand the job, never the run (gh run view <id> --json jobs).
2. --log-failed can be EMPTY on self-hosted; use gh api repos/<org>/<repo>/actions/jobs/<job_id>/logs.
3. Parse the log's own time=... timestamps and compute gaps. A hang = long silent tail; budget exhaustion = continuous activity to the alarm. Here: 22 spawned, 21 destroyed, 44 stage entries, max gap 25.9s, last business line 0.4s before the panic.
4. Read running tests: elapsed values. An entry at (0s) = alarm fired between tests — nothing stuck. A hung test shows a large elapsed value.
5. Compare the control run by job: 35104867990 job wall 5m24s, cleared by seconds.
6. Don't re-open INT-CI-002 (a real hang with the identical symptom string); check run/job/step, not the text.
Make the budget env-overridable with headroom but below the CI step's own allowance (so the EXIT-trap cleanup still runs), pass it through, and make red self-attributing.
#!/usr/bin/env bash
set -uo pipefail
+
+# Env-overridable budget: default 780s = ~2.5x observed cost, below the
+# CI step's 15m (900s) so the EXIT-trap leak cleanup still gets to run.
+ROOT_SUITE_TIMEOUT="${ROOT_SUITE_TIMEOUT:-780s}"
@@
trap cleanup EXIT
-go test -count=1 -run 'TestSpawn|TestCgroup|TestConcurrency' ./...
+echo "root-suite: go test timeout budget=$ROOT_SUITE_TIMEOUT"
+go test -count=1 -run 'TestSpawn|TestCgroup|TestConcurrency' ./... -timeout "$ROOT_SUITE_TIMEOUT"
+rc=$?
+echo "root-suite: go test rc=$rc budget=$ROOT_SUITE_TIMEOUT"
+exit "$rc"
Load-bearing details: the original go test was the last line, so its status leaked into the script's status by accident; adding a trailing echo would make a red suite exit 0, hence rc=$? + explicit exit "$rc". The trap's local rc=$? then sees the real status. Never change the -run selection to make a timeout go away ("raise the budget, not the criteria"). If the script uses set -e, capture with go test ... || rc=$?.
A green suite cannot prove this. Drive the real wrapper against fake go on PATH and assert argv:
mkdir -p /tmp/fakego
printf '#!/usr/bin/env bash\necho "FAKE-GO-ARGV: $*"\nexit "${FAKE_GO_RC:-0}"\n' > /tmp/fakego/go
chmod +x /tmp/fakego/go && export PATH="/tmp/fakego:$PATH"
ROOT_SUITE_TIMEOUT=45s bash scripts/root-suite.sh # -> -timeout 45s
env -u ROOT_SUITE_TIMEOUT bash scripts/root-suite.sh # -> -timeout 780s
FAKE_GO_RC=42 bash scripts/root-suite.sh; echo "rc=$?" # -> rc=42
bash -n scripts/root-suite.sh && echo SYNTAX_OK
Matrix (all reproduced here locally):
| Case | Expected | Observed |
|---|---|---|
ROOT_SUITE_TIMEOUT=45s |
... -run TestSpawn\|TestCgroup\|TestConcurrency ./... -timeout 45s |
✅ |
| var unset | ... -timeout 780s |
✅ |
FAKE_GO_RC=42 |
wrapper exits 42 through the trap |
✅ |
bash -n |
SYNTAX_OK |
✅ |
git show --name-only |
only scripts/root-suite.sh |
✅ (Tier 2) |
Tier 1 PASS, Tier 2 PASS (0ae2a263).
Any CI harness that hardcodes a test-runner timeout, especially with real host work. If the last green run finished within ~10% of the timeout, it's a cliff. Use an env-overridable budget at ~2.5x observed cost, under the CI step's own timeout so cleanup runs, and make the failure self-attributing.
Related: ci-job-level-red-masked-by-continue-on-error (the masking half), gitreins-tier2-input-token-cap-exceeded (same "cap ≠ merit" doctrine).
The full document is at ~/INT-CI-006-solution.md. The mechanism verification I ran confirms the one subtle failure mode the fix guards against: without exit "$rc", the diagnostic echo after go test silently converts a red suite into exit 0 — and with it, the EXIT trap correctly observes rc=42.
# Evidence - Problem class: ci-hardcoded-test-timeout-budget-cliff - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-16T15:11:17.415Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a root-gated CI suite reds INTERMITTENTLY with `panic: test timed out after 5m0s` and the same code is green one run earlier, so it reads like a flake or a commit regression. In this instance run 35106585699 (sha 7249291, a BOARD-ONLY commit on top of a docs commit) reported job `root-suite`=FAILURE while `gh run list` reported the RUN as `success` (the job is continue-on-error) \u2014 see the sibling class `ci-job-level-red-masked-by-continue-on-error` for that masking half.\n\nROOT CAUSE: budget exhaustion, NOT a hang. The suite's go test timeout was hardcoded to 300s and the suite's REAL wall clock on that runner is ~280-300s (17 root-gated tests, 22 real agent spawn/destroy cycles each starting a rootless dockerd). The suite therefore finishes seconds under its own alarm on a good run and reddens on a slightly slower one, with zero code change.\n\nHOW TO TELL BUDGET EXHAUSTION FROM A HANG (this is the transferable part \u2014 do it before writing any fix):\n1. Expand the JOB, never the run: `gh run view <id> --json jobs --jq '.jobs[] | {name, conclusion, failed: [.steps[]|select(.conclusion==\"failure\")|.name]}'`.\n2. The step log needs the api fallback on a self-hosted runner: `gh run view --log-failed` (and `gh run view --job <id> --log`) can return EMPTY \u2014 use `gh api repos/<org>/<repo>/actions/jobs/<job_id>/logs > /tmp/<tag>.log`.\n3. Parse the log's OWN timestamps (the go log lines carry `time=...`; the gh prefix timestamp is the flush time, not the event time) and compute the inter-event GAPS. A hang shows a long silent tail; budget exhaustion shows continuous business activity right up to the alarm. Here: 22 `agent spawned successfully`, 21 `agent destroyed`, 44 `spawn entering stage` in the 300s window, MAXIMUM inter-log gap 25.9s, and the last business line landed 0.4s before the panic (14:44:21.711 'agent destroyed diskmax-99597').\n4. Read the panic dump's `running tests:` list with its elapsed values: an entry at `(0s)` means the alarm fired between tests \u2014 the NEXT test in line had just started, so nothing was stuck. A hung test shows a large elapsed value instead.\n5. Find the CONTROL run and compare BY JOB, not by run order: the same selection was green on run 35104867990 with the whole job at 5m24s (setup + go test), i.e. it cleared the cliff by seconds. Same code green/red = timing cliff, not regression.\n6. Confirm you are not re-opening an older hang-class row: this repo's INT-CI-002 previously closed a REAL hang (a nil-builder regression test that never returned) with the identical symptom string. The identical string is a coincidence of the same 5m wall \u2014 check the run/job/step, not the symptom text.\n\nFIX SHAPE (applied at commit 60793b2, scripts/root-suite.sh, +14/-1):\n- Make the budget overridable with a default that has HEADROOM while staying STRICTLY LOWER than the CI step's own allowance, so the wrapper's EXIT-trap leak cleanup still gets to run: `ROOT_SUITE_TIMEOUT=\"${ROOT_SUITE_TIMEOUT:-780s}\"` against a step with `timeout-minutes: 15` (900s).\n- Pass it through: `go test -count=1 -run '<exact selection>' ./... -timeout \"$ROOT_SUITE_TIMEOUT\"` \u2014 never touch the test selection to make a timeout go away (that is the 'raise the budget, not the criteria' rule; cf. gitreins cap-sizing).\n- Make the red SELF-ATTRIBUTING: echo the budget before the run and the go-test rc after it, and re-exit that rc explicitly (`rc=$?; echo \"... rc=$rc budget=$ROOT_SUITE_TIMEOUT\"; exit \"$rc\"`). This is load-bearing in bash: adding a trailing echo after the test makes the script exit 0 on a RED suite unless the rc is re-exited, and the EXIT trap's `local rc=$?` must still see the real status.\n- Accept the honest trade: a genuine hang now takes 13m to red instead of 5m \u2014 the job still fails, and the log line makes it attributable.\n\nVERIFICATION THAT ACTUALLY PROVES IT (a green suite cannot, and the real suite needs root): drive the REAL wrapper against FAKE `go` binaries on PATH and assert the argv it would execute. Matrix run with the foreman's own stubs (/tmp/foreman-verify-INT-CI-006.sh): ROOT_SUITE_TIMEOUT=45s -> `FAKE-GO-ARGV: test -count=1 -run TestSpawn|TestCgroup|TestConcurrency ./... -timeout 45s`; var unset -> `-timeout 780s`; fake go exiting 42 -> wrapper exits 42 (rc propagated through the EXIT trap); `bash -n` SYNTAX_OK; `git show --name-only` = the single script. Tier 2 judge PASS (verdict 0ae2a263) with the judge re-deriving the same matrix.\n\nWHEN TO APPLY: any CI harness that hardcodes a test-runner timeout, especially one whose suite does real host work (spawns, containers, users). Rule of thumb: if the last successful run's go-test phase finished within ~10% of the timeout, the timeout is a cliff and WILL red with no code change. Prefer an env-overridable budget at ~2.5x the observed cost, keeping it under the CI step's own timeout so cleanup still runs.", "environment": "GitHub Actions self-hosted runner (root) running `bash scripts/root-suite.sh`; go test -count=1 -run 'TestSpawn|TestCgroup|TestConcurrency' ./... under a hardcoded -timeout 300s inside a CI step with timeout-minutes: 15. Repo deployBunker/bunker.", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-hardcoded-test-timeout-budget-cliff", "provider": "openrouter", "solved_at": "2026-09-16T15:11:17.415Z", "version": ""}All four cases pass — including the critical one: with FAKE_GO_RC=42 the EXIT trap sees rc=42 and the wrapper exits 42. Without the explicit exit "$rc", the trailing echo would have masked the failure as 0.
I've written the solution to ~/INT-CI-006-solution.md. Here it is:
root-suite panics test timed out after 5m0s: a hardcoded-timeout budget cliff, not a hangRepo: deployBunker/bunker · File fixed: scripts/root-suite.sh (commit 60793b2, +14/−1)
Task: INT-CI-006 · Verdict: Tier 1 PASS, Tier 2 PASS (0ae2a263)
Failing run: 35106585699 (sha 7249291), job root-suite = FAILURE
Control run: 35104867990 (sha 17318af), job root-suite = success, job wall 5m24s
Root-gated CI suite goes red intermittently with panic: test timed out after 5m0s, same code green on adjacent runs. The failing commit 7249291 was board-only — no code/test change. The timeout is a hardcoded go test -timeout 300s sitting seconds below the suite's real wall clock. A second bug masked it: the job is continue-on-error: true, so the run reported success while the job failed (sibling class ci-job-level-red-masked-by-continue-on-error).
The suite is 17 root-gated tests doing 22 real agent spawn/destroy cycles, each starting a rootless dockerd. Real wall clock ≈ 280–300s against a hardcoded 300s timeout — a cliff. A good run clears its own alarm by seconds; a slightly slower one reds with zero code change. Rule of thumb: if the last green run's test phase finished within ~10% of the timeout, the timeout is a cliff.
How to tell budget exhaustion from a hang (do this first):
1. Expand the job, never the run (gh run view <id> --json jobs).
2. --log-failed can be EMPTY on self-hosted; use gh api repos/<org>/<repo>/actions/jobs/<job_id>/logs.
3. Parse the log's own time=... timestamps and compute gaps. A hang = long silent tail; budget exhaustion = continuous activity to the alarm. Here: 22 spawned, 21 destroyed, 44 stage entries, max gap 25.9s, last business line 0.4s before the panic.
4. Read running tests: elapsed values. An entry at (0s) = alarm fired between tests — nothing stuck. A hung test shows a large elapsed value.
5. Compare the control run by job: 35104867990 job wall 5m24s, cleared by seconds.
6. Don't re-open INT-CI-002 (a real hang with the identical symptom string); check run/job/step, not the text.
Make the budget env-overridable with headroom but below the CI step's own allowance (so the EXIT-trap cleanup still runs), pass it through, and make red self-attributing.
#!/usr/bin/env bash
set -uo pipefail
+
+# Env-overridable budget: default 780s = ~2.5x observed cost, below the
+# CI step's 15m (900s) so the EXIT-trap leak cleanup still gets to run.
+ROOT_SUITE_TIMEOUT="${ROOT_SUITE_TIMEOUT:-780s}"
@@
trap cleanup EXIT
-go test -count=1 -run 'TestSpawn|TestCgroup|TestConcurrency' ./...
+echo "root-suite: go test timeout budget=$ROOT_SUITE_TIMEOUT"
+go test -count=1 -run 'TestSpawn|TestCgroup|TestConcurrency' ./... -timeout "$ROOT_SUITE_TIMEOUT"
+rc=$?
+echo "root-suite: go test rc=$rc budget=$ROOT_SUITE_TIMEOUT"
+exit "$rc"
Load-bearing details: the original go test was the last line, so its status leaked into the script's status by accident; adding a trailing echo would make a red suite exit 0, hence rc=$? + explicit exit "$rc". The trap's local rc=$? then sees the real status. Never change the -run selection to make a timeout go away ("raise the budget, not the criteria"). If the script uses set -e, capture with go test ... || rc=$?.
A green suite cannot prove this. Drive the real wrapper against fake go on PATH and assert argv:
mkdir -p /tmp/fakego
printf '#!/usr/bin/env bash\necho "FAKE-GO-ARGV: $*"\nexit "${FAKE_GO_RC:-0}"\n' > /tmp/fakego/go
chmod +x /tmp/fakego/go && export PATH="/tmp/fakego:$PATH"
ROOT_SUITE_TIMEOUT=45s bash scripts/root-suite.sh # -> -timeout 45s
env -u ROOT_SUITE_TIMEOUT bash scripts/root-suite.sh # -> -timeout 780s
FAKE_GO_RC=42 bash scripts/root-suite.sh; echo "rc=$?" # -> rc=42
bash -n scripts/root-suite.sh && echo SYNTAX_OK
Matrix (all reproduced here locally):
| Case | Expected | Observed |
|---|---|---|
ROOT_SUITE_TIMEOUT=45s |
... -run TestSpawn\|TestCgroup\|TestConcurrency ./... -timeout 45s |
✅ |
| var unset | ... -timeout 780s |
✅ |
FAKE_GO_RC=42 |
wrapper exits 42 through the trap |
✅ |
bash -n |
SYNTAX_OK |
✅ |
git show --name-only |
only scripts/root-suite.sh |
✅ (Tier 2) |
Tier 1 PASS, Tier 2 PASS (0ae2a263).
Any CI harness that hardcodes a test-runner timeout, especially with real host work. If the last green run finished within ~10% of the timeout, it's a cliff. Use an env-overridable budget at ~2.5x observed cost, under the CI step's own timeout so cleanup runs, and make the failure self-attributing.
Related: ci-job-level-red-masked-by-continue-on-error (the masking half), gitreins-tier2-input-token-cap-exceeded (same "cap ≠ merit" doctrine).
The full document is at ~/INT-CI-006-solution.md. The mechanism verification I ran confirms the one subtle failure mode the fix guards against: without exit "$rc", the diagnostic echo after go test silently converts a red suite into exit 0 — and with it, the EXIT trap correctly observes rc=42.
# Evidence - Problem class: ci-hardcoded-test-timeout-budget-cliff - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-16T15:11:17.415Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a root-gated CI suite reds INTERMITTENTLY with `panic: test timed out after 5m0s` and the same code is green one run earlier, so it reads like a flake or a commit regression. In this instance run 35106585699 (sha 7249291, a BOARD-ONLY commit on top of a docs commit) reported job `root-suite`=FAILURE while `gh run list` reported the RUN as `success` (the job is continue-on-error) \u2014 see the sibling class `ci-job-level-red-masked-by-continue-on-error` for that masking half.\n\nROOT CAUSE: budget exhaustion, NOT a hang. The suite's go test timeout was hardcoded to 300s and the suite's REAL wall clock on that runner is ~280-300s (17 root-gated tests, 22 real agent spawn/destroy cycles each starting a rootless dockerd). The suite therefore finishes seconds under its own alarm on a good run and reddens on a slightly slower one, with zero code change.\n\nHOW TO TELL BUDGET EXHAUSTION FROM A HANG (this is the transferable part \u2014 do it before writing any fix):\n1. Expand the JOB, never the run: `gh run view <id> --json jobs --jq '.jobs[] | {name, conclusion, failed: [.steps[]|select(.conclusion==\"failure\")|.name]}'`.\n2. The step log needs the api fallback on a self-hosted runner: `gh run view --log-failed` (and `gh run view --job <id> --log`) can return EMPTY \u2014 use `gh api repos/<org>/<repo>/actions/jobs/<job_id>/logs > /tmp/<tag>.log`.\n3. Parse the log's OWN timestamps (the go log lines carry `time=...`; the gh prefix timestamp is the flush time, not the event time) and compute the inter-event GAPS. A hang shows a long silent tail; budget exhaustion shows continuous business activity right up to the alarm. Here: 22 `agent spawned successfully`, 21 `agent destroyed`, 44 `spawn entering stage` in the 300s window, MAXIMUM inter-log gap 25.9s, and the last business line landed 0.4s before the panic (14:44:21.711 'agent destroyed diskmax-99597').\n4. Read the panic dump's `running tests:` list with its elapsed values: an entry at `(0s)` means the alarm fired between tests \u2014 the NEXT test in line had just started, so nothing was stuck. A hung test shows a large elapsed value instead.\n5. Find the CONTROL run and compare BY JOB, not by run order: the same selection was green on run 35104867990 with the whole job at 5m24s (setup + go test), i.e. it cleared the cliff by seconds. Same code green/red = timing cliff, not regression.\n6. Confirm you are not re-opening an older hang-class row: this repo's INT-CI-002 previously closed a REAL hang (a nil-builder regression test that never returned) with the identical symptom string. The identical string is a coincidence of the same 5m wall \u2014 check the run/job/step, not the symptom text.\n\nFIX SHAPE (applied at commit 60793b2, scripts/root-suite.sh, +14/-1):\n- Make the budget overridable with a default that has HEADROOM while staying STRICTLY LOWER than the CI step's own allowance, so the wrapper's EXIT-trap leak cleanup still gets to run: `ROOT_SUITE_TIMEOUT=\"${ROOT_SUITE_TIMEOUT:-780s}\"` against a step with `timeout-minutes: 15` (900s).\n- Pass it through: `go test -count=1 -run '<exact selection>' ./... -timeout \"$ROOT_SUITE_TIMEOUT\"` \u2014 never touch the test selection to make a timeout go away (that is the 'raise the budget, not the criteria' rule; cf. gitreins cap-sizing).\n- Make the red SELF-ATTRIBUTING: echo the budget before the run and the go-test rc after it, and re-exit that rc explicitly (`rc=$?; echo \"... rc=$rc budget=$ROOT_SUITE_TIMEOUT\"; exit \"$rc\"`). This is load-bearing in bash: adding a trailing echo after the test makes the script exit 0 on a RED suite unless the rc is re-exited, and the EXIT trap's `local rc=$?` must still see the real status.\n- Accept the honest trade: a genuine hang now takes 13m to red instead of 5m \u2014 the job still fails, and the log line makes it attributable.\n\nVERIFICATION THAT ACTUALLY PROVES IT (a green suite cannot, and the real suite needs root): drive the REAL wrapper against FAKE `go` binaries on PATH and assert the argv it would execute. Matrix run with the foreman's own stubs (/tmp/foreman-verify-INT-CI-006.sh): ROOT_SUITE_TIMEOUT=45s -> `FAKE-GO-ARGV: test -count=1 -run TestSpawn|TestCgroup|TestConcurrency ./... -timeout 45s`; var unset -> `-timeout 780s`; fake go exiting 42 -> wrapper exits 42 (rc propagated through the EXIT trap); `bash -n` SYNTAX_OK; `git show --name-only` = the single script. Tier 2 judge PASS (verdict 0ae2a263) with the judge re-deriving the same matrix.\n\nWHEN TO APPLY: any CI harness that hardcodes a test-runner timeout, especially one whose suite does real host work (spawns, containers, users). Rule of thumb: if the last successful run's go-test phase finished within ~10% of the timeout, the timeout is a cliff and WILL red with no code change. Prefer an env-overridable budget at ~2.5x the observed cost, keeping it under the CI step's own timeout so cleanup still runs.", "environment": "GitHub Actions self-hosted runner (root) running `bash scripts/root-suite.sh`; go test -count=1 -run 'TestSpawn|TestCgroup|TestConcurrency' ./... under a hardcoded -timeout 300s inside a CI step with timeout-minutes: 15. Repo deployBunker/bunker.", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-hardcoded-test-timeout-budget-cliff", "provider": "openrouter", "solved_at": "2026-09-16T15:11:17.415Z", "version": ""}