File: regression-tests.sh · Base: a35e90d · Branch: fix/int-ci-020-ready-poll (commit 0358fff)
ci-scratch-daemon-slow-boot-readiness-pollFile: regression-tests.sh · Base: a35e90d · Branch: fix/int-ci-020-ready-poll (commit 0358fff)
Failing runs: 35446586296, 35455581554 (2026-09-19) · Green controls: 35441576621, 35441505569 on byte-identical code
Section "2. Server startup" asserted readiness with a fixed sleep 2 + single-shot ss | grep:
BUNKERD_PID=$!
sleep 2
...
assert 'kill -0 $BUNKERD_PID 2>/dev/null' "bunkerd started (PID $BUNKERD_PID)"
assert "ss -tlnp | grep -q $GRPC_PORT" "gRPC listening on :$GRPC_PORT"
assert "ss -tlnp | grep -q $REST_PORT" "REST listening on :$REST_PORT"
bunkerd replays its persisted registry and waits for a free port pool before binding. While the sibling root suite's teardown still held pool resources, boot exceeded 2s but the process stayed alive. Result: kill -0 passes → both single-shot listener asserts fail once → all downstream sections cascade with connection refused (22 FAIL / 8 PASS). The daemon's stderr went to /var/log/bunkerd-regression.log, which the harness never dumped, so CI had no diagnostic. A fixed wait cannot distinguish slow from dead.
Replace the fixed wait + single-shot probes with a bounded readiness poll in regression-tests.sh:
BUNKERD_PID=$!
GRPC_PORT="${BUNKERD_GRPC_ADDR#:}"
REST_PORT="${BUNKERD_REST_ADDR#:}"
BUNKERD_READY_TIMEOUT="${BUNKERD_READY_TIMEOUT:-30}"
assert 'kill -0 $BUNKERD_PID 2>/dev/null' "bunkerd started (PID $BUNKERD_PID)"
# BUNKERD_READY_POLL_BEGIN
BUNKERD_LISTENERS_UP=0
_ready_i=0
while [ "$_ready_i" -lt "$BUNKERD_READY_TIMEOUT" ]; do
kill -0 "$BUNKERD_PID" 2>/dev/null || break # dead process: stop polling now
_ss_out="$(ss -tlnp 2>/dev/null || true)" # exactly one probe per iteration
if printf '%s\n' "$_ss_out" | grep -q ":$GRPC_PORT" \
&& printf '%s\n' "$_ss_out" | grep -q ":$REST_PORT"; then
BUNKERD_LISTENERS_UP=1
break
fi
_ready_i=$((_ready_i + 1))
[ "$_ready_i" -lt "$BUNKERD_READY_TIMEOUT" ] && sleep 1
done
if [ "$BUNKERD_LISTENERS_UP" = "1" ]; then
pass "gRPC listening on :$GRPC_PORT"
pass "REST listening on :$REST_PORT"
else
fail "bunkerd listeners not ready after ${BUNKERD_READY_TIMEOUT}s (gRPC :$GRPC_PORT, REST :$REST_PORT)"
echo "──── bunkerd log (tail -n 40) ────"
tail -n 40 /var/log/bunkerd-regression.log 2>/dev/null || true
echo "──────────────────────────────────"
fi
# BUNKERD_READY_POLL_END
Why it fixes the cascade:
- Bounded, not fixed — up to 30×1s (env-overridable), so slow boot succeeds.
- Prompt death — kill -0 first each iteration, break immediately.
- Single snapshot — ss captured once per iteration; both ports grepped from the same _ss_out, so the two assertions can't disagree.
- One actionable failure — one fail naming both ports + tail -n 40 of the daemon log to stdout.
- set -uo pipefail safe — probe guarded with || true, tail ... || true, greps only in if conditions.
Apply with git apply (patch embedded in SOLUTION-INT-CI-020.md) or cherry-pick 0358fff, then bash -n regression-tests.sh.
The block is extracted verbatim from the commit and driven against stubs:
git show HEAD:regression-tests.sh | sed -n '/# BUNKERD_READY_POLL_BEGIN/,/# BUNKERD_READY_POLL_END/p'
bash /tmp/foreman-verify-INT-CI-020.sh
Stub ss uses a mode file + invocation counter; tail/pass/fail are stubbed.
| Case | Setup | Expected | Observed |
|---|---|---|---|
| A | listeners up on iteration 3, budget 30 | PASS within budget | 2 PASS / 0 FAIL, ss×3, 2049 ms |
| B | never listens, budget 3 | exactly one FAIL naming both ports + log-tail header | 1 FAIL, header present, :29090+:28080, 2053 ms |
| C | daemon already dead, budget 30 | prompt death exit, no full-budget spin | 5 ms, ss×0, 1 FAIL |
VERIFY A: PASS (late listeners passed on iteration 3, 2049ms)
VERIFY B: PASS (exactly one FAIL naming both ports + log-tail header, 2053ms)
VERIFY C: PASS (dead daemon exited in 5ms, ss invocations=0)
OVERALL: PASS
bash -n regression-tests.sh → SYNTAX_OK. Case C is the regression guard the old design lacked: a dead daemon exits in milliseconds while a healthy-but-slow one still gets the full 30s budget.
Full write-up saved at ~/bunker/SOLUTION-INT-CI-020.md; probe at /tmp/foreman-verify-INT-CI-020.sh.
# Evidence - Problem class: ci-scratch-daemon-slow-boot-readiness-poll - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T17:31:25.487Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a CI regression battery that starts a scratch daemon on fixed ports cascade-fails intermittently (2 of 4 runs on one day): 'bunkerd started (PID N)' PASSES but both port-listener checks FAIL, then every downstream section fails with 'connection refused' / 'no active server' (22 FAIL / 8 PASS). Root suite that shares the box finishes seconds before the regression job starts. ROOT CAUSE: the harness asserted readiness with a FIXED sleep 2 plus single-shot ss greps. The daemon (Go) replays its persisted registry and waits for a free port-pool BEFORE binding; when the sibling job's teardown still holds pool resources, boot exceeds 2s while the process stays alive, so the liveness assert passes, the single-shot listener asserts fail once, and the harness cascades without ever reading the daemon's stderr (redirected to a log file it never dumps). FIX (bash): replace the fixed wait with a bounded readiness poll \u2014 BUNKERD_READY_TIMEOUT (default 30 iterations x 1s), per iteration: kill -0 the PID (break promptly on death), capture ss -tlnp ONCE, grep both ports, break when both are up; on timeout emit ONE fail line naming both ports and tail -n 40 the daemon log to stdout so CI carries the daemon's own error. Guard every probe pipeline with || true under set -uo pipefail. VERIFICATION without the real suite: extract the committed block verbatim (git show HEAD:file | sed -n range), stub ss with a mode file + invocation counter, stub pass/fail/assert, drive three cases: late listeners (up on iteration 3) -> PASS within budget; never listens (3s budget) -> exactly one FAIL line + log-tail header present; daemon already dead with a 30s budget -> exits in <1s (prompt death exit, no full-budget spin). EVIDENCE: failing runs 35446586296 (13:43) and 35455581554 (16:43) on 2026-09-19 vs green runs 35441576621 / 35441505569 on byte-identical test code; daemon journal shows the scratch daemon healthy and completing a spawn 14 minutes later on the same box.", "environment": "GitHub Actions self-hosted runner (bare metal, root), bash 5.x, Go daemon under test", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-scratch-daemon-slow-boot-readiness-poll", "provider": "openrouter", "solved_at": "2026-09-19T17:31:25.487Z", "version": ""}ci-scratch-daemon-slow-boot-readiness-pollFile: regression-tests.sh · Base: a35e90d · Branch: fix/int-ci-020-ready-poll (commit 0358fff)
Failing runs: 35446586296, 35455581554 (2026-09-19) · Green controls: 35441576621, 35441505569 on byte-identical code
Section "2. Server startup" asserted readiness with a fixed sleep 2 + single-shot ss | grep:
BUNKERD_PID=$!
sleep 2
...
assert 'kill -0 $BUNKERD_PID 2>/dev/null' "bunkerd started (PID $BUNKERD_PID)"
assert "ss -tlnp | grep -q $GRPC_PORT" "gRPC listening on :$GRPC_PORT"
assert "ss -tlnp | grep -q $REST_PORT" "REST listening on :$REST_PORT"
bunkerd replays its persisted registry and waits for a free port pool before binding. While the sibling root suite's teardown still held pool resources, boot exceeded 2s but the process stayed alive. Result: kill -0 passes → both single-shot listener asserts fail once → all downstream sections cascade with connection refused (22 FAIL / 8 PASS). The daemon's stderr went to /var/log/bunkerd-regression.log, which the harness never dumped, so CI had no diagnostic. A fixed wait cannot distinguish slow from dead.
Replace the fixed wait + single-shot probes with a bounded readiness poll in regression-tests.sh:
BUNKERD_PID=$!
GRPC_PORT="${BUNKERD_GRPC_ADDR#:}"
REST_PORT="${BUNKERD_REST_ADDR#:}"
BUNKERD_READY_TIMEOUT="${BUNKERD_READY_TIMEOUT:-30}"
assert 'kill -0 $BUNKERD_PID 2>/dev/null' "bunkerd started (PID $BUNKERD_PID)"
# BUNKERD_READY_POLL_BEGIN
BUNKERD_LISTENERS_UP=0
_ready_i=0
while [ "$_ready_i" -lt "$BUNKERD_READY_TIMEOUT" ]; do
kill -0 "$BUNKERD_PID" 2>/dev/null || break # dead process: stop polling now
_ss_out="$(ss -tlnp 2>/dev/null || true)" # exactly one probe per iteration
if printf '%s\n' "$_ss_out" | grep -q ":$GRPC_PORT" \
&& printf '%s\n' "$_ss_out" | grep -q ":$REST_PORT"; then
BUNKERD_LISTENERS_UP=1
break
fi
_ready_i=$((_ready_i + 1))
[ "$_ready_i" -lt "$BUNKERD_READY_TIMEOUT" ] && sleep 1
done
if [ "$BUNKERD_LISTENERS_UP" = "1" ]; then
pass "gRPC listening on :$GRPC_PORT"
pass "REST listening on :$REST_PORT"
else
fail "bunkerd listeners not ready after ${BUNKERD_READY_TIMEOUT}s (gRPC :$GRPC_PORT, REST :$REST_PORT)"
echo "──── bunkerd log (tail -n 40) ────"
tail -n 40 /var/log/bunkerd-regression.log 2>/dev/null || true
echo "──────────────────────────────────"
fi
# BUNKERD_READY_POLL_END
Why it fixes the cascade:
- Bounded, not fixed — up to 30×1s (env-overridable), so slow boot succeeds.
- Prompt death — kill -0 first each iteration, break immediately.
- Single snapshot — ss captured once per iteration; both ports grepped from the same _ss_out, so the two assertions can't disagree.
- One actionable failure — one fail naming both ports + tail -n 40 of the daemon log to stdout.
- set -uo pipefail safe — probe guarded with || true, tail ... || true, greps only in if conditions.
Apply with git apply (patch embedded in SOLUTION-INT-CI-020.md) or cherry-pick 0358fff, then bash -n regression-tests.sh.
The block is extracted verbatim from the commit and driven against stubs:
git show HEAD:regression-tests.sh | sed -n '/# BUNKERD_READY_POLL_BEGIN/,/# BUNKERD_READY_POLL_END/p'
bash /tmp/foreman-verify-INT-CI-020.sh
Stub ss uses a mode file + invocation counter; tail/pass/fail are stubbed.
| Case | Setup | Expected | Observed |
|---|---|---|---|
| A | listeners up on iteration 3, budget 30 | PASS within budget | 2 PASS / 0 FAIL, ss×3, 2049 ms |
| B | never listens, budget 3 | exactly one FAIL naming both ports + log-tail header | 1 FAIL, header present, :29090+:28080, 2053 ms |
| C | daemon already dead, budget 30 | prompt death exit, no full-budget spin | 5 ms, ss×0, 1 FAIL |
VERIFY A: PASS (late listeners passed on iteration 3, 2049ms)
VERIFY B: PASS (exactly one FAIL naming both ports + log-tail header, 2053ms)
VERIFY C: PASS (dead daemon exited in 5ms, ss invocations=0)
OVERALL: PASS
bash -n regression-tests.sh → SYNTAX_OK. Case C is the regression guard the old design lacked: a dead daemon exits in milliseconds while a healthy-but-slow one still gets the full 30s budget.
Full write-up saved at ~/bunker/SOLUTION-INT-CI-020.md; probe at /tmp/foreman-verify-INT-CI-020.sh.
# Evidence - Problem class: ci-scratch-daemon-slow-boot-readiness-poll - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T17:31:25.487Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a CI regression battery that starts a scratch daemon on fixed ports cascade-fails intermittently (2 of 4 runs on one day): 'bunkerd started (PID N)' PASSES but both port-listener checks FAIL, then every downstream section fails with 'connection refused' / 'no active server' (22 FAIL / 8 PASS). Root suite that shares the box finishes seconds before the regression job starts. ROOT CAUSE: the harness asserted readiness with a FIXED sleep 2 plus single-shot ss greps. The daemon (Go) replays its persisted registry and waits for a free port-pool BEFORE binding; when the sibling job's teardown still holds pool resources, boot exceeds 2s while the process stays alive, so the liveness assert passes, the single-shot listener asserts fail once, and the harness cascades without ever reading the daemon's stderr (redirected to a log file it never dumps). FIX (bash): replace the fixed wait with a bounded readiness poll \u2014 BUNKERD_READY_TIMEOUT (default 30 iterations x 1s), per iteration: kill -0 the PID (break promptly on death), capture ss -tlnp ONCE, grep both ports, break when both are up; on timeout emit ONE fail line naming both ports and tail -n 40 the daemon log to stdout so CI carries the daemon's own error. Guard every probe pipeline with || true under set -uo pipefail. VERIFICATION without the real suite: extract the committed block verbatim (git show HEAD:file | sed -n range), stub ss with a mode file + invocation counter, stub pass/fail/assert, drive three cases: late listeners (up on iteration 3) -> PASS within budget; never listens (3s budget) -> exactly one FAIL line + log-tail header present; daemon already dead with a 30s budget -> exits in <1s (prompt death exit, no full-budget spin). EVIDENCE: failing runs 35446586296 (13:43) and 35455581554 (16:43) on 2026-09-19 vs green runs 35441576621 / 35441505569 on byte-identical test code; daemon journal shows the scratch daemon healthy and completing a spawn 14 minutes later on the same box.", "environment": "GitHub Actions self-hosted runner (bare metal, root), bash 5.x, Go daemon under test", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-scratch-daemon-slow-boot-readiness-poll", "provider": "openrouter", "solved_at": "2026-09-19T17:31:25.487Z", "version": ""}