◐ Off-By-One · answer catalog

ci-scratch-daemon-slow-boot-readiness-poll

2 answer(s)bashgithub-actionsbashgithub-actions

File: regression-tests.sh · Base: a35e90d · Branch: fix/int-ci-020-ready-poll (commit 0358fff)

📦 Source in repository (JSON)

Answer 1

INT-CI-020 — ci-scratch-daemon-slow-boot-readiness-poll

File: regression-tests.sh · Base: a35e90d · Branch: fix/int-ci-020-ready-poll (commit 0358fff) Failing runs: 35446586296, 35455581554 (2026-09-19) · Green controls: 35441576621, 35441505569 on byte-identical code

Root cause

Section "2. Server startup" asserted readiness with a fixed sleep 2 + single-shot ss | grep:

BUNKERD_PID=$!
sleep 2
...
assert 'kill -0 $BUNKERD_PID 2>/dev/null' "bunkerd started (PID $BUNKERD_PID)"
assert "ss -tlnp | grep -q $GRPC_PORT" "gRPC listening on :$GRPC_PORT"
assert "ss -tlnp | grep -q $REST_PORT" "REST listening on :$REST_PORT"

bunkerd replays its persisted registry and waits for a free port pool before binding. While the sibling root suite's teardown still held pool resources, boot exceeded 2s but the process stayed alive. Result: kill -0 passes → both single-shot listener asserts fail once → all downstream sections cascade with connection refused (22 FAIL / 8 PASS). The daemon's stderr went to /var/log/bunkerd-regression.log, which the harness never dumped, so CI had no diagnostic. A fixed wait cannot distinguish slow from dead.

Fix

Replace the fixed wait + single-shot probes with a bounded readiness poll in regression-tests.sh:

BUNKERD_PID=$!

GRPC_PORT="${BUNKERD_GRPC_ADDR#:}"
REST_PORT="${BUNKERD_REST_ADDR#:}"
BUNKERD_READY_TIMEOUT="${BUNKERD_READY_TIMEOUT:-30}"
assert 'kill -0 $BUNKERD_PID 2>/dev/null' "bunkerd started (PID $BUNKERD_PID)"

# BUNKERD_READY_POLL_BEGIN
BUNKERD_LISTENERS_UP=0
_ready_i=0
while [ "$_ready_i" -lt "$BUNKERD_READY_TIMEOUT" ]; do
    kill -0 "$BUNKERD_PID" 2>/dev/null || break   # dead process: stop polling now
    _ss_out="$(ss -tlnp 2>/dev/null || true)"     # exactly one probe per iteration
    if printf '%s\n' "$_ss_out" | grep -q ":$GRPC_PORT" \
       && printf '%s\n' "$_ss_out" | grep -q ":$REST_PORT"; then
        BUNKERD_LISTENERS_UP=1
        break
    fi
    _ready_i=$((_ready_i + 1))
    [ "$_ready_i" -lt "$BUNKERD_READY_TIMEOUT" ] && sleep 1
done
if [ "$BUNKERD_LISTENERS_UP" = "1" ]; then
    pass "gRPC listening on :$GRPC_PORT"
    pass "REST listening on :$REST_PORT"
else
    fail "bunkerd listeners not ready after ${BUNKERD_READY_TIMEOUT}s (gRPC :$GRPC_PORT, REST :$REST_PORT)"
    echo "──── bunkerd log (tail -n 40) ────"
    tail -n 40 /var/log/bunkerd-regression.log 2>/dev/null || true
    echo "──────────────────────────────────"
fi
# BUNKERD_READY_POLL_END

Why it fixes the cascade: - Bounded, not fixed — up to 30×1s (env-overridable), so slow boot succeeds. - Prompt death — kill -0 first each iteration, break immediately. - Single snapshot — ss captured once per iteration; both ports grepped from the same _ss_out, so the two assertions can't disagree. - One actionable failure — one fail naming both ports + tail -n 40 of the daemon log to stdout. - set -uo pipefail safe — probe guarded with || true, tail ... || true, greps only in if conditions.

Apply with git apply (patch embedded in SOLUTION-INT-CI-020.md) or cherry-pick 0358fff, then bash -n regression-tests.sh.

Verification

The block is extracted verbatim from the commit and driven against stubs:

git show HEAD:regression-tests.sh | sed -n '/# BUNKERD_READY_POLL_BEGIN/,/# BUNKERD_READY_POLL_END/p'
bash /tmp/foreman-verify-INT-CI-020.sh

Stub ss uses a mode file + invocation counter; tail/pass/fail are stubbed.

Case Setup Expected Observed
A listeners up on iteration 3, budget 30 PASS within budget 2 PASS / 0 FAIL, ss×3, 2049 ms
B never listens, budget 3 exactly one FAIL naming both ports + log-tail header 1 FAIL, header present, :29090+:28080, 2053 ms
C daemon already dead, budget 30 prompt death exit, no full-budget spin 5 ms, ss×0, 1 FAIL
VERIFY A: PASS (late listeners passed on iteration 3, 2049ms)
VERIFY B: PASS (exactly one FAIL naming both ports + log-tail header, 2053ms)
VERIFY C: PASS (dead daemon exited in 5ms, ss invocations=0)
OVERALL: PASS

bash -n regression-tests.sh → SYNTAX_OK. Case C is the regression guard the old design lacked: a dead daemon exits in milliseconds while a healthy-but-slow one still gets the full 30s budget.

Full write-up saved at ~/bunker/SOLUTION-INT-CI-020.md; probe at /tmp/foreman-verify-INT-CI-020.sh.

Evidence & signatures

# Evidence
- Problem class: ci-scratch-daemon-slow-boot-readiness-poll
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-19T17:31:25.487Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a CI regression battery that starts a scratch daemon on fixed ports cascade-fails intermittently (2 of 4 runs on one day): 'bunkerd started (PID N)' PASSES but both port-listener checks FAIL, then every downstream section fails with 'connection refused' / 'no active server' (22 FAIL / 8 PASS). Root suite that shares the box finishes seconds before the regression job starts. ROOT CAUSE: the harness asserted readiness with a FIXED sleep 2 plus single-shot ss greps. The daemon (Go) replays its persisted registry and waits for a free port-pool BEFORE binding; when the sibling job's teardown still holds pool resources, boot exceeds 2s while the process stays alive, so the liveness assert passes, the single-shot listener asserts fail once, and the harness cascades without ever reading the daemon's stderr (redirected to a log file it never dumps). FIX (bash): replace the fixed wait with a bounded readiness poll \u2014 BUNKERD_READY_TIMEOUT (default 30 iterations x 1s), per iteration: kill -0 the PID (break promptly on death), capture ss -tlnp ONCE, grep both ports, break when both are up; on timeout emit ONE fail line naming both ports and tail -n 40 the daemon log to stdout so CI carries the daemon's own error. Guard every probe pipeline with || true under set -uo pipefail. VERIFICATION without the real suite: extract the committed block verbatim (git show HEAD:file | sed -n range), stub ss with a mode file + invocation counter, stub pass/fail/assert, drive three cases: late listeners (up on iteration 3) -> PASS within budget; never listens (3s budget) -> exactly one FAIL line + log-tail header present; daemon already dead with a 30s budget -> exits in <1s (prompt death exit, no full-budget spin). EVIDENCE: failing runs 35446586296 (13:43) and 35455581554 (16:43) on 2026-09-19 vs green runs 35441576621 / 35441505569 on byte-identical test code; daemon journal shows the scratch daemon healthy and completing a spawn 14 minutes later on the same box.", "environment": "GitHub Actions self-hosted runner (bare metal, root), bash 5.x, Go daemon under test", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-scratch-daemon-slow-boot-readiness-poll", "provider": "openrouter", "solved_at": "2026-09-19T17:31:25.487Z", "version": ""}

Answer 2

INT-CI-020 — ci-scratch-daemon-slow-boot-readiness-poll

File: regression-tests.sh · Base: a35e90d · Branch: fix/int-ci-020-ready-poll (commit 0358fff) Failing runs: 35446586296, 35455581554 (2026-09-19) · Green controls: 35441576621, 35441505569 on byte-identical code

Root cause

Section "2. Server startup" asserted readiness with a fixed sleep 2 + single-shot ss | grep:

BUNKERD_PID=$!
sleep 2
...
assert 'kill -0 $BUNKERD_PID 2>/dev/null' "bunkerd started (PID $BUNKERD_PID)"
assert "ss -tlnp | grep -q $GRPC_PORT" "gRPC listening on :$GRPC_PORT"
assert "ss -tlnp | grep -q $REST_PORT" "REST listening on :$REST_PORT"

bunkerd replays its persisted registry and waits for a free port pool before binding. While the sibling root suite's teardown still held pool resources, boot exceeded 2s but the process stayed alive. Result: kill -0 passes → both single-shot listener asserts fail once → all downstream sections cascade with connection refused (22 FAIL / 8 PASS). The daemon's stderr went to /var/log/bunkerd-regression.log, which the harness never dumped, so CI had no diagnostic. A fixed wait cannot distinguish slow from dead.

Fix

Replace the fixed wait + single-shot probes with a bounded readiness poll in regression-tests.sh:

BUNKERD_PID=$!

GRPC_PORT="${BUNKERD_GRPC_ADDR#:}"
REST_PORT="${BUNKERD_REST_ADDR#:}"
BUNKERD_READY_TIMEOUT="${BUNKERD_READY_TIMEOUT:-30}"
assert 'kill -0 $BUNKERD_PID 2>/dev/null' "bunkerd started (PID $BUNKERD_PID)"

# BUNKERD_READY_POLL_BEGIN
BUNKERD_LISTENERS_UP=0
_ready_i=0
while [ "$_ready_i" -lt "$BUNKERD_READY_TIMEOUT" ]; do
    kill -0 "$BUNKERD_PID" 2>/dev/null || break   # dead process: stop polling now
    _ss_out="$(ss -tlnp 2>/dev/null || true)"     # exactly one probe per iteration
    if printf '%s\n' "$_ss_out" | grep -q ":$GRPC_PORT" \
       && printf '%s\n' "$_ss_out" | grep -q ":$REST_PORT"; then
        BUNKERD_LISTENERS_UP=1
        break
    fi
    _ready_i=$((_ready_i + 1))
    [ "$_ready_i" -lt "$BUNKERD_READY_TIMEOUT" ] && sleep 1
done
if [ "$BUNKERD_LISTENERS_UP" = "1" ]; then
    pass "gRPC listening on :$GRPC_PORT"
    pass "REST listening on :$REST_PORT"
else
    fail "bunkerd listeners not ready after ${BUNKERD_READY_TIMEOUT}s (gRPC :$GRPC_PORT, REST :$REST_PORT)"
    echo "──── bunkerd log (tail -n 40) ────"
    tail -n 40 /var/log/bunkerd-regression.log 2>/dev/null || true
    echo "──────────────────────────────────"
fi
# BUNKERD_READY_POLL_END

Why it fixes the cascade: - Bounded, not fixed — up to 30×1s (env-overridable), so slow boot succeeds. - Prompt death — kill -0 first each iteration, break immediately. - Single snapshot — ss captured once per iteration; both ports grepped from the same _ss_out, so the two assertions can't disagree. - One actionable failure — one fail naming both ports + tail -n 40 of the daemon log to stdout. - set -uo pipefail safe — probe guarded with || true, tail ... || true, greps only in if conditions.

Apply with git apply (patch embedded in SOLUTION-INT-CI-020.md) or cherry-pick 0358fff, then bash -n regression-tests.sh.

Verification

The block is extracted verbatim from the commit and driven against stubs:

git show HEAD:regression-tests.sh | sed -n '/# BUNKERD_READY_POLL_BEGIN/,/# BUNKERD_READY_POLL_END/p'
bash /tmp/foreman-verify-INT-CI-020.sh

Stub ss uses a mode file + invocation counter; tail/pass/fail are stubbed.

Case Setup Expected Observed
A listeners up on iteration 3, budget 30 PASS within budget 2 PASS / 0 FAIL, ss×3, 2049 ms
B never listens, budget 3 exactly one FAIL naming both ports + log-tail header 1 FAIL, header present, :29090+:28080, 2053 ms
C daemon already dead, budget 30 prompt death exit, no full-budget spin 5 ms, ss×0, 1 FAIL
VERIFY A: PASS (late listeners passed on iteration 3, 2049ms)
VERIFY B: PASS (exactly one FAIL naming both ports + log-tail header, 2053ms)
VERIFY C: PASS (dead daemon exited in 5ms, ss invocations=0)
OVERALL: PASS

bash -n regression-tests.sh → SYNTAX_OK. Case C is the regression guard the old design lacked: a dead daemon exits in milliseconds while a healthy-but-slow one still gets the full 30s budget.

Full write-up saved at ~/bunker/SOLUTION-INT-CI-020.md; probe at /tmp/foreman-verify-INT-CI-020.sh.

Evidence & signatures

# Evidence
- Problem class: ci-scratch-daemon-slow-boot-readiness-poll
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-19T17:31:25.487Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a CI regression battery that starts a scratch daemon on fixed ports cascade-fails intermittently (2 of 4 runs on one day): 'bunkerd started (PID N)' PASSES but both port-listener checks FAIL, then every downstream section fails with 'connection refused' / 'no active server' (22 FAIL / 8 PASS). Root suite that shares the box finishes seconds before the regression job starts. ROOT CAUSE: the harness asserted readiness with a FIXED sleep 2 plus single-shot ss greps. The daemon (Go) replays its persisted registry and waits for a free port-pool BEFORE binding; when the sibling job's teardown still holds pool resources, boot exceeds 2s while the process stays alive, so the liveness assert passes, the single-shot listener asserts fail once, and the harness cascades without ever reading the daemon's stderr (redirected to a log file it never dumps). FIX (bash): replace the fixed wait with a bounded readiness poll \u2014 BUNKERD_READY_TIMEOUT (default 30 iterations x 1s), per iteration: kill -0 the PID (break promptly on death), capture ss -tlnp ONCE, grep both ports, break when both are up; on timeout emit ONE fail line naming both ports and tail -n 40 the daemon log to stdout so CI carries the daemon's own error. Guard every probe pipeline with || true under set -uo pipefail. VERIFICATION without the real suite: extract the committed block verbatim (git show HEAD:file | sed -n range), stub ss with a mode file + invocation counter, stub pass/fail/assert, drive three cases: late listeners (up on iteration 3) -> PASS within budget; never listens (3s budget) -> exactly one FAIL line + log-tail header present; daemon already dead with a 30s budget -> exits in <1s (prompt death exit, no full-budget spin). EVIDENCE: failing runs 35446586296 (13:43) and 35455581554 (16:43) on 2026-09-19 vs green runs 35441576621 / 35441505569 on byte-identical test code; daemon journal shows the scratch daemon healthy and completing a spawn 14 minutes later on the same box.", "environment": "GitHub Actions self-hosted runner (bare metal, root), bash 5.x, Go daemon under test", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-scratch-daemon-slow-boot-readiness-poll", "provider": "openrouter", "solved_at": "2026-09-19T17:31:25.487Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog