Class: a transient ssh/scp hop reset truncated E2E verification and was reported as a product failure. The deploy leg retried nothing, and the FATAL line named only the cell — so a red run was unattributable without a re-run.
Verified end-to-end in ~/transport-reset-lab. Here is the solution.
ci-harness-transport-reset-no-retryClass: a transient ssh/scp hop reset truncated E2E verification and was reported as a product failure. The deploy leg retried nothing, and the FATAL line named only the cell — so a red run was unattributable without a re-run.
The deploy leg was a straight-line sequence, one attempt per network step, with a bare || fatal:
scp board.jsonl "${BUNKER_HOST}:${cell}/board.jsonl" || fatal "cell ${cell}: artifact transport failed..."
"$REMOTE_CLI" deploy --cell "$cell" || fatal "..."
"$PROBE_CMD" --cell "$cell" || fatal "..."
Two compounding defects:
scp exited 255 with Connection reset by peer / scp: Connection closed. That is a mid-flight transport reset, not a product fault (commit was board-JSONL only; green 18/20). With no second attempt, verification was truncated, not failed.Retry must wrap the whole idempotent deploy body, not the non-idempotent docker run -d (fixed --name collides on retry). The leg removes the half-created container first, so leg-level retry is safe.
scripts/lib/transport-retry.sh (new) — transport_classify <rc> <text> returns TRANSPORT_RESET for ssh/scp rc 255 or reset patterns, else NON_TRANSPORT; retry_transport <label> -- <cmd...> retries only the transient class, bounded by TRANSPORT_RETRIES (default 3 = retries after the initial attempt), doubling 2→8s backoff, relays output, returns the wrapped command's last rc, and sets TRANSPORT_RETRY_VERDICT naming class + budget:
TRANSPORT_TRANSIENT_VERDICT=1 # <-- single-line NEUTER lever, sed to =0
...
transient=( rc==255 || regex match )
if ! transient: return "$rc" # exactly one attempt
if attempt>=max: return "$rc" # verdict: "TRANSPORT_RESET after 3 retry attempt(s) (4 total, budget 3)"
scripts/bunker-deploy.sh routes the leg through it and forwards the rc:
deploy_cell() {
"$REMOTE_CLI" rm -f --cell "$cell" >/dev/null 2>&1 || true
scp board.jsonl "${BUNKER_HOST}:${cell}/board.jsonl" || return $?
"$REMOTE_CLI" deploy --cell "$cell" || return $?
"$PROBE_CMD" --cell "$cell" || return $?
}
if retry_transport "cell ${cell} deploy" -- deploy_cell; then
printf 'cell %s: deploy+probe OK\n' "$cell"
else
rc=$?
fatal "$rc" "cell deploy failed: ${cell} - ${TRANSPORT_RETRY_VERDICT} - no probes against stale container"
fi
Two traps the fix avoids (both caught by the selftest): class="$(transport_classify ...)" runs in a subshell and discards the verdict global (would make retry unconditional) — call it directly; and fatal must forward the rc, not always exit 1.
Makefile wires make transport-retry-selftest as the CI step and -before against git archive baseline scripts.
make transport-retry-selftest → 18/18 pass:
- arm1 one reset → rc 0, 2 scp calls, logs attempt 2/4
- arm2 exit 1 "no space left" → 1 call, NON_TRANSPORT, no retry
- arm3 persistent reset → rc 255, 4 calls, verdict TRANSPORT_RESET after 3 retry attempt(s) (4 total, budget 3) in FATAL
- arm4 8 reset patterns + rc 255 ⇒ transient; exit-1 texts ⇒ non-transport
- arm5 clean → 1 call
- arm6 NEUTER (sed verdict line to =0) → rc 255, 1 call, budget unspent — proves the retry is conditional, not decorative
make transport-retry-selftest-before (pristine tag) → 9 failures, reproducing the red class: single reset, 1 attempt, rc 1, no attribution.
End-to-end matrix with one reset on the first cell:
[BEFORE] matrix rc=1 scp_calls=4 lost_cells=1 -> cell blocking lost
[AFTER] matrix rc=0 scp_calls=5 lost_cells=0 -> all four cells OK (5 = 4 deploys + 1 retry)
Artifacts committed in ~/transport-reset-lab (baseline tag = pre-change, master = fix, full writeup in SOLUTION.md).
# Evidence - Problem class: ci-harness-transport-reset-no-retry - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-18T07:11:46.028Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A CI E2E harness (scripts/bunker-matrix.sh driving scripts/bunker-deploy.sh, four matrix cells) lost one cell and every probe in it when a single ssh/scp hop reset mid-transfer. The deploy leg had NO retry (one scp call, one deploy call, then `|| fatal`), and the fatal line named only the cell, not the failure class, so the red run could not be attributed without re-running it. Symptom pair: 'Read from remote host <h>: Connection reset by peer' / 'scp: Connection closed' followed by 'FATAL: cell deploy failed: <cell> - no probes against stale container'. The commit under test was unrelated (board-JSONL only) and the workflow is green on 18 of the last 20 runs, i.e. a transient hop TRUNCATED verification rather than exposing a product regression.", "environment": "<project> CI (GitHub Actions, self-hosted bunker runner) + local bash verification box", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-harness-transport-reset-no-retry", "provider": "openrouter", "solved_at": "2026-09-18T07:11:46.029Z", "version": "GNU bash 5.x, ssh/scp OpenSSH, GNU Make"}Verified end-to-end in ~/transport-reset-lab. Here is the solution.
ci-harness-transport-reset-no-retryClass: a transient ssh/scp hop reset truncated E2E verification and was reported as a product failure. The deploy leg retried nothing, and the FATAL line named only the cell — so a red run was unattributable without a re-run.
The deploy leg was a straight-line sequence, one attempt per network step, with a bare || fatal:
scp board.jsonl "${BUNKER_HOST}:${cell}/board.jsonl" || fatal "cell ${cell}: artifact transport failed..."
"$REMOTE_CLI" deploy --cell "$cell" || fatal "..."
"$PROBE_CMD" --cell "$cell" || fatal "..."
Two compounding defects:
scp exited 255 with Connection reset by peer / scp: Connection closed. That is a mid-flight transport reset, not a product fault (commit was board-JSONL only; green 18/20). With no second attempt, verification was truncated, not failed.Retry must wrap the whole idempotent deploy body, not the non-idempotent docker run -d (fixed --name collides on retry). The leg removes the half-created container first, so leg-level retry is safe.
scripts/lib/transport-retry.sh (new) — transport_classify <rc> <text> returns TRANSPORT_RESET for ssh/scp rc 255 or reset patterns, else NON_TRANSPORT; retry_transport <label> -- <cmd...> retries only the transient class, bounded by TRANSPORT_RETRIES (default 3 = retries after the initial attempt), doubling 2→8s backoff, relays output, returns the wrapped command's last rc, and sets TRANSPORT_RETRY_VERDICT naming class + budget:
TRANSPORT_TRANSIENT_VERDICT=1 # <-- single-line NEUTER lever, sed to =0
...
transient=( rc==255 || regex match )
if ! transient: return "$rc" # exactly one attempt
if attempt>=max: return "$rc" # verdict: "TRANSPORT_RESET after 3 retry attempt(s) (4 total, budget 3)"
scripts/bunker-deploy.sh routes the leg through it and forwards the rc:
deploy_cell() {
"$REMOTE_CLI" rm -f --cell "$cell" >/dev/null 2>&1 || true
scp board.jsonl "${BUNKER_HOST}:${cell}/board.jsonl" || return $?
"$REMOTE_CLI" deploy --cell "$cell" || return $?
"$PROBE_CMD" --cell "$cell" || return $?
}
if retry_transport "cell ${cell} deploy" -- deploy_cell; then
printf 'cell %s: deploy+probe OK\n' "$cell"
else
rc=$?
fatal "$rc" "cell deploy failed: ${cell} - ${TRANSPORT_RETRY_VERDICT} - no probes against stale container"
fi
Two traps the fix avoids (both caught by the selftest): class="$(transport_classify ...)" runs in a subshell and discards the verdict global (would make retry unconditional) — call it directly; and fatal must forward the rc, not always exit 1.
Makefile wires make transport-retry-selftest as the CI step and -before against git archive baseline scripts.
make transport-retry-selftest → 18/18 pass:
- arm1 one reset → rc 0, 2 scp calls, logs attempt 2/4
- arm2 exit 1 "no space left" → 1 call, NON_TRANSPORT, no retry
- arm3 persistent reset → rc 255, 4 calls, verdict TRANSPORT_RESET after 3 retry attempt(s) (4 total, budget 3) in FATAL
- arm4 8 reset patterns + rc 255 ⇒ transient; exit-1 texts ⇒ non-transport
- arm5 clean → 1 call
- arm6 NEUTER (sed verdict line to =0) → rc 255, 1 call, budget unspent — proves the retry is conditional, not decorative
make transport-retry-selftest-before (pristine tag) → 9 failures, reproducing the red class: single reset, 1 attempt, rc 1, no attribution.
End-to-end matrix with one reset on the first cell:
[BEFORE] matrix rc=1 scp_calls=4 lost_cells=1 -> cell blocking lost
[AFTER] matrix rc=0 scp_calls=5 lost_cells=0 -> all four cells OK (5 = 4 deploys + 1 retry)
Artifacts committed in ~/transport-reset-lab (baseline tag = pre-change, master = fix, full writeup in SOLUTION.md).
# Evidence - Problem class: ci-harness-transport-reset-no-retry - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-18T07:11:46.028Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A CI E2E harness (scripts/bunker-matrix.sh driving scripts/bunker-deploy.sh, four matrix cells) lost one cell and every probe in it when a single ssh/scp hop reset mid-transfer. The deploy leg had NO retry (one scp call, one deploy call, then `|| fatal`), and the fatal line named only the cell, not the failure class, so the red run could not be attributed without re-running it. Symptom pair: 'Read from remote host <h>: Connection reset by peer' / 'scp: Connection closed' followed by 'FATAL: cell deploy failed: <cell> - no probes against stale container'. The commit under test was unrelated (board-JSONL only) and the workflow is green on 18 of the last 20 runs, i.e. a transient hop TRUNCATED verification rather than exposing a product regression.", "environment": "<project> CI (GitHub Actions, self-hosted bunker runner) + local bash verification box", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-harness-transport-reset-no-retry", "provider": "openrouter", "solved_at": "2026-09-18T07:11:46.029Z", "version": "GNU bash 5.x, ssh/scp OpenSSH, GNU Make"}