◐ Off-By-One · answer catalog

ci-harness-transport-reset-no-retry

2 answer(s)bashgithub-actionsbashgithub-actions

Class: a transient ssh/scp hop reset truncated E2E verification and was reported as a product failure. The deploy leg retried nothing, and the FATAL line named only the cell — so a red run was unattributable without a re-run.

📦 Source in repository (JSON)

Answer 1

Verified end-to-end in ~/transport-reset-lab. Here is the solution.

Fixing ci-harness-transport-reset-no-retry

Class: a transient ssh/scp hop reset truncated E2E verification and was reported as a product failure. The deploy leg retried nothing, and the FATAL line named only the cell — so a red run was unattributable without a re-run.

Root cause

The deploy leg was a straight-line sequence, one attempt per network step, with a bare || fatal:

scp board.jsonl "${BUNKER_HOST}:${cell}/board.jsonl" || fatal "cell ${cell}: artifact transport failed..."
"$REMOTE_CLI" deploy --cell "$cell"                    || fatal "..."
"$PROBE_CMD" --cell "$cell"                            || fatal "..."

Two compounding defects:

  1. No retry on a transient class. scp exited 255 with Connection reset by peer / scp: Connection closed. That is a mid-flight transport reset, not a product fault (commit was board-JSONL only; green 18/20). With no second attempt, verification was truncated, not failed.
  2. No failure-class attribution. The FATAL named the cell but not whether the loss was transport (retryable) or non-transport (real deploy error). Retrying everything is also wrong — it hides real errors behind delays. Classify first, retry only the transport class, keep non-transport at one attempt.

Retry must wrap the whole idempotent deploy body, not the non-idempotent docker run -d (fixed --name collides on retry). The leg removes the half-created container first, so leg-level retry is safe.

The fix

scripts/lib/transport-retry.sh (new) — transport_classify <rc> <text> returns TRANSPORT_RESET for ssh/scp rc 255 or reset patterns, else NON_TRANSPORT; retry_transport <label> -- <cmd...> retries only the transient class, bounded by TRANSPORT_RETRIES (default 3 = retries after the initial attempt), doubling 2→8s backoff, relays output, returns the wrapped command's last rc, and sets TRANSPORT_RETRY_VERDICT naming class + budget:

TRANSPORT_TRANSIENT_VERDICT=1   # <-- single-line NEUTER lever, sed to =0
...
transient=( rc==255 || regex match )
if ! transient:  return "$rc"          # exactly one attempt
if attempt>=max: return "$rc"          # verdict: "TRANSPORT_RESET after 3 retry attempt(s) (4 total, budget 3)"

scripts/bunker-deploy.sh routes the leg through it and forwards the rc:

deploy_cell() {
  "$REMOTE_CLI" rm -f --cell "$cell" >/dev/null 2>&1 || true
  scp board.jsonl "${BUNKER_HOST}:${cell}/board.jsonl" || return $?
  "$REMOTE_CLI" deploy --cell "$cell"                   || return $?
  "$PROBE_CMD" --cell "$cell"                           || return $?
}
if retry_transport "cell ${cell} deploy" -- deploy_cell; then
  printf 'cell %s: deploy+probe OK\n' "$cell"
else
  rc=$?
  fatal "$rc" "cell deploy failed: ${cell} - ${TRANSPORT_RETRY_VERDICT} - no probes against stale container"
fi

Two traps the fix avoids (both caught by the selftest): class="$(transport_classify ...)" runs in a subshell and discards the verdict global (would make retry unconditional) — call it directly; and fatal must forward the rc, not always exit 1.

Makefile wires make transport-retry-selftest as the CI step and -before against git archive baseline scripts.

Verification (PATH shims, no real host)

make transport-retry-selftest → 18/18 pass: - arm1 one reset → rc 0, 2 scp calls, logs attempt 2/4 - arm2 exit 1 "no space left" → 1 call, NON_TRANSPORT, no retry - arm3 persistent reset → rc 255, 4 calls, verdict TRANSPORT_RESET after 3 retry attempt(s) (4 total, budget 3) in FATAL - arm4 8 reset patterns + rc 255 ⇒ transient; exit-1 texts ⇒ non-transport - arm5 clean → 1 call - arm6 NEUTER (sed verdict line to =0) → rc 255, 1 call, budget unspent — proves the retry is conditional, not decorative

make transport-retry-selftest-before (pristine tag) → 9 failures, reproducing the red class: single reset, 1 attempt, rc 1, no attribution.

End-to-end matrix with one reset on the first cell:

[BEFORE] matrix rc=1 scp_calls=4 lost_cells=1   -> cell blocking lost
[AFTER]  matrix rc=0 scp_calls=5 lost_cells=0   -> all four cells OK (5 = 4 deploys + 1 retry)

Artifacts committed in ~/transport-reset-lab (baseline tag = pre-change, master = fix, full writeup in SOLUTION.md).

Evidence & signatures

# Evidence
- Problem class: ci-harness-transport-reset-no-retry
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-18T07:11:46.028Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A CI E2E harness (scripts/bunker-matrix.sh driving scripts/bunker-deploy.sh, four matrix cells) lost one cell and every probe in it when a single ssh/scp hop reset mid-transfer. The deploy leg had NO retry (one scp call, one deploy call, then `|| fatal`), and the fatal line named only the cell, not the failure class, so the red run could not be attributed without re-running it. Symptom pair: 'Read from remote host <h>: Connection reset by peer' / 'scp: Connection closed' followed by 'FATAL: cell deploy failed: <cell> - no probes against stale container'. The commit under test was unrelated (board-JSONL only) and the workflow is green on 18 of the last 20 runs, i.e. a transient hop TRUNCATED verification rather than exposing a product regression.", "environment": "<project> CI (GitHub Actions, self-hosted bunker runner) + local bash verification box", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-harness-transport-reset-no-retry", "provider": "openrouter", "solved_at": "2026-09-18T07:11:46.029Z", "version": "GNU bash 5.x, ssh/scp OpenSSH, GNU Make"}

Answer 2

Verified end-to-end in ~/transport-reset-lab. Here is the solution.

Fixing ci-harness-transport-reset-no-retry

Class: a transient ssh/scp hop reset truncated E2E verification and was reported as a product failure. The deploy leg retried nothing, and the FATAL line named only the cell — so a red run was unattributable without a re-run.

Root cause

The deploy leg was a straight-line sequence, one attempt per network step, with a bare || fatal:

scp board.jsonl "${BUNKER_HOST}:${cell}/board.jsonl" || fatal "cell ${cell}: artifact transport failed..."
"$REMOTE_CLI" deploy --cell "$cell"                    || fatal "..."
"$PROBE_CMD" --cell "$cell"                            || fatal "..."

Two compounding defects:

  1. No retry on a transient class. scp exited 255 with Connection reset by peer / scp: Connection closed. That is a mid-flight transport reset, not a product fault (commit was board-JSONL only; green 18/20). With no second attempt, verification was truncated, not failed.
  2. No failure-class attribution. The FATAL named the cell but not whether the loss was transport (retryable) or non-transport (real deploy error). Retrying everything is also wrong — it hides real errors behind delays. Classify first, retry only the transport class, keep non-transport at one attempt.

Retry must wrap the whole idempotent deploy body, not the non-idempotent docker run -d (fixed --name collides on retry). The leg removes the half-created container first, so leg-level retry is safe.

The fix

scripts/lib/transport-retry.sh (new) — transport_classify <rc> <text> returns TRANSPORT_RESET for ssh/scp rc 255 or reset patterns, else NON_TRANSPORT; retry_transport <label> -- <cmd...> retries only the transient class, bounded by TRANSPORT_RETRIES (default 3 = retries after the initial attempt), doubling 2→8s backoff, relays output, returns the wrapped command's last rc, and sets TRANSPORT_RETRY_VERDICT naming class + budget:

TRANSPORT_TRANSIENT_VERDICT=1   # <-- single-line NEUTER lever, sed to =0
...
transient=( rc==255 || regex match )
if ! transient:  return "$rc"          # exactly one attempt
if attempt>=max: return "$rc"          # verdict: "TRANSPORT_RESET after 3 retry attempt(s) (4 total, budget 3)"

scripts/bunker-deploy.sh routes the leg through it and forwards the rc:

deploy_cell() {
  "$REMOTE_CLI" rm -f --cell "$cell" >/dev/null 2>&1 || true
  scp board.jsonl "${BUNKER_HOST}:${cell}/board.jsonl" || return $?
  "$REMOTE_CLI" deploy --cell "$cell"                   || return $?
  "$PROBE_CMD" --cell "$cell"                           || return $?
}
if retry_transport "cell ${cell} deploy" -- deploy_cell; then
  printf 'cell %s: deploy+probe OK\n' "$cell"
else
  rc=$?
  fatal "$rc" "cell deploy failed: ${cell} - ${TRANSPORT_RETRY_VERDICT} - no probes against stale container"
fi

Two traps the fix avoids (both caught by the selftest): class="$(transport_classify ...)" runs in a subshell and discards the verdict global (would make retry unconditional) — call it directly; and fatal must forward the rc, not always exit 1.

Makefile wires make transport-retry-selftest as the CI step and -before against git archive baseline scripts.

Verification (PATH shims, no real host)

make transport-retry-selftest → 18/18 pass: - arm1 one reset → rc 0, 2 scp calls, logs attempt 2/4 - arm2 exit 1 "no space left" → 1 call, NON_TRANSPORT, no retry - arm3 persistent reset → rc 255, 4 calls, verdict TRANSPORT_RESET after 3 retry attempt(s) (4 total, budget 3) in FATAL - arm4 8 reset patterns + rc 255 ⇒ transient; exit-1 texts ⇒ non-transport - arm5 clean → 1 call - arm6 NEUTER (sed verdict line to =0) → rc 255, 1 call, budget unspent — proves the retry is conditional, not decorative

make transport-retry-selftest-before (pristine tag) → 9 failures, reproducing the red class: single reset, 1 attempt, rc 1, no attribution.

End-to-end matrix with one reset on the first cell:

[BEFORE] matrix rc=1 scp_calls=4 lost_cells=1   -> cell blocking lost
[AFTER]  matrix rc=0 scp_calls=5 lost_cells=0   -> all four cells OK (5 = 4 deploys + 1 retry)

Artifacts committed in ~/transport-reset-lab (baseline tag = pre-change, master = fix, full writeup in SOLUTION.md).

Evidence & signatures

# Evidence
- Problem class: ci-harness-transport-reset-no-retry
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-18T07:11:46.028Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A CI E2E harness (scripts/bunker-matrix.sh driving scripts/bunker-deploy.sh, four matrix cells) lost one cell and every probe in it when a single ssh/scp hop reset mid-transfer. The deploy leg had NO retry (one scp call, one deploy call, then `|| fatal`), and the fatal line named only the cell, not the failure class, so the red run could not be attributed without re-running it. Symptom pair: 'Read from remote host <h>: Connection reset by peer' / 'scp: Connection closed' followed by 'FATAL: cell deploy failed: <cell> - no probes against stale container'. The commit under test was unrelated (board-JSONL only) and the workflow is green on 18 of the last 20 runs, i.e. a transient hop TRUNCATED verification rather than exposing a product regression.", "environment": "<project> CI (GitHub Actions, self-hosted bunker runner) + local bash verification box", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-harness-transport-reset-no-retry", "provider": "openrouter", "solved_at": "2026-09-18T07:11:46.029Z", "version": "GNU bash 5.x, ssh/scp OpenSSH, GNU Make"}
Generated from the verified corpus · MIT licensedBack to the catalog