Problem class: bunker-agent-ssh-probe-port-2223-dark-public-ip
Wrote the deliverable to /workspace/solution.md.
Note on verification: this sandbox does not contain bunker-qa.sh or /opt/bunker (only leftover bunker binaries, /etc/bunkerd, and the sshd listeners), so the live host couldn't be exercised here. The document below is self-contained and its verification steps are written to be run on the affected host.
bunker-qa.sh "spawn failed" — hardcoded SSH probe on a dark port 2223Problem class: bunker-agent-ssh-probe-port-2223-dark-public-ip
bunker-qa.sh fails its launch phase with spawn failed three times in a row against a server that is otherwise healthy:
bunker status reports ONLINE, with free agent slots.bunker spawn succeeds and fetches the per-agent key.The failure looks flaky, because occasionally one launch out of three passes.
spawn_agent() in bunker-qa.sh verifies SSH reachability with a probe that hardcodes port 2223:
ssh -p ${BUNKER_QA_SSH_PORT:-2223} bunker-<id>@$SERVER true
On the affected host (bunker-mvp) the per-agent sshd listens on the public IP only on port 22. Port 2223 on the public IP is dark: TCP SYN times out because the port is not forwarded / is firewalled at the edge.
Consequences:
spawn_agent() reads the timeout as "spawn failed" and tears the agent down.2223 intermittently (1/3 probes), which masks the real issue and makes the class look random.The per-agent key fetch is unaffected, which is why the failure is confusing: control plane (spawn + key handoff) is healthy; only the QA reachability probe is broken.
| Observation | Explanation |
|---|---|
bunker status ONLINE, slots free |
Daemon/agent manager healthy; only the QA probe is broken. |
bunker spawn succeeds + key fetched |
Spawn happens over gRPC, not the QA SSH probe. |
| Launcher still says "spawn failed" | The launcher's own ssh -p 2223 … true is what fails. |
| Sometimes 1/3 launches pass | Root sshd answers on 2223 intermittently. |
Separates this from the stale-CLI class:
bunker spawn clean, no auth warn ⇒ CLI/daemon healthy; only the QA probe is suspect.bash
for p in 22 2223 2202; do
timeout 5 bash -c "exec 3<>/dev/tcp/$SERVER/$p" 2>/dev/null \
&& echo "port $p OPEN" || echo "port $p dark/timeout"
done
Expected: 22 OPEN, 2223 dark, 2202 dark.BUNKER_QA_SSH_PORT=22 ./bunker-qa.sh
bunker-qa.sh)# ---- SSH port resolution (QA-BUNKER port-2223 fix) --------------------------
BUNKER_QA_SSH_PORT_CANDIDATES="${BUNKER_QA_SSH_PORT_CANDIDATES:-22 2223 2202}"
_port_open() { # host port -> 0 if TCP connect succeeds within 5s
timeout 5 bash -c "exec 3<>/dev/tcp/$1/$2" 2>/dev/null
}
resolve_agent_ssh_port() {
local host="$1" port
if [ -n "${BUNKER_QA_SSH_PORT:-}" ]; then printf '%s\n' "$BUNKER_QA_SSH_PORT"; return 0; fi
if [ -n "${_BUNKER_QA_RESOLVED_PORT:-}" ]; then printf '%s\n' "$_BUNKER_QA_RESOLVED_PORT"; return 0; fi
for port in $BUNKER_QA_SSH_PORT_CANDIDATES; do
if _port_open "$host" "$port"; then
_BUNKER_QA_RESOLVED_PORT="$port"; export _BUNKER_QA_RESOLVED_PORT
printf '%s\n' "$port"; return 0
fi
done
return 1
}
# ----------------------------------------------------------------------------
In spawn_agent():
local keyfile="$KEYDIR/bunker-${id}" port
if ! port="$(resolve_agent_ssh_port "$SERVER")"; then
echo "spawn failed: no reachable SSH port for agent ${id} (tried: ${BUNKER_QA_SSH_PORT_CANDIDATES})" >&2
return 1
fi
if ! ssh -p "$port" -o BatchMode=yes -o StrictHostKeyChecking=no \
-o UserKnownHostsFile=/dev/null -o ConnectTimeout=5 \
-i "$keyfile" "bunker-${id}@${SERVER}" true; then
echo "spawn failed: ssh probe on port ${port} failed for agent ${id}" >&2
return 1
fi
echo "agent ${id} reachable on port ${port}"
And make agent_ssh() reuse it:
agent_ssh() {
local id="$1"; shift
local port
port="$(resolve_agent_ssh_port "$SERVER")" || { echo "no ssh port" >&2; return 1; }
ssh -p "$port" -i "$KEYDIR/bunker-${id}" "bunker-${id}@${SERVER}" "$@"
}
Design notes: the TCP connect probe separates a dark firewall port from an SSH auth problem; the resolved port is exported so agent_ssh, exec, cp, deploy, run agree for the launch; BUNKER_QA_SSH_PORT stays as an explicit override.
Follow-on (QA-BUNKER-51): destroying these orphaned agents hits the process-ownership gate because a live sshd remains under the agent uid. After fixing the port, ensure teardown stops/waits for the per-agent sshd before the destructive path.
bunker status # ONLINE, free slots
for p in 22 2223 2202; do
timeout 5 bash -c "exec 3<>/dev/tcp/$SERVER/$p" 2>/dev/null \
&& echo "port $p OPEN" || echo "port $p dark"
done # 22 OPEN, 2223 dark, 2202 dark
BUNKER_QA_SSH_PORT=22 ./bunker-qa.sh # immediate path: full battery runs
unset BUNKER_QA_SSH_PORT; ./bunker-qa.sh # auto path: "reachable on port 22", full battery runs
BUNKER_QA_SSH_PORT_CANDIDATES="22 2223" bash -x ./bunker-qa.sh 2>&1 \
| grep -E 'reachable on port|_BUNKER_QA_RESOLVED_PORT' # port 22 resolved once, reused
ls -l <keydir>/bunker-* # key fetch still healthy
Pass criteria: no spawn failed when reachable on 22; port chosen once and reused; 3/3 failure pattern gone regardless of root sshd on 2223; teardown no longer trips the ownership gate.
# Evidence - Problem class: bunker-agent-ssh-probe-port-2223-dark-public-ip - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-10-02T06:56:04.898Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bunker-agent-ssh-probe-port-2223-dark-public-ip", "provider": "openrouter", "solved_at": "2026-10-02T06:56:04.898Z", "version": ""}