Problem class: df-off-by-one-1-live-probe
:8766 Live Probe — Does a Fresh Server Complete a Solve?Problem class: df-off-by-one-1-live-probe
Verdict: ✅ Yes. The freshly started off-by-one.service on *:8766 is healthy and completes the full solve chain end-to-end. This solve is itself the proof: it is running inside a sandbox spawned by that server.
The probe is a canary/self-test of the whole pre-solve pipeline, not of a single function. A "solve" only completes if every link below is green:
client / fleet
│ POST /api/v1/problems/submit
▼
off-by-one.service ── listens on *:8766 (HTTP API + catalog UI)
│ dequeue
▼
bwrap sandbox ── --unshare-all --share-net, workspace bind-mounted rw
│ exec
▼
~/.local/bin/pi-agent solve --problem-file … --output … --model …
│ spawn + stdin EOF
▼
pi ── provider/model (deepseek-v4-flash) with the selected API key
│ stdout
▼
/workspace/solution.md + evidence.md + signatures.json (bind mount)
│
▼
GitReins judge → verified → publish to catalog + repo
Any red link reproduces the symptom "the server accepts the submission but never produces a solved answer." The -live-probe variant is the server asking itself that question with context: {"probe": true}.
All commands below are read-only and were run from inside the probe sandbox.
$ curl -s http://<ip-address>:8766/health
{"status":"ok","uptime":"9m51s"}
$ curl -s http://<ip-address>:8766/api/v1/stats
{"total_problems":1484,"total_answers":1663,"verified_answers":1635,
"queue_depth":2,"hit_rate":0.98316,"coverage":1.10175,
"avg_solve_time":"2m40s","readonly":false,"solver_available":true}
$ ss -ltnp | grep 8766
LISTEN 0 4096 *:8766 *:*
$ ps -ef | grep -E 'bwrap|pi-agent| pi$'
/usr/bin/bwrap --unshare-all --share-net --die-with-parent --new-session \
--ro-bind /usr /usr --ro-bind /etc /etc ... \
--ro-bind ~/.local/bin ~/.local/bin \
--ro-bind /tmp/pi /tmp/pi \
--bind /tmp/off-by-one-sandbox-sub_c048bd /workspace \
-- ~/.local/bin/pi-agent solve \
--problem-file /workspace/problem.json \
--output /workspace/solution.md \
--evidence /workspace/evidence.md \
--signatures /workspace/signatures.json \
--model deepseek-v4-flash
node ~/.local/bin/pi-agent solve ...
pi
$ cat /proc/2/environ | tr '\0' '\n' | grep off-by-one
MEMORY_PRESSURE_WATCH=/sys/fs/cgroup/system.slice/off-by-one.service/memory.pressure
Interpretation:
solver_available:true, uptime advancing, and *:8766 in LISTEN → the server process is up on a fresh boot.off-by-one.service cgroup → it was launched by the service (not by a stray shell)./workspace is the server-created bind mount /tmp/off-by-one-sandbox-sub_c048bd, so this run's output lands where the server/judge can read it.sub_9b0880 … complete (1m42s). Historical failed rows are the chain regressions the probe exists to catch.Conclusion of the probe: a fresh :8766 server does complete a solve.
There is no defect to repair on the happy path. The value of the probe is that it turns any of the following into a hard failure. These are the real root causes behind a "server won't complete a solve" report, in the order they actually bite:
| # | Broken link | Symptom | Root cause |
|---|---|---|---|
| 1 | pi discovery | pi-agent: pi binary not found |
PI_BIN/$PI_HOME unset and neither /tmp/pi/pi nor the npm dist/cli.js layout exists |
| 2 | provider key | No API key found |
placeholder key passed isUsableKey, or provider-hygiene deleted the only key after PI_MODEL forced a lane |
| 3 | model id | cannot resolve model / wrong provider |
bare deepseek-v4-flash not mapped to a provider-qualified id |
| 4 | stdin | pi hangs until the whole deadline | pi waits for EOF on a non-TTY stdin; must be a real pipe that is closed (spawn + child.stdin.end()), not execFile/ignore |
| 5 | TLS env | error setting certificate file |
host-only SSL_CERT_FILE/CURL_CA_BUNDLE leaked into the sandbox |
| 6 | agent store | wrong-provider auth skipped | stale ~/.pi/agent store; must isolate via PI_CODING_AGENT_DIR |
| 7 | output path | solve succeeds, no solution.md |
workspace bind missing/read-only |
| 8 | timeouts | killed mid-answer | pi-agent cap (25 min) vs server cap (30 min) vs a slow thinking model |
The off-by-one packaging of this class refers to the two boundary conditions that are easy to get wrong and that the probe pins down:
position starting at 1 and queue_depth excludes completed rows.25 min < 30 min), or the last solve of a burst is SIGKILLed exactly when it would otherwise finish.Nothing in the live deployment needs changing. The reproducible "fix" is the deterministic bring-up + readiness gate that guarantees a fresh :8766 server completes solves. Run as the service user (kara) on the host:
# 0. One-time prerequisites
install -d -m 0755 ~/off-by-one
# ~/off-by-one/off-by-one <- server binary (mode 0755)
# ~/off-by-one/.env <- provider + tuning env
cat > ~/off-by-one/.env <<'EOF'
OPENROUTER_API_KEY=sk-or-v1-... # real key, >20 chars, no "example"/"changeme"
PI_MODEL=openrouter/deepseek/deepseek-v4.1-flash
EOF
chmod 600 ~/off-by-one/.env
# 1. Make the bridge discoverable and sane (matches the live lab layout)
install -d -m 0755 ~/.local/bin
test -x ~/.local/bin/pi-agent || install -m 0755 pi-agent ~/.local/bin/pi-agent
export PI_BIN=/tmp/pi/pi # ELF, or PI_HOME for the npm layout
test -x "$PI_BIN" || { echo "pi binary missing"; exit 1; }
# 2. (Re)start the service and wait for health
sudo systemctl daemon-reload
sudo systemctl restart off-by-one.service
for i in $(seq 1 30); do
curl -fsS http://<ip-address>:8766/health >/dev/null 2>&1 && break
sleep 1
done
curl -fsS http://<ip-address>:8766/health
Hardening that prevents the failure modes in §3 (already implemented in pi-agent; keep it that way):
delete env.SSL_CERT_FILE; delete env.CURL_CA_BUNDLE;spawn + child.stdin.end() for pi; never execFile.PI_CODING_AGENT_DIR=$(mktemp -d).inner_timeout (25m) < outer_timeout (30m).$ curl -fsS http://<ip-address>:8766/health # {"status":"ok",...}
$ curl -fsS http://<ip-address>:8766/api/v1/stats \
| grep -o '"solver_available":[a-z]*' # "solver_available":true
$ ss -ltn | grep -q ':8766' && echo "listening"
complete)SUBMIT=$(curl -fsS -X POST http://<ip-address>:8766/api/v1/problems/submit \
-H 'Content-Type: application/json' \
-d '{"problem_class":"probe-8766-liveness",
"description":"Return a one-line proof that the solve chain works.",
"environment":"linux","language":"","version":"","cadence":"once"}')
echo "$SUBMIT"
# -> {"problem_class":"...","status":"queued","position":N}
# poll until status is complete (or failed)
while :; do
curl -fsS http://<ip-address>:8766/api/v1/queue \
| tr ',' '\n' | grep -A2 '"problem_class":"probe-8766-liveness"' && break
sleep 10
done
Expected terminal state: "status":"complete","stage":"done", with solution.md, evidence.md, and signatures.json present in the workspace that the server created (/tmp/off-by-one-sandbox-sub_<id>).
The probe you are reading was dispatched exactly as:
off-by-one.service
└─ bwrap … --bind /tmp/off-by-one-sandbox-sub_c048bd /workspace
└─ pi-agent solve --problem-file /workspace/problem.json
--output /workspace/solution.md
--evidence /workspace/evidence.md
--signatures /workspace/signatures.json
--model deepseek-v4-flash
└─ pi (provider=openrouter, model=deepseek/deepseek-v4.1-flash)
In-sandbox checks that prove completion:
test -s /workspace/solution.md && echo "solution.md written"
test -s /workspace/evidence.md && echo "evidence.md written"
python3 -c 'import json;json.load(open("/workspace/signatures.json"));print("signatures.json valid")'
When those three pass and /health returns {"status":"ok"}, the answer to the probe is yes: a fresh server on :8766 completes a solve.
# Evidence - Problem class: df-off-by-one-1-live-probe - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-11T10:22:02.405Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Live probe: does a fresh server on :8766 complete a solve?", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "df-off-by-one-1-live-probe", "provider": "openrouter", "solved_at": "2026-09-11T10:22:02.406Z", "version": ""}:8766 Live Probe — Does a Fresh Server Complete a Solve?Problem class: df-off-by-one-1-live-probe
Verdict: ✅ Yes. The freshly started off-by-one.service on *:8766 is healthy and completes the full solve chain end-to-end. This solve is itself the proof: it is running inside a sandbox spawned by that server.
The probe is a canary/self-test of the whole pre-solve pipeline, not of a single function. A "solve" only completes if every link below is green:
client / fleet
│ POST /api/v1/problems/submit
▼
off-by-one.service ── listens on *:8766 (HTTP API + catalog UI)
│ dequeue
▼
bwrap sandbox ── --unshare-all --share-net, workspace bind-mounted rw
│ exec
▼
~/.local/bin/pi-agent solve --problem-file … --output … --model …
│ spawn + stdin EOF
▼
pi ── provider/model (deepseek-v4-flash) with the selected API key
│ stdout
▼
/workspace/solution.md + evidence.md + signatures.json (bind mount)
│
▼
GitReins judge → verified → publish to catalog + repo
Any red link reproduces the symptom "the server accepts the submission but never produces a solved answer." The -live-probe variant is the server asking itself that question with context: {"probe": true}.
All commands below are read-only and were run from inside the probe sandbox.
$ curl -s http://<ip-address>:8766/health
{"status":"ok","uptime":"9m51s"}
$ curl -s http://<ip-address>:8766/api/v1/stats
{"total_problems":1484,"total_answers":1663,"verified_answers":1635,
"queue_depth":2,"hit_rate":0.98316,"coverage":1.10175,
"avg_solve_time":"2m40s","readonly":false,"solver_available":true}
$ ss -ltnp | grep 8766
LISTEN 0 4096 *:8766 *:*
$ ps -ef | grep -E 'bwrap|pi-agent| pi$'
/usr/bin/bwrap --unshare-all --share-net --die-with-parent --new-session \
--ro-bind /usr /usr --ro-bind /etc /etc ... \
--ro-bind ~/.local/bin ~/.local/bin \
--ro-bind /tmp/pi /tmp/pi \
--bind /tmp/off-by-one-sandbox-sub_c048bd /workspace \
-- ~/.local/bin/pi-agent solve \
--problem-file /workspace/problem.json \
--output /workspace/solution.md \
--evidence /workspace/evidence.md \
--signatures /workspace/signatures.json \
--model deepseek-v4-flash
node ~/.local/bin/pi-agent solve ...
pi
$ cat /proc/2/environ | tr '\0' '\n' | grep off-by-one
MEMORY_PRESSURE_WATCH=/sys/fs/cgroup/system.slice/off-by-one.service/memory.pressure
Interpretation:
solver_available:true, uptime advancing, and *:8766 in LISTEN → the server process is up on a fresh boot.off-by-one.service cgroup → it was launched by the service (not by a stray shell)./workspace is the server-created bind mount /tmp/off-by-one-sandbox-sub_c048bd, so this run's output lands where the server/judge can read it.sub_9b0880 … complete (1m42s). Historical failed rows are the chain regressions the probe exists to catch.Conclusion of the probe: a fresh :8766 server does complete a solve.
There is no defect to repair on the happy path. The value of the probe is that it turns any of the following into a hard failure. These are the real root causes behind a "server won't complete a solve" report, in the order they actually bite:
| # | Broken link | Symptom | Root cause |
|---|---|---|---|
| 1 | pi discovery | pi-agent: pi binary not found |
PI_BIN/$PI_HOME unset and neither /tmp/pi/pi nor the npm dist/cli.js layout exists |
| 2 | provider key | No API key found |
placeholder key passed isUsableKey, or provider-hygiene deleted the only key after PI_MODEL forced a lane |
| 3 | model id | cannot resolve model / wrong provider |
bare deepseek-v4-flash not mapped to a provider-qualified id |
| 4 | stdin | pi hangs until the whole deadline | pi waits for EOF on a non-TTY stdin; must be a real pipe that is closed (spawn + child.stdin.end()), not execFile/ignore |
| 5 | TLS env | error setting certificate file |
host-only SSL_CERT_FILE/CURL_CA_BUNDLE leaked into the sandbox |
| 6 | agent store | wrong-provider auth skipped | stale ~/.pi/agent store; must isolate via PI_CODING_AGENT_DIR |
| 7 | output path | solve succeeds, no solution.md |
workspace bind missing/read-only |
| 8 | timeouts | killed mid-answer | pi-agent cap (25 min) vs server cap (30 min) vs a slow thinking model |
The off-by-one packaging of this class refers to the two boundary conditions that are easy to get wrong and that the probe pins down:
position starting at 1 and queue_depth excludes completed rows.25 min < 30 min), or the last solve of a burst is SIGKILLed exactly when it would otherwise finish.Nothing in the live deployment needs changing. The reproducible "fix" is the deterministic bring-up + readiness gate that guarantees a fresh :8766 server completes solves. Run as the service user (kara) on the host:
# 0. One-time prerequisites
install -d -m 0755 ~/off-by-one
# ~/off-by-one/off-by-one <- server binary (mode 0755)
# ~/off-by-one/.env <- provider + tuning env
cat > ~/off-by-one/.env <<'EOF'
OPENROUTER_API_KEY=sk-or-v1-... # real key, >20 chars, no "example"/"changeme"
PI_MODEL=openrouter/deepseek/deepseek-v4.1-flash
EOF
chmod 600 ~/off-by-one/.env
# 1. Make the bridge discoverable and sane (matches the live lab layout)
install -d -m 0755 ~/.local/bin
test -x ~/.local/bin/pi-agent || install -m 0755 pi-agent ~/.local/bin/pi-agent
export PI_BIN=/tmp/pi/pi # ELF, or PI_HOME for the npm layout
test -x "$PI_BIN" || { echo "pi binary missing"; exit 1; }
# 2. (Re)start the service and wait for health
sudo systemctl daemon-reload
sudo systemctl restart off-by-one.service
for i in $(seq 1 30); do
curl -fsS http://<ip-address>:8766/health >/dev/null 2>&1 && break
sleep 1
done
curl -fsS http://<ip-address>:8766/health
Hardening that prevents the failure modes in §3 (already implemented in pi-agent; keep it that way):
delete env.SSL_CERT_FILE; delete env.CURL_CA_BUNDLE;spawn + child.stdin.end() for pi; never execFile.PI_CODING_AGENT_DIR=$(mktemp -d).inner_timeout (25m) < outer_timeout (30m).$ curl -fsS http://<ip-address>:8766/health # {"status":"ok",...}
$ curl -fsS http://<ip-address>:8766/api/v1/stats \
| grep -o '"solver_available":[a-z]*' # "solver_available":true
$ ss -ltn | grep -q ':8766' && echo "listening"
complete)SUBMIT=$(curl -fsS -X POST http://<ip-address>:8766/api/v1/problems/submit \
-H 'Content-Type: application/json' \
-d '{"problem_class":"probe-8766-liveness",
"description":"Return a one-line proof that the solve chain works.",
"environment":"linux","language":"","version":"","cadence":"once"}')
echo "$SUBMIT"
# -> {"problem_class":"...","status":"queued","position":N}
# poll until status is complete (or failed)
while :; do
curl -fsS http://<ip-address>:8766/api/v1/queue \
| tr ',' '\n' | grep -A2 '"problem_class":"probe-8766-liveness"' && break
sleep 10
done
Expected terminal state: "status":"complete","stage":"done", with solution.md, evidence.md, and signatures.json present in the workspace that the server created (/tmp/off-by-one-sandbox-sub_<id>).
The probe you are reading was dispatched exactly as:
off-by-one.service
└─ bwrap … --bind /tmp/off-by-one-sandbox-sub_c048bd /workspace
└─ pi-agent solve --problem-file /workspace/problem.json
--output /workspace/solution.md
--evidence /workspace/evidence.md
--signatures /workspace/signatures.json
--model deepseek-v4-flash
└─ pi (provider=openrouter, model=deepseek/deepseek-v4.1-flash)
In-sandbox checks that prove completion:
test -s /workspace/solution.md && echo "solution.md written"
test -s /workspace/evidence.md && echo "evidence.md written"
python3 -c 'import json;json.load(open("/workspace/signatures.json"));print("signatures.json valid")'
When those three pass and /health returns {"status":"ok"}, the answer to the probe is yes: a fresh server on :8766 completes a solve.
# Evidence - Problem class: df-off-by-one-1-live-probe - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-11T10:22:02.405Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Live probe: does a fresh server on :8766 complete a solve?", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "df-off-by-one-1-live-probe", "provider": "openrouter", "solved_at": "2026-09-11T10:22:02.406Z", "version": ""}