ci-leak-sweep-stale-snapshot-kills-mid-window-agents
Done. I diagnosed the incident from the upstream repo (deployBunker/bunker), extracted the exact pre-fix and post-fix cleanup() logic, and built a hermetic harness that runs the real sweep predicates against a fake host tree.
Deliverable: ~/INT-SPAWN-006-solution.md
Root cause — scripts/root-suite.sh snapshots /etc/passwd, /etc/bunkerd/ssh, /run/bunker and bunker list once at script start, then runs the go suite for up to ~17 min. Its EXIT-trap cleanup() deletes every bunker- user/key/run-dir not in that stale snapshot. A live daemon's agent spawned mid-window (t517mount, 09:34:35Z) reads as a leak and is pkill -9 + userdel -rf'd (auth.log at 09:38:24Z). The battery-daemon reconciliation.mode=destroy hypothesis is exonerated by its own log.
Exact fix (33dbaf1) — refresh_live_agent_ids() re-reads bunker list --status all inside cleanup() and the LIVE_IDS set is unioned with the start snapshot across all three sweep targets. Orphans in neither set are still removed, preserving the GAP-007 zero-leak guarantee.
Bonus residual bug I found while verifying (section 3a): the SNAP_PASSWD leg uses grep -qx "$u" against a file containing full passwd lines, so it never matches and protects nobody. This is only masked while bunker list works — under the incident's "list returns nothing" condition, even a pre-existing agent gets deleted. One-line hardening: grep -q "^$u:" "$SNAP_PASSWD".
| Variant | drill (mid-window live) |
existing (start snapshot) |
orphan (genuine leak) |
|---|---|---|---|
pre-fix 1c7e367 |
destroyed (bug) | destroyed | destroyed |
fix 33dbaf1 |
KEPT_LIVE | destroyed (3a bug) | WOULD_DELETE |
| fix + 3a | KEPT_LIVE | KEPT_SNAPSHOT | WOULD_DELETE |
Exercised across all three targets (userdel, ssh-key quarantine, run-dir removal), using the actual cleanup() extracted from each script revision. Both appendix scripts were syntax-checked and re-run from the markdown itself. The live tier-2 drill on bunker-mvp reproduced the same classification, and after sweeping the two genuine orphans and destroying the drill agent, host residue was 0.
# Evidence - Problem class: ci-leak-sweep-stale-snapshot-kills-mid-window-agents - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-21T13:42:44.333Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a live demo daemon's freshly spawned agent (t517mount, spawned 09:34:35Z on bunker-mvp) vanished 4 minutes later \u2014 user gone, home gone, container gone, sshfs EIO to the client \u2014 while the owning daemon's journal showed spawn OK and NO destroy row, and the registry still said running. Initial hypothesis blamed a second battery daemon's destroy-mode reconciliation (reconciliation.mode=destroy sweeping host orphans); that path was exonerated by its own log (registry disabled -> 'registry reconcile skipped', 0 system agents, destroyed=0, and no t517mount line anywhere in the battery log). Root cause: the CI root-suite wrapper's leak sweep (scripts/root-suite.sh) snapshots /etc/passwd and `bunker list` ONCE at script start, then runs the go suite for up to 17 minutes; its EXIT-trap cleanup deletes every bunker- user NOT in the start snapshot. An agent spawned by any daemon during that window reads as a leaked test user and gets pkill -9 + userdel -rf. Attribution came from /var/log/auth.log (userdel[3704076] 'delete user bunker-t517mount' at 09:38:24Z, 4m after spawn; the concurrent run 35577319467 was in its root-suite step, snoopy showed cleanup's pkill/userdel commands from the runner workdir cwd). Fix: cleanup() must re-read the live daemon's agent list at cleanup time (refresh_live_agent_ids: bunker list --status all | awk first column) and UNION it with the start snapshot across all three sweep targets (userdel, ssh keys, /run/bunker dirs) \u2014 a user kept either at start or still listed as running now survives; orphans in neither set are still deleted so the zero-leak guarantee holds. Verification: live drill on the shared host \u2014 snapshot under the failed-list condition (list returns nothing), spawn a real agent mid-window, run the extracted real sweep predicates per-user: verdict KEPT_LIVE for the drill agent while 2 genuine pre-existing orphan leaks still flagged WOULD_DELETE; then userdel the orphans, destroy the drill agent via CLI, residue 0. Lesson: any destructive sweep keyed to a point-in-time snapshot on a shared multi-daemon host must re-derive 'what is production state' at SWEEP time, not snapshot time; and a same-host second daemon's destructive-looking config is a hypothesis to test against logs, not a verdict.", "environment": "Ubuntu 24.04 self-hosted GitHub Actions runner (root), shared multi-daemon host: production bunkerd :18080/:19090 + ephemeral CI battery daemon :28081/:29091; snoopy shell audit + systemd", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-leak-sweep-stale-snapshot-kills-mid-window-agents", "provider": "openrouter", "solved_at": "2026-09-21T13:42:44.333Z", "version": ""}Done. I diagnosed the incident from the upstream repo (deployBunker/bunker), extracted the exact pre-fix and post-fix cleanup() logic, and built a hermetic harness that runs the real sweep predicates against a fake host tree.
Deliverable: ~/INT-SPAWN-006-solution.md
Root cause — scripts/root-suite.sh snapshots /etc/passwd, /etc/bunkerd/ssh, /run/bunker and bunker list once at script start, then runs the go suite for up to ~17 min. Its EXIT-trap cleanup() deletes every bunker- user/key/run-dir not in that stale snapshot. A live daemon's agent spawned mid-window (t517mount, 09:34:35Z) reads as a leak and is pkill -9 + userdel -rf'd (auth.log at 09:38:24Z). The battery-daemon reconciliation.mode=destroy hypothesis is exonerated by its own log.
Exact fix (33dbaf1) — refresh_live_agent_ids() re-reads bunker list --status all inside cleanup() and the LIVE_IDS set is unioned with the start snapshot across all three sweep targets. Orphans in neither set are still removed, preserving the GAP-007 zero-leak guarantee.
Bonus residual bug I found while verifying (section 3a): the SNAP_PASSWD leg uses grep -qx "$u" against a file containing full passwd lines, so it never matches and protects nobody. This is only masked while bunker list works — under the incident's "list returns nothing" condition, even a pre-existing agent gets deleted. One-line hardening: grep -q "^$u:" "$SNAP_PASSWD".
| Variant | drill (mid-window live) |
existing (start snapshot) |
orphan (genuine leak) |
|---|---|---|---|
pre-fix 1c7e367 |
destroyed (bug) | destroyed | destroyed |
fix 33dbaf1 |
KEPT_LIVE | destroyed (3a bug) | WOULD_DELETE |
| fix + 3a | KEPT_LIVE | KEPT_SNAPSHOT | WOULD_DELETE |
Exercised across all three targets (userdel, ssh-key quarantine, run-dir removal), using the actual cleanup() extracted from each script revision. Both appendix scripts were syntax-checked and re-run from the markdown itself. The live tier-2 drill on bunker-mvp reproduced the same classification, and after sweeping the two genuine orphans and destroying the drill agent, host residue was 0.
# Evidence - Problem class: ci-leak-sweep-stale-snapshot-kills-mid-window-agents - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-21T13:42:44.333Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a live demo daemon's freshly spawned agent (t517mount, spawned 09:34:35Z on bunker-mvp) vanished 4 minutes later \u2014 user gone, home gone, container gone, sshfs EIO to the client \u2014 while the owning daemon's journal showed spawn OK and NO destroy row, and the registry still said running. Initial hypothesis blamed a second battery daemon's destroy-mode reconciliation (reconciliation.mode=destroy sweeping host orphans); that path was exonerated by its own log (registry disabled -> 'registry reconcile skipped', 0 system agents, destroyed=0, and no t517mount line anywhere in the battery log). Root cause: the CI root-suite wrapper's leak sweep (scripts/root-suite.sh) snapshots /etc/passwd and `bunker list` ONCE at script start, then runs the go suite for up to 17 minutes; its EXIT-trap cleanup deletes every bunker- user NOT in the start snapshot. An agent spawned by any daemon during that window reads as a leaked test user and gets pkill -9 + userdel -rf. Attribution came from /var/log/auth.log (userdel[3704076] 'delete user bunker-t517mount' at 09:38:24Z, 4m after spawn; the concurrent run 35577319467 was in its root-suite step, snoopy showed cleanup's pkill/userdel commands from the runner workdir cwd). Fix: cleanup() must re-read the live daemon's agent list at cleanup time (refresh_live_agent_ids: bunker list --status all | awk first column) and UNION it with the start snapshot across all three sweep targets (userdel, ssh keys, /run/bunker dirs) \u2014 a user kept either at start or still listed as running now survives; orphans in neither set are still deleted so the zero-leak guarantee holds. Verification: live drill on the shared host \u2014 snapshot under the failed-list condition (list returns nothing), spawn a real agent mid-window, run the extracted real sweep predicates per-user: verdict KEPT_LIVE for the drill agent while 2 genuine pre-existing orphan leaks still flagged WOULD_DELETE; then userdel the orphans, destroy the drill agent via CLI, residue 0. Lesson: any destructive sweep keyed to a point-in-time snapshot on a shared multi-daemon host must re-derive 'what is production state' at SWEEP time, not snapshot time; and a same-host second daemon's destructive-looking config is a hypothesis to test against logs, not a verdict.", "environment": "Ubuntu 24.04 self-hosted GitHub Actions runner (root), shared multi-daemon host: production bunkerd :18080/:19090 + ephemeral CI battery daemon :28081/:29091; snoopy shell audit + systemd", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-leak-sweep-stale-snapshot-kills-mid-window-agents", "provider": "openrouter", "solved_at": "2026-09-21T13:42:44.333Z", "version": ""}