Problem class: bunkerd-pre-deploy-orphan-triage
Problem class: bunkerd-pre-deploy-orphan-triage
Scope: any host about to run a bunkerd build that includes the GAP‑070 on-disk agent registry (reconcile defaults to destroying unmanaged bunker-* users).
Goal: triage every unmanaged bunker-* user as LIVE or corpse using read-only checks, preserve live orphans (agent.reconciliation.mode=adopt), reap only corpses, pre-create the registry file before restart, and prove the result.
The GAP‑070 registry makes the daemon durable: /var/lib/bunkerd/agents.jsonl becomes the source of truth for which agents exist. On startup the daemon reconciles the host against that file. Any bunker-* user the registry does not know about ("unmanaged") is destroyed by default. Three things make that dangerous:
/var/lib/bunkerd/agents.jsonl, so the file must exist with mode 0600 before restart.resource.Tracker (in-memory AgentRecords) only, never /home/<user>/.bunker/ports. An orphan holding 30000–30099 is invisible, so a fresh spawn gets the same range and collides.Fix: triage first (read-only), set reconciliation.mode=adopt if any live orphan exists, reap only proven corpses manually, pre-create the registry 0600, update the port ledger, then restart and verify.
GAP‑070 adds /var/lib/bunkerd/agents.jsonl as a persisted JSONL of AgentRecords and a startup reconcile pass. The host scan keys on the bunker- username prefix and the bunker-docker-<id>.service unit prefix. Registry membership decides ownership:
| Host state | Registry state | Result |
|---|---|---|
bunker-<id> exists |
record present | managed, untouched |
bunker-<id> exists |
no record | unmanaged → default destroy |
| user deleted | record present | stale record, pruned |
agent.reconciliation.mode selects the unmanaged-user policy:
- destroy (default in GAP‑070) — delete user + runtime state.
- adopt — import the existing user into the registry and keep it.
Verify the exact accepted values of your build with
bunkerd --help/ the release notes before rollout; any value other thanadoptshould be treated as destructive for live orphans until proven otherwise.
The daemon opens the registry at boot and aborts if the open fails. It does not reliably create it for you. Restart without the file ⇒ crash-loop, which on a production host looks like a failed deploy and tempts operators to "just destroy everything".
Grounded in the installed binary (go version -m /usr/local/bin/bunkerd shows github.com/deployBunker/bunker), the allocation path is:
internal/resource.(*PortAllocator).Allocate/.Free
internal/resource.(*Tracker).Register/.Unregister/.Get/.List
internal/resource.AgentRecord
PortAllocator is reserved from Tracker, and Tracker is populated only by Register calls made by this daemon process. The disk artifact written for the user (/home/<user>/.bunker/ports, written by write port range file) is not read back on startup in 0.1.3. Therefore:
kara-lair holds 30000–30099;30000–30099;GAP‑070's persisted registry is exactly the missing seed: restart can rebuild the tracker from agents.jsonl and reserve the orphan's range. Until the registry is populated with the orphan (adopt), the collision persists.
On the QA-battery host:
- bunker list --server <host> was non-empty (the daemon's live set);
- journalctl -u bunkerd showed recent SpawnAgent POSTs;
- the accounts were minutes old.
bunker list is authoritative for agents this instance tracks, but an empty/non-matching result is not proof of death. Triage must combine multiple independent liveness signals (next section).
bunker CLI available on the host; BUNKER_SERVER or --server points at the host's bunkerd endpoint./etc/bunkerd/config.yaml/var/lib/bunkerd/agents.jsonl/run/bunker/<id>/~<id>/.bunker/portsbunker-docker-<id>.service<id> = the agent id (usually a UUID) = username with the bunker- prefix removed.set -euo pipefail
STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
WORK="/var/tmp/bunker-triage-$STAMP"; mkdir -p "$WORK"
CONFIG=/etc/bunkerd/config.yaml
REGISTRY=/var/lib/bunkerd/agents.jsonl
SERVER_HOST="$(hostname -f)"
cp -a "$CONFIG" "$WORK/config.yaml.pre"
getent passwd | awk -F: '$1 ~ /^bunker-/ {print}' > "$WORK/users.pre.txt"
bunker list --server "$SERVER_HOST" > "$WORK/agents.list.pre.txt" 2>&1 || true
systemctl list-units --all --no-legend 'bunker-docker-*' > "$WORK/units.pre.txt" 2>&1 || true
echo "snapshot in $WORK"
bunker-* users# Every host user the daemon does NOT already manage.
bunker list --server "$SERVER_HOST" | awk '{for(i=1;i<=NF;i++) if($i ~ /^[0-9a-f-]{8,}$/) print $i}' \
| sort -u > "$WORK/managed.ids"
getent passwd | awk -F: '$1 ~ /^bunker-/ {sub(/^bunker-/,"",$1); print $1}' \
| sort -u > "$WORK/host.ids"
comm -23 "$WORK/host.ids" "$WORK/managed.ids" > "$WORK/unmanaged.ids"
echo "Unmanaged bunker-* agents:"; cat "$WORK/unmanaged.ids"
Each check is safe to run. A user is LIVE if any strong signal fires; it is a corpse only if all strong signals are absent, and after a short re-confirmation (to avoid reaping a spawn that is still in flight).
#!/usr/bin/env bash
# triage-orphans.sh -- read-only LIVE/corpse classification
set -u
SERVER_HOST="$(hostname -f)"
RUN_DIR=/run/bunker
WINDOW="120 min ago"
OUT="${1:-/var/tmp/bunker-triage.tsv}"
printf 'id\tuser\tunit_active\tsock_ps\tprocs\tports_file\tjournal_spawn\tin_list\tverdict\n' > "$OUT"
while read -r id; do
[ -n "$id" ] || continue
user="bunker-$id"
unit="bunker-docker-${id}"
sock="unix://${RUN_DIR}/${id}/docker.sock"
unit_active=no; systemctl is-active --quiet "$unit" && unit_active=yes
sock_ps="n/a"
if [ -S "${RUN_DIR}/${id}/docker.sock" ]; then
if out="$(docker -H "$sock" ps -a --format '{{.ID}} {{.Status}}' 2>/dev/null)"; then
[ -n "$out" ] && sock_ps="containers" || sock_ps="empty"
else
sock_ps="unreachable"
fi
else
sock_ps="nosock"
fi
procs=no; pgrep -u "$user" >/dev/null 2>&1 && procs=yes
ports_file="absent"
[ -f "/home/${user}/.bunker/ports" ] && ports_file="$(tr '\n' ',' < "/home/${user}/.bunker/ports")"
journal_spawn=no
journalctl -u bunkerd --since "$WINDOW" --no-pager 2>/dev/null | grep -q "$id" && journal_spawn=yes
in_list=no
bunker list --server "$SERVER_HOST" 2>/dev/null | grep -q "$id" && in_list=yes
# Any strong signal => LIVE. port file alone is NOT proof of life.
verdict=CORPSE
if [ "$unit_active" = yes ] || [ "$sock_ps" = containers ] || \
[ "$procs" = yes ] || [ "$journal_spawn" = yes ] || [ "$in_list" = yes ]; then
verdict=LIVE
fi
printf '%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\n' \
"$id" "$user" "$unit_active" "$sock_ps" "$procs" "$ports_file" "$journal_spawn" "$in_list" "$verdict" | tee -a "$OUT"
done < "$2"
chmod +x triage-orphans.sh
./triage-orphans.sh "$WORK/triage.tsv" "$WORK/unmanaged.ids"
Re-confirm before reaping:
# Corpse candidates: re-probe 60s later; anything that woke up flips to LIVE.
sleep 60
./triage-orphans.sh "$WORK/triage.confirm.tsv" "$WORK/unmanaged.ids"
awk -F'\t' 'NR>1 && $9=="CORPSE"{print $1}' "$WORK/triage.confirm.tsv" | sort -u > "$WORK/corpses.ids"
awk -F'\t' 'NR>1 && $9=="LIVE" {print $1}' "$WORK/triage.confirm.tsv" | sort -u > "$WORK/live.ids"
echo "LIVE:"; cat "$WORK/live.ids"; echo "CORPSES:"; cat "$WORK/corpses.ids"
On an active battery host, also treat a fresh
created_at/spawn timestamp as LIVE even if the process check is momentarily empty — a spawn may be mid-provision. The 60 s re-probe covers this.
if [ -s "$WORK/live.ids" ]; then
echo ">>> LIVE orphan(s) present -> reconciliation.mode MUST be 'adopt'."
else
echo ">>> No live orphans; corpses may be reaped. Adopt is still safe."
fi
Rule: never restart with the default destroy while live.ids is non-empty and the owner has not signed off on destruction.
while read -r id; do
[ -n "$id" ] || continue
user="bunker-$id"; unit="bunker-docker-${id}"
echo "reaping corpse $id ($user)"
systemctl stop "$unit" 2>/dev/null || true
systemctl disable "$unit" 2>/dev/null || true
loginctl terminate-user "$user" 2>/dev/null || true
pkill -9 -u "$user" 2>/dev/null || true
sleep 1
userdel -r "$user" 2>/dev/null || true
rm -rf "/run/bunker/$id"
done < "$WORK/corpses.ids"
Do not touch ids in $WORK/live.ids.
install -d -m 0755 /var/lib/bunkerd
# Create empty registry; owner must match the daemon's runtime user.
install -m 0600 /dev/null /var/lib/bunkerd/agents.jsonl
# If bunkerd.service runs as a service user (e.g. User=bunkerd), chown it:
# chown bunkerd:bunkerd /var/lib/bunkerd/agents.jsonl
stat -c '%a %U:%G %n' /var/lib/bunkerd/agents.jsonl # expect: 600 <owner> /var/lib/bunkerd/agents.jsonl
agent.reconciliation.mode=adopt (required if any LIVE orphan)Preferred (yq):
cp -a "$CONFIG" "$CONFIG.bak.$(date -u +%s)"
yq -i '.agent.reconciliation.mode = "adopt"' "$CONFIG"
Fallback (no yq) — inserts the nested keys once, idempotently:
if ! grep -qE '^[[:space:]]+reconciliation:' "$CONFIG"; then
awk '
BEGIN{done=0}
/^agent:[[:space:]]*$/ && !done {
print; print " reconciliation:"; print " mode: adopt"; done=1; next
}
{print}
END{ if(!done){ print ""; print "agent:"; print " reconciliation:"; print " mode: adopt" } }
' "$CONFIG" > "$CONFIG.tmp" && mv "$CONFIG.tmp" "$CONFIG"
fi
The allocator will only see live orphans after adopt imports them. Before that, compute the union of on-disk claims and refuse to deploy if any overlap the active range.
echo "== On-disk port claims =="
for f in ~*/.bunker/ports; do
[ -f "$f" ] || continue
printf '%s\t%s\n' "$(basename "$(dirname "$(dirname "$f")")")" "$(tr '\n' ',' < "$f")"
done | sort
echo "== Pairwise overlap check =="
for f in ~*/.bunker/ports; do
[ -f "$f" ] || continue
cat "$f"
done | grep -oE '[0-9]+-[0-9]+' | sort -t- -k1,1n | \
awk -F- 'NR>1 && $1<=prev_end {print "OVERLAP: "prev" and "$0} {prev=$0; prev_end=$2}'
If an overlap exists involving a live orphan:
- it is preserved by adopt;
- after restart, confirm the registry seeded the tracker and the fresh agent received a different range (Step 8);
- if the build still double-allocates after adopt, temporarily pin the orphan's range out of the fresh pool (or move the fresh spawn) and file the allocator bug — the durable fix is registry-seeding of PortAllocator.
systemctl daemon-reload
systemctl restart bunkerd
systemctl is-active bunkerd
journalctl -u bunkerd --since '2 min ago' --no-pager | tail -40
Run every check; all must pass.
echo "== daemon =="
systemctl is-active bunkerd # active
echo "== registry =="
stat -c '%a %U:%G %s %n' /var/lib/bunkerd/agents.jsonl # 600, non-zero after adopt
grep -c . /var/lib/bunkerd/agents.jsonl 2>/dev/null || true
echo "== adopted orphans still present =="
bunker list --server "$SERVER_HOST"
while read -r id; do
[ -n "$id" ] || continue
getent passwd "bunker-$id" >/dev/null && echo "PRESERVED bunker-$id" || echo "MISSING bunker-$id"
grep -q "$id" /var/lib/bunkerd/agents.jsonl && echo "IN REGISTRY $id" || echo "NOT IN REGISTRY $id"
done < "$WORK/live.ids"
echo "== corpses gone =="
while read -r id; do
[ -n "$id" ] || continue
getent passwd "bunker-$id" >/dev/null && echo "STILL PRESENT $id (FAIL)" || echo "reaped $id"
done < "$WORK/corpses.ids"
echo "== no port double allocation =="
for f in ~*/.bunker/ports; do
[ -f "$f" ] || continue
cat "$f"
done | grep -oE '[0-9]+-[0-9]+' | sort -t- -k1,1n | \
awk -F- 'NR>1 && $1<=prev_end {print "OVERLAP: "prev" and "$0; bad=1} {prev=$0; prev_end=$2} END{exit bad?1:0}' \
&& echo "no overlaps"
echo "== reconcile log =="
journalctl -u bunkerd --since '10 min ago' --no-pager | grep -iE 'reconcil|adopt|registry|unmanaged' || true
Expected:
- bunkerd active, registry 0600.
- Every id from live.ids still has its user and a registry record.
- Every id from corpses.ids is gone.
- No overlapping port ranges.
- Reconcile log shows adopt, not destroy.
systemctl stop bunkerd
cp -a "$WORK/config.yaml.pre" "$CONFIG"
cp -a /var/lib/bunkerd/agents.jsonl "$WORK/agents.jsonl.rollback" 2>/dev/null || true
rm -f /var/lib/bunkerd/agents.jsonl # if a clean empty state is required
# recreate preserved users from backups / re-run provisioning as needed
systemctl start bunkerd
Keep $WORK (users, units, registry, triage TSV) until the owner confirms the adoption list.
The sandbox confirms the mechanism and the missing pieces:
go version -m /usr/local/bin/bunkerd
# path github.com/deployBunker/bunker/cmd/bunkerd
# mod github.com/deployBunker/bunker v0.1.4-0.20260825202354-33c9e0083a9e+dirty
strings -n6 /usr/local/bin/bunkerd | grep -oE 'internal/resource\.[A-Za-z0-9_()*.]+' | sort -u
# internal/resource.(*PortAllocator).Allocate/.Free
# internal/resource.(*Tracker).Register/.Unregister/.Get/.List
# internal/resource.AgentRecord
strings -n6 /usr/local/bin/bunkerd | grep -oE 'mapstructure:"[^"]*"' | sort -u
# agent.max_agents, agent.port_range_start/end/per_agent, agent.ssh_dir, ...
# NOTE: no agent.reconciliation key in the currently installed build -> the
# GAP-070 registry/adopt schema is the piece being rolled out.
And a live-start test with a scratch config proved the 0.1.3/0.1.4 daemon starts from an empty state with an in-memory tracker and writes no on-disk registry:
/usr/local/bin/bunkerd -c /tmp/bt/config.yaml # starts; /tmp/bt/data contains no agents.jsonl
That is precisely why an orphan holding 30000–30099 is invisible to PortAllocator, and why the GAP‑070 registry plus a pre-deploy LIVE/corpse triage is mandatory before enabling the default-destroy reconcile.
# Evidence - Problem class: bunkerd-pre-deploy-orphan-triage - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-13T03:25:19.988Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "TR-041 tick 39 task-router: before deploying bunkerd with GAP-070 registry (reconcile default=DESTROYS unmanaged bunker-* users), triage every unmanaged bunker-* user as LIVE-or-corpse first. Read-only checks: systemctl status bunker-docker-<id>, docker -H unix:///run/bunker/<id>/docker.sock ps -a, cat /home/<user>/.bunker/ports, journalctl -u bunkerd spawn lines, and bunker list --server <host>. On an active QA-battery host, unknown users turned out to be LIVE battery agents spawned minutes earlier (bunker list was non-empty; daemon journal showed SpawnAgent POSTs). Reap only corpses; set agent.reconciliation.mode=adopt in /etc/bunkerd/config.yaml when any live orphan must be preserved (owner decision pending); pre-create /var/lib/bunkerd/agents.jsonl (0600) BEFORE restart \u2014 daemon refuses to start if it cannot open the registry file. Also found 0.1.3 double-allocated port range 30000-30099 to a fresh spawn while orphaned kara-lair still held it \u2014 allocator ignores users the daemon memory does not track.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bunkerd-pre-deploy-orphan-triage", "provider": "openrouter", "solved_at": "2026-09-13T03:25:19.988Z", "version": ""}Problem class: bunkerd-pre-deploy-orphan-triage
Scope: any host about to run a bunkerd build that includes the GAP‑070 on-disk agent registry (reconcile defaults to destroying unmanaged bunker-* users).
Goal: triage every unmanaged bunker-* user as LIVE or corpse using read-only checks, preserve live orphans (agent.reconciliation.mode=adopt), reap only corpses, pre-create the registry file before restart, and prove the result.
The GAP‑070 registry makes the daemon durable: /var/lib/bunkerd/agents.jsonl becomes the source of truth for which agents exist. On startup the daemon reconciles the host against that file. Any bunker-* user the registry does not know about ("unmanaged") is destroyed by default. Three things make that dangerous:
/var/lib/bunkerd/agents.jsonl, so the file must exist with mode 0600 before restart.resource.Tracker (in-memory AgentRecords) only, never /home/<user>/.bunker/ports. An orphan holding 30000–30099 is invisible, so a fresh spawn gets the same range and collides.Fix: triage first (read-only), set reconciliation.mode=adopt if any live orphan exists, reap only proven corpses manually, pre-create the registry 0600, update the port ledger, then restart and verify.
GAP‑070 adds /var/lib/bunkerd/agents.jsonl as a persisted JSONL of AgentRecords and a startup reconcile pass. The host scan keys on the bunker- username prefix and the bunker-docker-<id>.service unit prefix. Registry membership decides ownership:
| Host state | Registry state | Result |
|---|---|---|
bunker-<id> exists |
record present | managed, untouched |
bunker-<id> exists |
no record | unmanaged → default destroy |
| user deleted | record present | stale record, pruned |
agent.reconciliation.mode selects the unmanaged-user policy:
- destroy (default in GAP‑070) — delete user + runtime state.
- adopt — import the existing user into the registry and keep it.
Verify the exact accepted values of your build with
bunkerd --help/ the release notes before rollout; any value other thanadoptshould be treated as destructive for live orphans until proven otherwise.
The daemon opens the registry at boot and aborts if the open fails. It does not reliably create it for you. Restart without the file ⇒ crash-loop, which on a production host looks like a failed deploy and tempts operators to "just destroy everything".
Grounded in the installed binary (go version -m /usr/local/bin/bunkerd shows github.com/deployBunker/bunker), the allocation path is:
internal/resource.(*PortAllocator).Allocate/.Free
internal/resource.(*Tracker).Register/.Unregister/.Get/.List
internal/resource.AgentRecord
PortAllocator is reserved from Tracker, and Tracker is populated only by Register calls made by this daemon process. The disk artifact written for the user (/home/<user>/.bunker/ports, written by write port range file) is not read back on startup in 0.1.3. Therefore:
kara-lair holds 30000–30099;30000–30099;GAP‑070's persisted registry is exactly the missing seed: restart can rebuild the tracker from agents.jsonl and reserve the orphan's range. Until the registry is populated with the orphan (adopt), the collision persists.
On the QA-battery host:
- bunker list --server <host> was non-empty (the daemon's live set);
- journalctl -u bunkerd showed recent SpawnAgent POSTs;
- the accounts were minutes old.
bunker list is authoritative for agents this instance tracks, but an empty/non-matching result is not proof of death. Triage must combine multiple independent liveness signals (next section).
bunker CLI available on the host; BUNKER_SERVER or --server points at the host's bunkerd endpoint./etc/bunkerd/config.yaml/var/lib/bunkerd/agents.jsonl/run/bunker/<id>/~<id>/.bunker/portsbunker-docker-<id>.service<id> = the agent id (usually a UUID) = username with the bunker- prefix removed.set -euo pipefail
STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
WORK="/var/tmp/bunker-triage-$STAMP"; mkdir -p "$WORK"
CONFIG=/etc/bunkerd/config.yaml
REGISTRY=/var/lib/bunkerd/agents.jsonl
SERVER_HOST="$(hostname -f)"
cp -a "$CONFIG" "$WORK/config.yaml.pre"
getent passwd | awk -F: '$1 ~ /^bunker-/ {print}' > "$WORK/users.pre.txt"
bunker list --server "$SERVER_HOST" > "$WORK/agents.list.pre.txt" 2>&1 || true
systemctl list-units --all --no-legend 'bunker-docker-*' > "$WORK/units.pre.txt" 2>&1 || true
echo "snapshot in $WORK"
bunker-* users# Every host user the daemon does NOT already manage.
bunker list --server "$SERVER_HOST" | awk '{for(i=1;i<=NF;i++) if($i ~ /^[0-9a-f-]{8,}$/) print $i}' \
| sort -u > "$WORK/managed.ids"
getent passwd | awk -F: '$1 ~ /^bunker-/ {sub(/^bunker-/,"",$1); print $1}' \
| sort -u > "$WORK/host.ids"
comm -23 "$WORK/host.ids" "$WORK/managed.ids" > "$WORK/unmanaged.ids"
echo "Unmanaged bunker-* agents:"; cat "$WORK/unmanaged.ids"
Each check is safe to run. A user is LIVE if any strong signal fires; it is a corpse only if all strong signals are absent, and after a short re-confirmation (to avoid reaping a spawn that is still in flight).
#!/usr/bin/env bash
# triage-orphans.sh -- read-only LIVE/corpse classification
set -u
SERVER_HOST="$(hostname -f)"
RUN_DIR=/run/bunker
WINDOW="120 min ago"
OUT="${1:-/var/tmp/bunker-triage.tsv}"
printf 'id\tuser\tunit_active\tsock_ps\tprocs\tports_file\tjournal_spawn\tin_list\tverdict\n' > "$OUT"
while read -r id; do
[ -n "$id" ] || continue
user="bunker-$id"
unit="bunker-docker-${id}"
sock="unix://${RUN_DIR}/${id}/docker.sock"
unit_active=no; systemctl is-active --quiet "$unit" && unit_active=yes
sock_ps="n/a"
if [ -S "${RUN_DIR}/${id}/docker.sock" ]; then
if out="$(docker -H "$sock" ps -a --format '{{.ID}} {{.Status}}' 2>/dev/null)"; then
[ -n "$out" ] && sock_ps="containers" || sock_ps="empty"
else
sock_ps="unreachable"
fi
else
sock_ps="nosock"
fi
procs=no; pgrep -u "$user" >/dev/null 2>&1 && procs=yes
ports_file="absent"
[ -f "/home/${user}/.bunker/ports" ] && ports_file="$(tr '\n' ',' < "/home/${user}/.bunker/ports")"
journal_spawn=no
journalctl -u bunkerd --since "$WINDOW" --no-pager 2>/dev/null | grep -q "$id" && journal_spawn=yes
in_list=no
bunker list --server "$SERVER_HOST" 2>/dev/null | grep -q "$id" && in_list=yes
# Any strong signal => LIVE. port file alone is NOT proof of life.
verdict=CORPSE
if [ "$unit_active" = yes ] || [ "$sock_ps" = containers ] || \
[ "$procs" = yes ] || [ "$journal_spawn" = yes ] || [ "$in_list" = yes ]; then
verdict=LIVE
fi
printf '%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\n' \
"$id" "$user" "$unit_active" "$sock_ps" "$procs" "$ports_file" "$journal_spawn" "$in_list" "$verdict" | tee -a "$OUT"
done < "$2"
chmod +x triage-orphans.sh
./triage-orphans.sh "$WORK/triage.tsv" "$WORK/unmanaged.ids"
Re-confirm before reaping:
# Corpse candidates: re-probe 60s later; anything that woke up flips to LIVE.
sleep 60
./triage-orphans.sh "$WORK/triage.confirm.tsv" "$WORK/unmanaged.ids"
awk -F'\t' 'NR>1 && $9=="CORPSE"{print $1}' "$WORK/triage.confirm.tsv" | sort -u > "$WORK/corpses.ids"
awk -F'\t' 'NR>1 && $9=="LIVE" {print $1}' "$WORK/triage.confirm.tsv" | sort -u > "$WORK/live.ids"
echo "LIVE:"; cat "$WORK/live.ids"; echo "CORPSES:"; cat "$WORK/corpses.ids"
On an active battery host, also treat a fresh
created_at/spawn timestamp as LIVE even if the process check is momentarily empty — a spawn may be mid-provision. The 60 s re-probe covers this.
if [ -s "$WORK/live.ids" ]; then
echo ">>> LIVE orphan(s) present -> reconciliation.mode MUST be 'adopt'."
else
echo ">>> No live orphans; corpses may be reaped. Adopt is still safe."
fi
Rule: never restart with the default destroy while live.ids is non-empty and the owner has not signed off on destruction.
while read -r id; do
[ -n "$id" ] || continue
user="bunker-$id"; unit="bunker-docker-${id}"
echo "reaping corpse $id ($user)"
systemctl stop "$unit" 2>/dev/null || true
systemctl disable "$unit" 2>/dev/null || true
loginctl terminate-user "$user" 2>/dev/null || true
pkill -9 -u "$user" 2>/dev/null || true
sleep 1
userdel -r "$user" 2>/dev/null || true
rm -rf "/run/bunker/$id"
done < "$WORK/corpses.ids"
Do not touch ids in $WORK/live.ids.
install -d -m 0755 /var/lib/bunkerd
# Create empty registry; owner must match the daemon's runtime user.
install -m 0600 /dev/null /var/lib/bunkerd/agents.jsonl
# If bunkerd.service runs as a service user (e.g. User=bunkerd), chown it:
# chown bunkerd:bunkerd /var/lib/bunkerd/agents.jsonl
stat -c '%a %U:%G %n' /var/lib/bunkerd/agents.jsonl # expect: 600 <owner> /var/lib/bunkerd/agents.jsonl
agent.reconciliation.mode=adopt (required if any LIVE orphan)Preferred (yq):
cp -a "$CONFIG" "$CONFIG.bak.$(date -u +%s)"
yq -i '.agent.reconciliation.mode = "adopt"' "$CONFIG"
Fallback (no yq) — inserts the nested keys once, idempotently:
if ! grep -qE '^[[:space:]]+reconciliation:' "$CONFIG"; then
awk '
BEGIN{done=0}
/^agent:[[:space:]]*$/ && !done {
print; print " reconciliation:"; print " mode: adopt"; done=1; next
}
{print}
END{ if(!done){ print ""; print "agent:"; print " reconciliation:"; print " mode: adopt" } }
' "$CONFIG" > "$CONFIG.tmp" && mv "$CONFIG.tmp" "$CONFIG"
fi
The allocator will only see live orphans after adopt imports them. Before that, compute the union of on-disk claims and refuse to deploy if any overlap the active range.
echo "== On-disk port claims =="
for f in ~*/.bunker/ports; do
[ -f "$f" ] || continue
printf '%s\t%s\n' "$(basename "$(dirname "$(dirname "$f")")")" "$(tr '\n' ',' < "$f")"
done | sort
echo "== Pairwise overlap check =="
for f in ~*/.bunker/ports; do
[ -f "$f" ] || continue
cat "$f"
done | grep -oE '[0-9]+-[0-9]+' | sort -t- -k1,1n | \
awk -F- 'NR>1 && $1<=prev_end {print "OVERLAP: "prev" and "$0} {prev=$0; prev_end=$2}'
If an overlap exists involving a live orphan:
- it is preserved by adopt;
- after restart, confirm the registry seeded the tracker and the fresh agent received a different range (Step 8);
- if the build still double-allocates after adopt, temporarily pin the orphan's range out of the fresh pool (or move the fresh spawn) and file the allocator bug — the durable fix is registry-seeding of PortAllocator.
systemctl daemon-reload
systemctl restart bunkerd
systemctl is-active bunkerd
journalctl -u bunkerd --since '2 min ago' --no-pager | tail -40
Run every check; all must pass.
echo "== daemon =="
systemctl is-active bunkerd # active
echo "== registry =="
stat -c '%a %U:%G %s %n' /var/lib/bunkerd/agents.jsonl # 600, non-zero after adopt
grep -c . /var/lib/bunkerd/agents.jsonl 2>/dev/null || true
echo "== adopted orphans still present =="
bunker list --server "$SERVER_HOST"
while read -r id; do
[ -n "$id" ] || continue
getent passwd "bunker-$id" >/dev/null && echo "PRESERVED bunker-$id" || echo "MISSING bunker-$id"
grep -q "$id" /var/lib/bunkerd/agents.jsonl && echo "IN REGISTRY $id" || echo "NOT IN REGISTRY $id"
done < "$WORK/live.ids"
echo "== corpses gone =="
while read -r id; do
[ -n "$id" ] || continue
getent passwd "bunker-$id" >/dev/null && echo "STILL PRESENT $id (FAIL)" || echo "reaped $id"
done < "$WORK/corpses.ids"
echo "== no port double allocation =="
for f in ~*/.bunker/ports; do
[ -f "$f" ] || continue
cat "$f"
done | grep -oE '[0-9]+-[0-9]+' | sort -t- -k1,1n | \
awk -F- 'NR>1 && $1<=prev_end {print "OVERLAP: "prev" and "$0; bad=1} {prev=$0; prev_end=$2} END{exit bad?1:0}' \
&& echo "no overlaps"
echo "== reconcile log =="
journalctl -u bunkerd --since '10 min ago' --no-pager | grep -iE 'reconcil|adopt|registry|unmanaged' || true
Expected:
- bunkerd active, registry 0600.
- Every id from live.ids still has its user and a registry record.
- Every id from corpses.ids is gone.
- No overlapping port ranges.
- Reconcile log shows adopt, not destroy.
systemctl stop bunkerd
cp -a "$WORK/config.yaml.pre" "$CONFIG"
cp -a /var/lib/bunkerd/agents.jsonl "$WORK/agents.jsonl.rollback" 2>/dev/null || true
rm -f /var/lib/bunkerd/agents.jsonl # if a clean empty state is required
# recreate preserved users from backups / re-run provisioning as needed
systemctl start bunkerd
Keep $WORK (users, units, registry, triage TSV) until the owner confirms the adoption list.
The sandbox confirms the mechanism and the missing pieces:
go version -m /usr/local/bin/bunkerd
# path github.com/deployBunker/bunker/cmd/bunkerd
# mod github.com/deployBunker/bunker v0.1.4-0.20260825202354-33c9e0083a9e+dirty
strings -n6 /usr/local/bin/bunkerd | grep -oE 'internal/resource\.[A-Za-z0-9_()*.]+' | sort -u
# internal/resource.(*PortAllocator).Allocate/.Free
# internal/resource.(*Tracker).Register/.Unregister/.Get/.List
# internal/resource.AgentRecord
strings -n6 /usr/local/bin/bunkerd | grep -oE 'mapstructure:"[^"]*"' | sort -u
# agent.max_agents, agent.port_range_start/end/per_agent, agent.ssh_dir, ...
# NOTE: no agent.reconciliation key in the currently installed build -> the
# GAP-070 registry/adopt schema is the piece being rolled out.
And a live-start test with a scratch config proved the 0.1.3/0.1.4 daemon starts from an empty state with an in-memory tracker and writes no on-disk registry:
/usr/local/bin/bunkerd -c /tmp/bt/config.yaml # starts; /tmp/bt/data contains no agents.jsonl
That is precisely why an orphan holding 30000–30099 is invisible to PortAllocator, and why the GAP‑070 registry plus a pre-deploy LIVE/corpse triage is mandatory before enabling the default-destroy reconcile.
# Evidence - Problem class: bunkerd-pre-deploy-orphan-triage - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-13T03:25:19.988Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "TR-041 tick 39 task-router: before deploying bunkerd with GAP-070 registry (reconcile default=DESTROYS unmanaged bunker-* users), triage every unmanaged bunker-* user as LIVE-or-corpse first. Read-only checks: systemctl status bunker-docker-<id>, docker -H unix:///run/bunker/<id>/docker.sock ps -a, cat /home/<user>/.bunker/ports, journalctl -u bunkerd spawn lines, and bunker list --server <host>. On an active QA-battery host, unknown users turned out to be LIVE battery agents spawned minutes earlier (bunker list was non-empty; daemon journal showed SpawnAgent POSTs). Reap only corpses; set agent.reconciliation.mode=adopt in /etc/bunkerd/config.yaml when any live orphan must be preserved (owner decision pending); pre-create /var/lib/bunkerd/agents.jsonl (0600) BEFORE restart \u2014 daemon refuses to start if it cannot open the registry file. Also found 0.1.3 double-allocated port range 30000-30099 to a fresh spawn while orphaned kara-lair still held it \u2014 allocator ignores users the daemon memory does not track.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bunkerd-pre-deploy-orphan-triage", "provider": "openrouter", "solved_at": "2026-09-13T03:25:19.988Z", "version": ""}