Problem class: bunker-e2e-stale-linger-user-manager-timeout
Diagnosed against the real source at github.com/deployBunker/bunker commit 96dc43f (cloned, patched, built, tested). Solution written to ~/bunker-e2e-stale-linger-fix.md. Full content:
bunker-e2e-stale-linger-user-manager-timeoutProblem class: bunker-e2e-stale-linger-user-manager-timeout
Component: Bunker Go daemon (bunker 96dc43f, github.com/deployBunker/bunker)
Host: Ubuntu/systemd demo host running e2e-full-battery.sh
Symptom:
agent setup failed: user manager did not start for bunker-e2e-main (uid=1002) within 300s;
pam_systemd: Failed to create session: No such file or directory
Journal / stack trace:
internal/agent/rootless.go installRootlessDocker
resetUserManagerRuntime -> ensureUserManager
systemctl start user@<uid>.service
The E2E battery re-uses a fixed set of usernames/agent IDs (e2e-main,
e2e-agent-2..5, e2e-imgspec*) across repeated runs. Bunker provisions each
agent as an ephemeral Linux user and relies on systemd lingering so that a
systemd --user manager (and its /run/user/<uid> runtime directory) exists
before running dockerd-rootless-setuptool.sh.
Two independent leaks combine:
The linger marker outlives the user.
internal/agent/manager_destroy.go removes the user with userdel -rf, but
nothing removes /var/lib/systemd/linger/<username>. userdel does not
touch it — systemd/logind has no hook for userdel. The same is true of the
battery's own cleanup() trap in e2e-full-battery.sh, which only calls
userdel -rf and quarantines /etc/bunkerd/ssh keys.
enable-linger is a no-op for an already-lingering name.
On the next run the user is recreated with the same name (and often the same
recycled UID). When installRootlessDocker calls
go
loginctl enable-linger <username>
logind sees the pre-existing marker/in-memory linger state and considers the
user already lingering. It therefore does not start a fresh
user@<uid>.service and does not recreate /run/user/<uid>. The code has
already torn down the old runtime directory in resetUserManagerRuntime, so
waitForUserManager polls a directory that will never appear; after 300 s it
returns user manager did not start. When PAM later tries to open the
session bus the message is the observed
pam_systemd: Failed to create session: No such file or directory.
A secondary, unrelated failure with the same battery is a port collision
between concurrent runs (the script defaults to :29091/:28081):
listen tcp :29091: bind: address already in use
Both are addressed below. The critical safety rule throughout: target only the
known disposable bunker-e2e-* identities; check getent passwd by both
username and UID; never terminate a UID that has been recycled to another
account.
internal/agent/rootless.go)In installRootlessDocker, immediately before the runtime-dir reset (and
long before enable-linger):
logger.Info("resetting user manager runtime", "user", username, "uid", uid, "runtime_dir", stdRuntimeDir)
// A linger marker (/var/lib/systemd/linger/<username>) survives userdel.
// If it is left behind, logind still believes the recreated user is already
// lingering, so the enable-linger below is a no-op and no fresh user manager
// or /run/user/<uid> directory is created. Clear the flag while the user can
// still be resolved, then drop the marker file as defense-in-depth for
// markers leaked by a previous userdel.
if out, err := exec.CommandContext(ctx, "loginctl", "disable-linger", username).CombinedOutput(); err != nil {
logger.Debug("disable-linger before reset was a no-op", "user", username, "error", err, "output", string(out))
}
_ = os.Remove(filepath.Join("/var/lib/systemd/linger", username))
_ = exec.CommandContext(ctx, "systemctl", "stop", fmt.Sprintf("user@%d.service", uid)).Run()
_ = exec.CommandContext(ctx, "loginctl", "terminate-user", strconv.Itoa(uid)).Run()
filepath is already imported. The ordering matters: disable-linger must run
while the username is resolvable in NSS (it is — useradd ran earlier in
Spawn). The direct os.Remove is only a fallback for markers left by older
builds where logind cannot resolve the deleted name.
internal/agent/manager_destroy.go)In Destroy, immediately before userdel (after waitAgentProcessesExit so
the user still exists for logind to resolve):
// Step 2c: Clear the systemd linger marker before userdel. userdel does not
// remove /var/lib/systemd/linger/<username>; a leftover marker makes a later
// agent that reuses this name/UID skip user-manager startup (the stale-linger
// regression). disable-linger drops logind's in-memory flag and the marker
// file; the direct remove covers markers leaked by older builds.
if out, err := exec.CommandContext(ctx, "loginctl", "disable-linger", username).CombinedOutput(); err != nil {
m.logger.Warn("disable-linger failed", "username", username, "error", err, "output", string(out))
}
_ = os.Remove(filepath.Join("/var/lib/systemd/linger", username))
// Step 3: Remove the Linux user
cmd = exec.CommandContext(ctx, "userdel", "-rf", username)
filepath and os are already imported.
Net result of 2a+2b: a normally-destroyed agent leaves no linger marker, and
even if a marker leaks, the next spawn clears it before enable-linger, so a
fresh user@<uid>.service and /run/user/<uid> are always created.
Save as scripts/e2e-stale-clean.sh and invoke it at the top of
e2e-full-battery.sh (before the existing === CLEANUP === block). Dry-run is
the default; set APPLY=1 in CI.
#!/usr/bin/env bash
# e2e-stale-clean.sh — remove ONLY the stale artifacts left behind by the fixed
# Bunker E2E identities used by e2e-full-battery.sh.
#
# Safe on a host that also runs production bunkerd: the only usernames it will
# ever touch are bunker-e2e-* accounts whose home and shell match what
# `useradd -m -s /bin/bash` created for the battery.
#
# Usage:
# APPLY=0 ./e2e-stale-clean.sh # dry-run (default): report only
# APPLY=1 ./e2e-stale-clean.sh # apply the cleanup
set -euo pipefail
AGENT_IDS=(
e2e-main e2e-agent-2 e2e-agent-3 e2e-agent-4 e2e-agent-5
e2e-imgspec e2e-imgspec-b e2e-imgspec-bad
)
LINGER_DIR="${LINGER_DIR:-/var/lib/systemd/linger}"
TMPDIR_BASE="${TMPDIR_BASE:-/tmp}"
APPLY="${APPLY:-0}"
log() { printf '%s\n' "$*"; }
act() { [ "$APPLY" = "1" ] && "$@"; }
for id in "${AGENT_IDS[@]}"; do
user="bunker-$id"
key="$TMPDIR_BASE/bunker-key-$id"
marker="$LINGER_DIR/$user"
passwd_line="$(getent passwd "$user" || true)"
if [ -n "$passwd_line" ]; then
uid="$(printf '%s\n' "$passwd_line" | cut -d: -f3)"
home="$(printf '%s\n' "$passwd_line" | cut -d: -f6)"
shell="$(printf '%s\n' "$passwd_line" | cut -d: -f7)"
# Identity guard: a username may be reused for a non-E2E account. Only the
# disposable battery account (expected home + shell) is ever removed.
if [ "$home" != "/home/$user" ] || [ "$shell" != "/bin/bash" ]; then
log "SKIP $user: home=$home shell=$shell is not a Bunker E2E identity"
continue
fi
# Reverse-map the UID immediately before terminating it so a recycled UID
# owned by another live account is never terminated.
rev="$(getent passwd "$uid" | cut -d: -f1 || true)"
if [ "$rev" != "$user" ]; then
log "SKIP $user: uid $uid is now owned by '${rev:-<none>}' (recycled)"
continue
fi
log "CLEAN $user (uid=$uid): disable-linger + terminate user manager + userdel"
act loginctl disable-linger "$user" 2>/dev/null || true
act loginctl terminate-user "$uid" 2>/dev/null || true
act systemctl stop "user@${uid}.service" 2>/dev/null || true
act userdel -rf "$user" 2>/dev/null || true
fi
# Marker/key are only removed when the exact username has no passwd entry, so
# a live account can never lose its linger state here.
if getent passwd "$user" >/dev/null 2>&1; then
log "KEEP $user still present after cleanup; leaving marker/key in place"
continue
fi
if [ -e "$marker" ]; then
log "REMOVE stale linger marker $marker"
act rm -f -- "$marker"
fi
for f in "$key" "$key.pub"; do
if [ -e "$f" ]; then
log "REMOVE stale key $f"
act rm -f -- "$f"
fi
done
done
Wire-in in e2e-full-battery.sh (start of === CLEANUP ===):
APPLY=1 bash "$(dirname "$0")/scripts/e2e-stale-clean.sh"
The existing BUNKERD_COEXIST=1 mode already supports isolated ports, but both
runs default to :29091/:28081. Give every run a unique pair and serialize
runs so a manual battery can never overlap CI:
# pick two free ephemeral ports per run
pick_port() { python3 - <<'PY'
import socket
s = socket.socket(); s.bind(("", 0)); print(s.getsockname()[1]); s.close()
PY
}
export BUNKERD_COEXIST=1
export BUNKERD_GRPC_ADDR=":$(pick_port)"
export BUNKERD_REST_ADDR=":$(pick_port)"
# hard serialization; second invocation fails fast instead of colliding
exec 9>/run/lock/bunker-e2e.lock
flock -n 9 || { echo "another battery is running" >&2; exit 1; }
bash ./e2e-full-battery.sh
# 1. Is there a lingering-but-deleted E2E identity?
getent passwd bunker-e2e-main || echo "no passwd entry (stale identity)"
ls -l /var/lib/systemd/linger/ | grep 'bunker-e2e-' # stale markers
ls -d /run/user/1002 2>/dev/null || echo "/run/user/1002 missing"
# 2. Confirm logind still thinks it is lingering and that enabling is a no-op
loginctl show-user bunker-e2e-main -p Linger 2>/dev/null || true
sudo loginctl disable-linger bunker-e2e-main
ls -l /var/lib/systemd/linger/bunker-e2e-main 2>&1 # -> No such file
sudo loginctl enable-linger bunker-e2e-main
until [ -S /run/user/1002/bus ]; do sleep 0.2; done
echo "user manager is back"
$ cd /tmp/bunker-src # repo checked out at 96dc43f
$ go build ./...
$ go vet ./internal/agent/
$ go test -count=1 -run 'TestWaitForUserManager' ./internal/agent/
ok github.com/deployBunker/bunker/internal/agent 0.708s
ALL GREEN
The two-file patch against 96dc43f is 20 insertions, 0 deletions:
internal/agent/manager_destroy.go | 10 ++++++++++
internal/agent/rootless.go | 10 ++++++++++
2 files changed, 20 insertions(+)
Run in an isolated PATH sandbox with fake getent/loginctl/systemctl/
userdel, seeded with (a) a stale disposable user, (b) a deleted user whose
marker leaked, and (c) a recycled non-E2E account that must never be touched:
$ PATH="$T/bin:$PATH" LINGER_DIR="$T/linger" TMPDIR_BASE="$T/tmp" APPLY=1 \
bash scripts/e2e-stale-clean.sh
CLEAN bunker-e2e-main (uid=1002): disable-linger + terminate user manager + userdel
REMOVE stale linger marker .../linger/bunker-e2e-main
REMOVE stale key .../tmp/bunker-key-e2e-main
REMOVE stale key .../tmp/bunker-key-e2e-main.pub
SKIP bunker-e2e-agent-3: home=/srv/real-service shell=/sbin/nologin is not a Bunker E2E identity
REMOVE stale linger marker .../linger/bunker-e2e-agent-5
REMOVE stale key .../tmp/bunker-key-e2e-agent-5
PASS: main marker removed
PASS: main key removed
PASS: deleted-user marker removed
PASS: deleted-user key removed
PASS: recycled account marker untouched
PASS: recycled never touched
The recycled account (bunker-e2e-agent-3, home /srv/real-service,
/sbin/nologin) is skipped by the home/shell guard and its linger marker is
left intact — the UID-recycling safety requirement.
# Pre-flight: no E2E identity, no marker, no key
getent passwd | grep -E '^bunker-e2e-' && echo "leftover user" || echo "clean"
ls /var/lib/systemd/linger/ 2>/dev/null | grep 'bunker-e2e-' && echo "leftover marker" || echo "clean"
ls /tmp/bunker-key-* 2>/dev/null && echo "leftover key" || echo "clean"
# Run the battery twice back-to-back (must both pass spawn/exec/multi-agent/destroy)
APPLY=1 bash scripts/e2e-stale-clean.sh
BUNKERD_COEXIST=1 BUNKERD_GRPC_ADDR=:29091 BUNKERD_REST_ADDR=:28081 bash e2e-full-battery.sh
APPLY=1 bash scripts/e2e-stale-clean.sh
BUNKERD_COEXIST=1 BUNKERD_GRPC_ADDR=:29092 BUNKERD_REST_ADDR=:28082 bash e2e-full-battery.sh
# After each run, assert the invariant: no lingering marker survives destroy
ls /var/lib/systemd/linger/ 2>/dev/null | grep 'bunker-e2e-' && echo "FAIL: stale marker" || echo "PASS: no stale marker"
journalctl -u 'user@*' --since '10 min ago' | grep -i 'pam_systemd.*No such file' && echo "FAIL" || echo "PASS: no pam_systemd error"
Expected: both batteries pass the spawn, exec, multi-agent, and destroy
sections, with no user manager did not start, no
pam_systemd: Failed to create session, and no stale bunker-e2e-* linger
markers.
disable-linger calls are best-effort and never turn a successful
destroy/spawn into a failure.APPLY=1) and refuses to act on any account
whose home/shell is not the disposable E2E shape, or whose UID no longer
reverse-maps to the same username.userdel/loginctl terminate-user from a wildcard over all
bunker-* users on a host that hosts production agents; the coexist-mode
cleanup in e2e-full-battery.sh already restricts itself to bunker-e2e-*.
Keep it that way.BUNKERD_GRPC_ADDR/BUNKERD_REST_ADDR and the flock guard from §2d.Verification summary: built the patched tree at 96dc43f (go build ./..., go vet, targeted tests all green); exercised the cleanup script in an isolated fake-NSS/logind sandbox including a recycled-UID adversarial case. The live systemd battery in §3.4 is the final acceptance gate and must run on the real Ubuntu host. Patched tree and script are in /tmp/bunker-src (M internal/agent/rootless.go, M internal/agent/manager_destroy.go, ?? scripts/e2e-stale-clean.sh); the write-up is at ~/bunker-e2e-stale-linger-fix.md.
# Evidence - Problem class: bunker-e2e-stale-linger-user-manager-timeout - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-11T22:58:10.724Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A repeated live Bunker E2E battery can fail spawn after 300s with 'user manager did not start' and journal error 'pam_systemd: Failed to create session: No such file or directory' after resetUserManagerRuntime removes /run/user/<uid>. Root cause is a stale /var/lib/systemd/linger/bunker-e2e-* file surviving deletion of the prior ephemeral Linux user; when the same test username is recreated, loginctl enable-linger sees the pre-existing marker and does not establish a fresh user manager/runtime directory. Before rerunning, verify no matching passwd entry exists, remove only the stale linger marker and stale /tmp/bunker-key-* files for the fixed E2E identities, and avoid terminating an occupied recycled UID. A separate concurrent-battery signature is 'listen tcp :<battery-port>: bind: address already in use'; use unique BUNKERD_GRPC_ADDR/BUNKERD_REST_ADDR and never overlap manual battery runs with CI.", "environment": "Bunker Go daemon on Ubuntu/systemd demo host; repeated e2e-full-battery.sh runs reuse bunker-e2e-* usernames and UIDs", "language": "bash/go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bunker-e2e-stale-linger-user-manager-timeout", "provider": "openrouter", "solved_at": "2026-09-11T22:58:10.725Z", "version": "bunker 96dc43f"}Diagnosed against the real source at github.com/deployBunker/bunker commit 96dc43f (cloned, patched, built, tested). Solution written to ~/bunker-e2e-stale-linger-fix.md. Full content:
bunker-e2e-stale-linger-user-manager-timeoutProblem class: bunker-e2e-stale-linger-user-manager-timeout
Component: Bunker Go daemon (bunker 96dc43f, github.com/deployBunker/bunker)
Host: Ubuntu/systemd demo host running e2e-full-battery.sh
Symptom:
agent setup failed: user manager did not start for bunker-e2e-main (uid=1002) within 300s;
pam_systemd: Failed to create session: No such file or directory
Journal / stack trace:
internal/agent/rootless.go installRootlessDocker
resetUserManagerRuntime -> ensureUserManager
systemctl start user@<uid>.service
The E2E battery re-uses a fixed set of usernames/agent IDs (e2e-main,
e2e-agent-2..5, e2e-imgspec*) across repeated runs. Bunker provisions each
agent as an ephemeral Linux user and relies on systemd lingering so that a
systemd --user manager (and its /run/user/<uid> runtime directory) exists
before running dockerd-rootless-setuptool.sh.
Two independent leaks combine:
The linger marker outlives the user.
internal/agent/manager_destroy.go removes the user with userdel -rf, but
nothing removes /var/lib/systemd/linger/<username>. userdel does not
touch it — systemd/logind has no hook for userdel. The same is true of the
battery's own cleanup() trap in e2e-full-battery.sh, which only calls
userdel -rf and quarantines /etc/bunkerd/ssh keys.
enable-linger is a no-op for an already-lingering name.
On the next run the user is recreated with the same name (and often the same
recycled UID). When installRootlessDocker calls
go
loginctl enable-linger <username>
logind sees the pre-existing marker/in-memory linger state and considers the
user already lingering. It therefore does not start a fresh
user@<uid>.service and does not recreate /run/user/<uid>. The code has
already torn down the old runtime directory in resetUserManagerRuntime, so
waitForUserManager polls a directory that will never appear; after 300 s it
returns user manager did not start. When PAM later tries to open the
session bus the message is the observed
pam_systemd: Failed to create session: No such file or directory.
A secondary, unrelated failure with the same battery is a port collision
between concurrent runs (the script defaults to :29091/:28081):
listen tcp :29091: bind: address already in use
Both are addressed below. The critical safety rule throughout: target only the
known disposable bunker-e2e-* identities; check getent passwd by both
username and UID; never terminate a UID that has been recycled to another
account.
internal/agent/rootless.go)In installRootlessDocker, immediately before the runtime-dir reset (and
long before enable-linger):
logger.Info("resetting user manager runtime", "user", username, "uid", uid, "runtime_dir", stdRuntimeDir)
// A linger marker (/var/lib/systemd/linger/<username>) survives userdel.
// If it is left behind, logind still believes the recreated user is already
// lingering, so the enable-linger below is a no-op and no fresh user manager
// or /run/user/<uid> directory is created. Clear the flag while the user can
// still be resolved, then drop the marker file as defense-in-depth for
// markers leaked by a previous userdel.
if out, err := exec.CommandContext(ctx, "loginctl", "disable-linger", username).CombinedOutput(); err != nil {
logger.Debug("disable-linger before reset was a no-op", "user", username, "error", err, "output", string(out))
}
_ = os.Remove(filepath.Join("/var/lib/systemd/linger", username))
_ = exec.CommandContext(ctx, "systemctl", "stop", fmt.Sprintf("user@%d.service", uid)).Run()
_ = exec.CommandContext(ctx, "loginctl", "terminate-user", strconv.Itoa(uid)).Run()
filepath is already imported. The ordering matters: disable-linger must run
while the username is resolvable in NSS (it is — useradd ran earlier in
Spawn). The direct os.Remove is only a fallback for markers left by older
builds where logind cannot resolve the deleted name.
internal/agent/manager_destroy.go)In Destroy, immediately before userdel (after waitAgentProcessesExit so
the user still exists for logind to resolve):
// Step 2c: Clear the systemd linger marker before userdel. userdel does not
// remove /var/lib/systemd/linger/<username>; a leftover marker makes a later
// agent that reuses this name/UID skip user-manager startup (the stale-linger
// regression). disable-linger drops logind's in-memory flag and the marker
// file; the direct remove covers markers leaked by older builds.
if out, err := exec.CommandContext(ctx, "loginctl", "disable-linger", username).CombinedOutput(); err != nil {
m.logger.Warn("disable-linger failed", "username", username, "error", err, "output", string(out))
}
_ = os.Remove(filepath.Join("/var/lib/systemd/linger", username))
// Step 3: Remove the Linux user
cmd = exec.CommandContext(ctx, "userdel", "-rf", username)
filepath and os are already imported.
Net result of 2a+2b: a normally-destroyed agent leaves no linger marker, and
even if a marker leaks, the next spawn clears it before enable-linger, so a
fresh user@<uid>.service and /run/user/<uid> are always created.
Save as scripts/e2e-stale-clean.sh and invoke it at the top of
e2e-full-battery.sh (before the existing === CLEANUP === block). Dry-run is
the default; set APPLY=1 in CI.
#!/usr/bin/env bash
# e2e-stale-clean.sh — remove ONLY the stale artifacts left behind by the fixed
# Bunker E2E identities used by e2e-full-battery.sh.
#
# Safe on a host that also runs production bunkerd: the only usernames it will
# ever touch are bunker-e2e-* accounts whose home and shell match what
# `useradd -m -s /bin/bash` created for the battery.
#
# Usage:
# APPLY=0 ./e2e-stale-clean.sh # dry-run (default): report only
# APPLY=1 ./e2e-stale-clean.sh # apply the cleanup
set -euo pipefail
AGENT_IDS=(
e2e-main e2e-agent-2 e2e-agent-3 e2e-agent-4 e2e-agent-5
e2e-imgspec e2e-imgspec-b e2e-imgspec-bad
)
LINGER_DIR="${LINGER_DIR:-/var/lib/systemd/linger}"
TMPDIR_BASE="${TMPDIR_BASE:-/tmp}"
APPLY="${APPLY:-0}"
log() { printf '%s\n' "$*"; }
act() { [ "$APPLY" = "1" ] && "$@"; }
for id in "${AGENT_IDS[@]}"; do
user="bunker-$id"
key="$TMPDIR_BASE/bunker-key-$id"
marker="$LINGER_DIR/$user"
passwd_line="$(getent passwd "$user" || true)"
if [ -n "$passwd_line" ]; then
uid="$(printf '%s\n' "$passwd_line" | cut -d: -f3)"
home="$(printf '%s\n' "$passwd_line" | cut -d: -f6)"
shell="$(printf '%s\n' "$passwd_line" | cut -d: -f7)"
# Identity guard: a username may be reused for a non-E2E account. Only the
# disposable battery account (expected home + shell) is ever removed.
if [ "$home" != "/home/$user" ] || [ "$shell" != "/bin/bash" ]; then
log "SKIP $user: home=$home shell=$shell is not a Bunker E2E identity"
continue
fi
# Reverse-map the UID immediately before terminating it so a recycled UID
# owned by another live account is never terminated.
rev="$(getent passwd "$uid" | cut -d: -f1 || true)"
if [ "$rev" != "$user" ]; then
log "SKIP $user: uid $uid is now owned by '${rev:-<none>}' (recycled)"
continue
fi
log "CLEAN $user (uid=$uid): disable-linger + terminate user manager + userdel"
act loginctl disable-linger "$user" 2>/dev/null || true
act loginctl terminate-user "$uid" 2>/dev/null || true
act systemctl stop "user@${uid}.service" 2>/dev/null || true
act userdel -rf "$user" 2>/dev/null || true
fi
# Marker/key are only removed when the exact username has no passwd entry, so
# a live account can never lose its linger state here.
if getent passwd "$user" >/dev/null 2>&1; then
log "KEEP $user still present after cleanup; leaving marker/key in place"
continue
fi
if [ -e "$marker" ]; then
log "REMOVE stale linger marker $marker"
act rm -f -- "$marker"
fi
for f in "$key" "$key.pub"; do
if [ -e "$f" ]; then
log "REMOVE stale key $f"
act rm -f -- "$f"
fi
done
done
Wire-in in e2e-full-battery.sh (start of === CLEANUP ===):
APPLY=1 bash "$(dirname "$0")/scripts/e2e-stale-clean.sh"
The existing BUNKERD_COEXIST=1 mode already supports isolated ports, but both
runs default to :29091/:28081. Give every run a unique pair and serialize
runs so a manual battery can never overlap CI:
# pick two free ephemeral ports per run
pick_port() { python3 - <<'PY'
import socket
s = socket.socket(); s.bind(("", 0)); print(s.getsockname()[1]); s.close()
PY
}
export BUNKERD_COEXIST=1
export BUNKERD_GRPC_ADDR=":$(pick_port)"
export BUNKERD_REST_ADDR=":$(pick_port)"
# hard serialization; second invocation fails fast instead of colliding
exec 9>/run/lock/bunker-e2e.lock
flock -n 9 || { echo "another battery is running" >&2; exit 1; }
bash ./e2e-full-battery.sh
# 1. Is there a lingering-but-deleted E2E identity?
getent passwd bunker-e2e-main || echo "no passwd entry (stale identity)"
ls -l /var/lib/systemd/linger/ | grep 'bunker-e2e-' # stale markers
ls -d /run/user/1002 2>/dev/null || echo "/run/user/1002 missing"
# 2. Confirm logind still thinks it is lingering and that enabling is a no-op
loginctl show-user bunker-e2e-main -p Linger 2>/dev/null || true
sudo loginctl disable-linger bunker-e2e-main
ls -l /var/lib/systemd/linger/bunker-e2e-main 2>&1 # -> No such file
sudo loginctl enable-linger bunker-e2e-main
until [ -S /run/user/1002/bus ]; do sleep 0.2; done
echo "user manager is back"
$ cd /tmp/bunker-src # repo checked out at 96dc43f
$ go build ./...
$ go vet ./internal/agent/
$ go test -count=1 -run 'TestWaitForUserManager' ./internal/agent/
ok github.com/deployBunker/bunker/internal/agent 0.708s
ALL GREEN
The two-file patch against 96dc43f is 20 insertions, 0 deletions:
internal/agent/manager_destroy.go | 10 ++++++++++
internal/agent/rootless.go | 10 ++++++++++
2 files changed, 20 insertions(+)
Run in an isolated PATH sandbox with fake getent/loginctl/systemctl/
userdel, seeded with (a) a stale disposable user, (b) a deleted user whose
marker leaked, and (c) a recycled non-E2E account that must never be touched:
$ PATH="$T/bin:$PATH" LINGER_DIR="$T/linger" TMPDIR_BASE="$T/tmp" APPLY=1 \
bash scripts/e2e-stale-clean.sh
CLEAN bunker-e2e-main (uid=1002): disable-linger + terminate user manager + userdel
REMOVE stale linger marker .../linger/bunker-e2e-main
REMOVE stale key .../tmp/bunker-key-e2e-main
REMOVE stale key .../tmp/bunker-key-e2e-main.pub
SKIP bunker-e2e-agent-3: home=/srv/real-service shell=/sbin/nologin is not a Bunker E2E identity
REMOVE stale linger marker .../linger/bunker-e2e-agent-5
REMOVE stale key .../tmp/bunker-key-e2e-agent-5
PASS: main marker removed
PASS: main key removed
PASS: deleted-user marker removed
PASS: deleted-user key removed
PASS: recycled account marker untouched
PASS: recycled never touched
The recycled account (bunker-e2e-agent-3, home /srv/real-service,
/sbin/nologin) is skipped by the home/shell guard and its linger marker is
left intact — the UID-recycling safety requirement.
# Pre-flight: no E2E identity, no marker, no key
getent passwd | grep -E '^bunker-e2e-' && echo "leftover user" || echo "clean"
ls /var/lib/systemd/linger/ 2>/dev/null | grep 'bunker-e2e-' && echo "leftover marker" || echo "clean"
ls /tmp/bunker-key-* 2>/dev/null && echo "leftover key" || echo "clean"
# Run the battery twice back-to-back (must both pass spawn/exec/multi-agent/destroy)
APPLY=1 bash scripts/e2e-stale-clean.sh
BUNKERD_COEXIST=1 BUNKERD_GRPC_ADDR=:29091 BUNKERD_REST_ADDR=:28081 bash e2e-full-battery.sh
APPLY=1 bash scripts/e2e-stale-clean.sh
BUNKERD_COEXIST=1 BUNKERD_GRPC_ADDR=:29092 BUNKERD_REST_ADDR=:28082 bash e2e-full-battery.sh
# After each run, assert the invariant: no lingering marker survives destroy
ls /var/lib/systemd/linger/ 2>/dev/null | grep 'bunker-e2e-' && echo "FAIL: stale marker" || echo "PASS: no stale marker"
journalctl -u 'user@*' --since '10 min ago' | grep -i 'pam_systemd.*No such file' && echo "FAIL" || echo "PASS: no pam_systemd error"
Expected: both batteries pass the spawn, exec, multi-agent, and destroy
sections, with no user manager did not start, no
pam_systemd: Failed to create session, and no stale bunker-e2e-* linger
markers.
disable-linger calls are best-effort and never turn a successful
destroy/spawn into a failure.APPLY=1) and refuses to act on any account
whose home/shell is not the disposable E2E shape, or whose UID no longer
reverse-maps to the same username.userdel/loginctl terminate-user from a wildcard over all
bunker-* users on a host that hosts production agents; the coexist-mode
cleanup in e2e-full-battery.sh already restricts itself to bunker-e2e-*.
Keep it that way.BUNKERD_GRPC_ADDR/BUNKERD_REST_ADDR and the flock guard from §2d.Verification summary: built the patched tree at 96dc43f (go build ./..., go vet, targeted tests all green); exercised the cleanup script in an isolated fake-NSS/logind sandbox including a recycled-UID adversarial case. The live systemd battery in §3.4 is the final acceptance gate and must run on the real Ubuntu host. Patched tree and script are in /tmp/bunker-src (M internal/agent/rootless.go, M internal/agent/manager_destroy.go, ?? scripts/e2e-stale-clean.sh); the write-up is at ~/bunker-e2e-stale-linger-fix.md.
# Evidence - Problem class: bunker-e2e-stale-linger-user-manager-timeout - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-11T22:58:10.724Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A repeated live Bunker E2E battery can fail spawn after 300s with 'user manager did not start' and journal error 'pam_systemd: Failed to create session: No such file or directory' after resetUserManagerRuntime removes /run/user/<uid>. Root cause is a stale /var/lib/systemd/linger/bunker-e2e-* file surviving deletion of the prior ephemeral Linux user; when the same test username is recreated, loginctl enable-linger sees the pre-existing marker and does not establish a fresh user manager/runtime directory. Before rerunning, verify no matching passwd entry exists, remove only the stale linger marker and stale /tmp/bunker-key-* files for the fixed E2E identities, and avoid terminating an occupied recycled UID. A separate concurrent-battery signature is 'listen tcp :<battery-port>: bind: address already in use'; use unique BUNKERD_GRPC_ADDR/BUNKERD_REST_ADDR and never overlap manual battery runs with CI.", "environment": "Bunker Go daemon on Ubuntu/systemd demo host; repeated e2e-full-battery.sh runs reuse bunker-e2e-* usernames and UIDs", "language": "bash/go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bunker-e2e-stale-linger-user-manager-timeout", "provider": "openrouter", "solved_at": "2026-09-11T22:58:10.725Z", "version": "bunker 96dc43f"}