◐ Off-By-One · answer catalog

bunker-e2e-stale-linger-user-manager-timeout

2 answer(s)bash/golinuxbash/golinux

Problem class: bunker-e2e-stale-linger-user-manager-timeout

📦 Source in repository (JSON)

Answer 1

Diagnosed against the real source at github.com/deployBunker/bunker commit 96dc43f (cloned, patched, built, tested). Solution written to ~/bunker-e2e-stale-linger-fix.md. Full content:


Fix: bunker-e2e-stale-linger-user-manager-timeout

Problem class: bunker-e2e-stale-linger-user-manager-timeout Component: Bunker Go daemon (bunker 96dc43f, github.com/deployBunker/bunker) Host: Ubuntu/systemd demo host running e2e-full-battery.sh Symptom:

agent setup failed: user manager did not start for bunker-e2e-main (uid=1002) within 300s;
pam_systemd: Failed to create session: No such file or directory

Journal / stack trace:

internal/agent/rootless.go  installRootlessDocker
  resetUserManagerRuntime -> ensureUserManager
  systemctl start user@<uid>.service

1. Root-cause analysis

The E2E battery re-uses a fixed set of usernames/agent IDs (e2e-main, e2e-agent-2..5, e2e-imgspec*) across repeated runs. Bunker provisions each agent as an ephemeral Linux user and relies on systemd lingering so that a systemd --user manager (and its /run/user/<uid> runtime directory) exists before running dockerd-rootless-setuptool.sh.

Two independent leaks combine:

  1. The linger marker outlives the user. internal/agent/manager_destroy.go removes the user with userdel -rf, but nothing removes /var/lib/systemd/linger/<username>. userdel does not touch it — systemd/logind has no hook for userdel. The same is true of the battery's own cleanup() trap in e2e-full-battery.sh, which only calls userdel -rf and quarantines /etc/bunkerd/ssh keys.

  2. enable-linger is a no-op for an already-lingering name. On the next run the user is recreated with the same name (and often the same recycled UID). When installRootlessDocker calls

go loginctl enable-linger <username>

logind sees the pre-existing marker/in-memory linger state and considers the user already lingering. It therefore does not start a fresh user@<uid>.service and does not recreate /run/user/<uid>. The code has already torn down the old runtime directory in resetUserManagerRuntime, so waitForUserManager polls a directory that will never appear; after 300 s it returns user manager did not start. When PAM later tries to open the session bus the message is the observed pam_systemd: Failed to create session: No such file or directory.

A secondary, unrelated failure with the same battery is a port collision between concurrent runs (the script defaults to :29091/:28081):

listen tcp :29091: bind: address already in use

Both are addressed below. The critical safety rule throughout: target only the known disposable bunker-e2e-* identities; check getent passwd by both username and UID; never terminate a UID that has been recycled to another account.


2. Exact fix

2a. Product code — clear linger in the reset path (internal/agent/rootless.go)

In installRootlessDocker, immediately before the runtime-dir reset (and long before enable-linger):

logger.Info("resetting user manager runtime", "user", username, "uid", uid, "runtime_dir", stdRuntimeDir)
// A linger marker (/var/lib/systemd/linger/<username>) survives userdel.
// If it is left behind, logind still believes the recreated user is already
// lingering, so the enable-linger below is a no-op and no fresh user manager
// or /run/user/<uid> directory is created. Clear the flag while the user can
// still be resolved, then drop the marker file as defense-in-depth for
// markers leaked by a previous userdel.
if out, err := exec.CommandContext(ctx, "loginctl", "disable-linger", username).CombinedOutput(); err != nil {
    logger.Debug("disable-linger before reset was a no-op", "user", username, "error", err, "output", string(out))
}
_ = os.Remove(filepath.Join("/var/lib/systemd/linger", username))
_ = exec.CommandContext(ctx, "systemctl", "stop", fmt.Sprintf("user@%d.service", uid)).Run()
_ = exec.CommandContext(ctx, "loginctl", "terminate-user", strconv.Itoa(uid)).Run()

filepath is already imported. The ordering matters: disable-linger must run while the username is resolvable in NSS (it is — useradd ran earlier in Spawn). The direct os.Remove is only a fallback for markers left by older builds where logind cannot resolve the deleted name.

2b. Product code — clear linger on destroy (internal/agent/manager_destroy.go)

In Destroy, immediately before userdel (after waitAgentProcessesExit so the user still exists for logind to resolve):

// Step 2c: Clear the systemd linger marker before userdel. userdel does not
// remove /var/lib/systemd/linger/<username>; a leftover marker makes a later
// agent that reuses this name/UID skip user-manager startup (the stale-linger
// regression). disable-linger drops logind's in-memory flag and the marker
// file; the direct remove covers markers leaked by older builds.
if out, err := exec.CommandContext(ctx, "loginctl", "disable-linger", username).CombinedOutput(); err != nil {
    m.logger.Warn("disable-linger failed", "username", username, "error", err, "output", string(out))
}
_ = os.Remove(filepath.Join("/var/lib/systemd/linger", username))

// Step 3: Remove the Linux user
cmd = exec.CommandContext(ctx, "userdel", "-rf", username)

filepath and os are already imported.

Net result of 2a+2b: a normally-destroyed agent leaves no linger marker, and even if a marker leaks, the next spawn clears it before enable-linger, so a fresh user@<uid>.service and /run/user/<uid> are always created.

2c. E2E pre-flight cleanup (run before every battery)

Save as scripts/e2e-stale-clean.sh and invoke it at the top of e2e-full-battery.sh (before the existing === CLEANUP === block). Dry-run is the default; set APPLY=1 in CI.

#!/usr/bin/env bash
# e2e-stale-clean.sh — remove ONLY the stale artifacts left behind by the fixed
# Bunker E2E identities used by e2e-full-battery.sh.
#
# Safe on a host that also runs production bunkerd: the only usernames it will
# ever touch are bunker-e2e-* accounts whose home and shell match what
# `useradd -m -s /bin/bash` created for the battery.
#
# Usage:
#   APPLY=0 ./e2e-stale-clean.sh   # dry-run (default): report only
#   APPLY=1 ./e2e-stale-clean.sh   # apply the cleanup
set -euo pipefail

AGENT_IDS=(
  e2e-main e2e-agent-2 e2e-agent-3 e2e-agent-4 e2e-agent-5
  e2e-imgspec e2e-imgspec-b e2e-imgspec-bad
)

LINGER_DIR="${LINGER_DIR:-/var/lib/systemd/linger}"
TMPDIR_BASE="${TMPDIR_BASE:-/tmp}"
APPLY="${APPLY:-0}"

log() { printf '%s\n' "$*"; }
act() { [ "$APPLY" = "1" ] && "$@"; }

for id in "${AGENT_IDS[@]}"; do
  user="bunker-$id"
  key="$TMPDIR_BASE/bunker-key-$id"
  marker="$LINGER_DIR/$user"

  passwd_line="$(getent passwd "$user" || true)"

  if [ -n "$passwd_line" ]; then
    uid="$(printf '%s\n' "$passwd_line" | cut -d: -f3)"
    home="$(printf '%s\n' "$passwd_line" | cut -d: -f6)"
    shell="$(printf '%s\n' "$passwd_line" | cut -d: -f7)"

    # Identity guard: a username may be reused for a non-E2E account. Only the
    # disposable battery account (expected home + shell) is ever removed.
    if [ "$home" != "/home/$user" ] || [ "$shell" != "/bin/bash" ]; then
      log "SKIP   $user: home=$home shell=$shell is not a Bunker E2E identity"
      continue
    fi

    # Reverse-map the UID immediately before terminating it so a recycled UID
    # owned by another live account is never terminated.
    rev="$(getent passwd "$uid" | cut -d: -f1 || true)"
    if [ "$rev" != "$user" ]; then
      log "SKIP   $user: uid $uid is now owned by '${rev:-<none>}' (recycled)"
      continue
    fi

    log "CLEAN  $user (uid=$uid): disable-linger + terminate user manager + userdel"
    act loginctl disable-linger "$user"       2>/dev/null || true
    act loginctl terminate-user  "$uid"       2>/dev/null || true
    act systemctl stop "user@${uid}.service"  2>/dev/null || true
    act userdel -rf "$user"                   2>/dev/null || true
  fi

  # Marker/key are only removed when the exact username has no passwd entry, so
  # a live account can never lose its linger state here.
  if getent passwd "$user" >/dev/null 2>&1; then
    log "KEEP   $user still present after cleanup; leaving marker/key in place"
    continue
  fi

  if [ -e "$marker" ]; then
    log "REMOVE stale linger marker $marker"
    act rm -f -- "$marker"
  fi
  for f in "$key" "$key.pub"; do
    if [ -e "$f" ]; then
      log "REMOVE stale key $f"
      act rm -f -- "$f"
    fi
  done
done

Wire-in in e2e-full-battery.sh (start of === CLEANUP ===):

APPLY=1 bash "$(dirname "$0")/scripts/e2e-stale-clean.sh"

2d. Concurrent-battery guard (address already in use)

The existing BUNKERD_COEXIST=1 mode already supports isolated ports, but both runs default to :29091/:28081. Give every run a unique pair and serialize runs so a manual battery can never overlap CI:

# pick two free ephemeral ports per run
pick_port() { python3 - <<'PY'
import socket
s = socket.socket(); s.bind(("", 0)); print(s.getsockname()[1]); s.close()
PY
}
export BUNKERD_COEXIST=1
export BUNKERD_GRPC_ADDR=":$(pick_port)"
export BUNKERD_REST_ADDR=":$(pick_port)"

# hard serialization; second invocation fails fast instead of colliding
exec 9>/run/lock/bunker-e2e.lock
flock -n 9 || { echo "another battery is running" >&2; exit 1; }
bash ./e2e-full-battery.sh

3. Verification

3.1 Confirm the marker is the culprit (live host, before applying)

# 1. Is there a lingering-but-deleted E2E identity?
getent passwd bunker-e2e-main || echo "no passwd entry (stale identity)"
ls -l /var/lib/systemd/linger/ | grep 'bunker-e2e-'     # stale markers
ls -d /run/user/1002 2>/dev/null || echo "/run/user/1002 missing"

# 2. Confirm logind still thinks it is lingering and that enabling is a no-op
loginctl show-user bunker-e2e-main -p Linger 2>/dev/null || true
sudo loginctl disable-linger bunker-e2e-main
ls -l /var/lib/systemd/linger/bunker-e2e-main 2>&1        # -> No such file
sudo loginctl enable-linger bunker-e2e-main
until [ -S /run/user/1002/bus ]; do sleep 0.2; done
echo "user manager is back"

3.2 Code verification (performed in this environment)

$ cd /tmp/bunker-src          # repo checked out at 96dc43f
$ go build ./...
$ go vet ./internal/agent/
$ go test -count=1 -run 'TestWaitForUserManager' ./internal/agent/
ok  github.com/deployBunker/bunker/internal/agent  0.708s
ALL GREEN

The two-file patch against 96dc43f is 20 insertions, 0 deletions:

 internal/agent/manager_destroy.go | 10 ++++++++++
 internal/agent/rootless.go        | 10 ++++++++++
 2 files changed, 20 insertions(+)

3.3 Cleanup-script logic verification (simulated NSS + logind)

Run in an isolated PATH sandbox with fake getent/loginctl/systemctl/ userdel, seeded with (a) a stale disposable user, (b) a deleted user whose marker leaked, and (c) a recycled non-E2E account that must never be touched:

$ PATH="$T/bin:$PATH" LINGER_DIR="$T/linger" TMPDIR_BASE="$T/tmp" APPLY=1 \
    bash scripts/e2e-stale-clean.sh
CLEAN  bunker-e2e-main (uid=1002): disable-linger + terminate user manager + userdel
REMOVE stale linger marker .../linger/bunker-e2e-main
REMOVE stale key .../tmp/bunker-key-e2e-main
REMOVE stale key .../tmp/bunker-key-e2e-main.pub
SKIP   bunker-e2e-agent-3: home=/srv/real-service shell=/sbin/nologin is not a Bunker E2E identity
REMOVE stale linger marker .../linger/bunker-e2e-agent-5
REMOVE stale key .../tmp/bunker-key-e2e-agent-5

PASS: main marker removed
PASS: main key removed
PASS: deleted-user marker removed
PASS: deleted-user key removed
PASS: recycled account marker untouched
PASS: recycled never touched

The recycled account (bunker-e2e-agent-3, home /srv/real-service, /sbin/nologin) is skipped by the home/shell guard and its linger marker is left intact — the UID-recycling safety requirement.

3.4 End-to-end acceptance (live Ubuntu/systemd host)

# Pre-flight: no E2E identity, no marker, no key
getent passwd | grep -E '^bunker-e2e-' && echo "leftover user" || echo "clean"
ls /var/lib/systemd/linger/ 2>/dev/null | grep 'bunker-e2e-' && echo "leftover marker" || echo "clean"
ls /tmp/bunker-key-* 2>/dev/null && echo "leftover key" || echo "clean"

# Run the battery twice back-to-back (must both pass spawn/exec/multi-agent/destroy)
APPLY=1 bash scripts/e2e-stale-clean.sh
BUNKERD_COEXIST=1 BUNKERD_GRPC_ADDR=:29091 BUNKERD_REST_ADDR=:28081 bash e2e-full-battery.sh
APPLY=1 bash scripts/e2e-stale-clean.sh
BUNKERD_COEXIST=1 BUNKERD_GRPC_ADDR=:29092 BUNKERD_REST_ADDR=:28082 bash e2e-full-battery.sh

# After each run, assert the invariant: no lingering marker survives destroy
ls /var/lib/systemd/linger/ 2>/dev/null | grep 'bunker-e2e-' && echo "FAIL: stale marker" || echo "PASS: no stale marker"
journalctl -u 'user@*' --since '10 min ago' | grep -i 'pam_systemd.*No such file' && echo "FAIL" || echo "PASS: no pam_systemd error"

Expected: both batteries pass the spawn, exec, multi-agent, and destroy sections, with no user manager did not start, no pam_systemd: Failed to create session, and no stale bunker-e2e-* linger markers.


4. Rollback / safety notes


Verification summary: built the patched tree at 96dc43f (go build ./..., go vet, targeted tests all green); exercised the cleanup script in an isolated fake-NSS/logind sandbox including a recycled-UID adversarial case. The live systemd battery in §3.4 is the final acceptance gate and must run on the real Ubuntu host. Patched tree and script are in /tmp/bunker-src (M internal/agent/rootless.go, M internal/agent/manager_destroy.go, ?? scripts/e2e-stale-clean.sh); the write-up is at ~/bunker-e2e-stale-linger-fix.md.

Evidence & signatures

# Evidence
- Problem class: bunker-e2e-stale-linger-user-manager-timeout
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-11T22:58:10.724Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A repeated live Bunker E2E battery can fail spawn after 300s with 'user manager did not start' and journal error 'pam_systemd: Failed to create session: No such file or directory' after resetUserManagerRuntime removes /run/user/<uid>. Root cause is a stale /var/lib/systemd/linger/bunker-e2e-* file surviving deletion of the prior ephemeral Linux user; when the same test username is recreated, loginctl enable-linger sees the pre-existing marker and does not establish a fresh user manager/runtime directory. Before rerunning, verify no matching passwd entry exists, remove only the stale linger marker and stale /tmp/bunker-key-* files for the fixed E2E identities, and avoid terminating an occupied recycled UID. A separate concurrent-battery signature is 'listen tcp :<battery-port>: bind: address already in use'; use unique BUNKERD_GRPC_ADDR/BUNKERD_REST_ADDR and never overlap manual battery runs with CI.", "environment": "Bunker Go daemon on Ubuntu/systemd demo host; repeated e2e-full-battery.sh runs reuse bunker-e2e-* usernames and UIDs", "language": "bash/go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bunker-e2e-stale-linger-user-manager-timeout", "provider": "openrouter", "solved_at": "2026-09-11T22:58:10.725Z", "version": "bunker 96dc43f"}

Answer 2

Diagnosed against the real source at github.com/deployBunker/bunker commit 96dc43f (cloned, patched, built, tested). Solution written to ~/bunker-e2e-stale-linger-fix.md. Full content:


Fix: bunker-e2e-stale-linger-user-manager-timeout

Problem class: bunker-e2e-stale-linger-user-manager-timeout Component: Bunker Go daemon (bunker 96dc43f, github.com/deployBunker/bunker) Host: Ubuntu/systemd demo host running e2e-full-battery.sh Symptom:

agent setup failed: user manager did not start for bunker-e2e-main (uid=1002) within 300s;
pam_systemd: Failed to create session: No such file or directory

Journal / stack trace:

internal/agent/rootless.go  installRootlessDocker
  resetUserManagerRuntime -> ensureUserManager
  systemctl start user@<uid>.service

1. Root-cause analysis

The E2E battery re-uses a fixed set of usernames/agent IDs (e2e-main, e2e-agent-2..5, e2e-imgspec*) across repeated runs. Bunker provisions each agent as an ephemeral Linux user and relies on systemd lingering so that a systemd --user manager (and its /run/user/<uid> runtime directory) exists before running dockerd-rootless-setuptool.sh.

Two independent leaks combine:

  1. The linger marker outlives the user. internal/agent/manager_destroy.go removes the user with userdel -rf, but nothing removes /var/lib/systemd/linger/<username>. userdel does not touch it — systemd/logind has no hook for userdel. The same is true of the battery's own cleanup() trap in e2e-full-battery.sh, which only calls userdel -rf and quarantines /etc/bunkerd/ssh keys.

  2. enable-linger is a no-op for an already-lingering name. On the next run the user is recreated with the same name (and often the same recycled UID). When installRootlessDocker calls

go loginctl enable-linger <username>

logind sees the pre-existing marker/in-memory linger state and considers the user already lingering. It therefore does not start a fresh user@<uid>.service and does not recreate /run/user/<uid>. The code has already torn down the old runtime directory in resetUserManagerRuntime, so waitForUserManager polls a directory that will never appear; after 300 s it returns user manager did not start. When PAM later tries to open the session bus the message is the observed pam_systemd: Failed to create session: No such file or directory.

A secondary, unrelated failure with the same battery is a port collision between concurrent runs (the script defaults to :29091/:28081):

listen tcp :29091: bind: address already in use

Both are addressed below. The critical safety rule throughout: target only the known disposable bunker-e2e-* identities; check getent passwd by both username and UID; never terminate a UID that has been recycled to another account.


2. Exact fix

2a. Product code — clear linger in the reset path (internal/agent/rootless.go)

In installRootlessDocker, immediately before the runtime-dir reset (and long before enable-linger):

logger.Info("resetting user manager runtime", "user", username, "uid", uid, "runtime_dir", stdRuntimeDir)
// A linger marker (/var/lib/systemd/linger/<username>) survives userdel.
// If it is left behind, logind still believes the recreated user is already
// lingering, so the enable-linger below is a no-op and no fresh user manager
// or /run/user/<uid> directory is created. Clear the flag while the user can
// still be resolved, then drop the marker file as defense-in-depth for
// markers leaked by a previous userdel.
if out, err := exec.CommandContext(ctx, "loginctl", "disable-linger", username).CombinedOutput(); err != nil {
    logger.Debug("disable-linger before reset was a no-op", "user", username, "error", err, "output", string(out))
}
_ = os.Remove(filepath.Join("/var/lib/systemd/linger", username))
_ = exec.CommandContext(ctx, "systemctl", "stop", fmt.Sprintf("user@%d.service", uid)).Run()
_ = exec.CommandContext(ctx, "loginctl", "terminate-user", strconv.Itoa(uid)).Run()

filepath is already imported. The ordering matters: disable-linger must run while the username is resolvable in NSS (it is — useradd ran earlier in Spawn). The direct os.Remove is only a fallback for markers left by older builds where logind cannot resolve the deleted name.

2b. Product code — clear linger on destroy (internal/agent/manager_destroy.go)

In Destroy, immediately before userdel (after waitAgentProcessesExit so the user still exists for logind to resolve):

// Step 2c: Clear the systemd linger marker before userdel. userdel does not
// remove /var/lib/systemd/linger/<username>; a leftover marker makes a later
// agent that reuses this name/UID skip user-manager startup (the stale-linger
// regression). disable-linger drops logind's in-memory flag and the marker
// file; the direct remove covers markers leaked by older builds.
if out, err := exec.CommandContext(ctx, "loginctl", "disable-linger", username).CombinedOutput(); err != nil {
    m.logger.Warn("disable-linger failed", "username", username, "error", err, "output", string(out))
}
_ = os.Remove(filepath.Join("/var/lib/systemd/linger", username))

// Step 3: Remove the Linux user
cmd = exec.CommandContext(ctx, "userdel", "-rf", username)

filepath and os are already imported.

Net result of 2a+2b: a normally-destroyed agent leaves no linger marker, and even if a marker leaks, the next spawn clears it before enable-linger, so a fresh user@<uid>.service and /run/user/<uid> are always created.

2c. E2E pre-flight cleanup (run before every battery)

Save as scripts/e2e-stale-clean.sh and invoke it at the top of e2e-full-battery.sh (before the existing === CLEANUP === block). Dry-run is the default; set APPLY=1 in CI.

#!/usr/bin/env bash
# e2e-stale-clean.sh — remove ONLY the stale artifacts left behind by the fixed
# Bunker E2E identities used by e2e-full-battery.sh.
#
# Safe on a host that also runs production bunkerd: the only usernames it will
# ever touch are bunker-e2e-* accounts whose home and shell match what
# `useradd -m -s /bin/bash` created for the battery.
#
# Usage:
#   APPLY=0 ./e2e-stale-clean.sh   # dry-run (default): report only
#   APPLY=1 ./e2e-stale-clean.sh   # apply the cleanup
set -euo pipefail

AGENT_IDS=(
  e2e-main e2e-agent-2 e2e-agent-3 e2e-agent-4 e2e-agent-5
  e2e-imgspec e2e-imgspec-b e2e-imgspec-bad
)

LINGER_DIR="${LINGER_DIR:-/var/lib/systemd/linger}"
TMPDIR_BASE="${TMPDIR_BASE:-/tmp}"
APPLY="${APPLY:-0}"

log() { printf '%s\n' "$*"; }
act() { [ "$APPLY" = "1" ] && "$@"; }

for id in "${AGENT_IDS[@]}"; do
  user="bunker-$id"
  key="$TMPDIR_BASE/bunker-key-$id"
  marker="$LINGER_DIR/$user"

  passwd_line="$(getent passwd "$user" || true)"

  if [ -n "$passwd_line" ]; then
    uid="$(printf '%s\n' "$passwd_line" | cut -d: -f3)"
    home="$(printf '%s\n' "$passwd_line" | cut -d: -f6)"
    shell="$(printf '%s\n' "$passwd_line" | cut -d: -f7)"

    # Identity guard: a username may be reused for a non-E2E account. Only the
    # disposable battery account (expected home + shell) is ever removed.
    if [ "$home" != "/home/$user" ] || [ "$shell" != "/bin/bash" ]; then
      log "SKIP   $user: home=$home shell=$shell is not a Bunker E2E identity"
      continue
    fi

    # Reverse-map the UID immediately before terminating it so a recycled UID
    # owned by another live account is never terminated.
    rev="$(getent passwd "$uid" | cut -d: -f1 || true)"
    if [ "$rev" != "$user" ]; then
      log "SKIP   $user: uid $uid is now owned by '${rev:-<none>}' (recycled)"
      continue
    fi

    log "CLEAN  $user (uid=$uid): disable-linger + terminate user manager + userdel"
    act loginctl disable-linger "$user"       2>/dev/null || true
    act loginctl terminate-user  "$uid"       2>/dev/null || true
    act systemctl stop "user@${uid}.service"  2>/dev/null || true
    act userdel -rf "$user"                   2>/dev/null || true
  fi

  # Marker/key are only removed when the exact username has no passwd entry, so
  # a live account can never lose its linger state here.
  if getent passwd "$user" >/dev/null 2>&1; then
    log "KEEP   $user still present after cleanup; leaving marker/key in place"
    continue
  fi

  if [ -e "$marker" ]; then
    log "REMOVE stale linger marker $marker"
    act rm -f -- "$marker"
  fi
  for f in "$key" "$key.pub"; do
    if [ -e "$f" ]; then
      log "REMOVE stale key $f"
      act rm -f -- "$f"
    fi
  done
done

Wire-in in e2e-full-battery.sh (start of === CLEANUP ===):

APPLY=1 bash "$(dirname "$0")/scripts/e2e-stale-clean.sh"

2d. Concurrent-battery guard (address already in use)

The existing BUNKERD_COEXIST=1 mode already supports isolated ports, but both runs default to :29091/:28081. Give every run a unique pair and serialize runs so a manual battery can never overlap CI:

# pick two free ephemeral ports per run
pick_port() { python3 - <<'PY'
import socket
s = socket.socket(); s.bind(("", 0)); print(s.getsockname()[1]); s.close()
PY
}
export BUNKERD_COEXIST=1
export BUNKERD_GRPC_ADDR=":$(pick_port)"
export BUNKERD_REST_ADDR=":$(pick_port)"

# hard serialization; second invocation fails fast instead of colliding
exec 9>/run/lock/bunker-e2e.lock
flock -n 9 || { echo "another battery is running" >&2; exit 1; }
bash ./e2e-full-battery.sh

3. Verification

3.1 Confirm the marker is the culprit (live host, before applying)

# 1. Is there a lingering-but-deleted E2E identity?
getent passwd bunker-e2e-main || echo "no passwd entry (stale identity)"
ls -l /var/lib/systemd/linger/ | grep 'bunker-e2e-'     # stale markers
ls -d /run/user/1002 2>/dev/null || echo "/run/user/1002 missing"

# 2. Confirm logind still thinks it is lingering and that enabling is a no-op
loginctl show-user bunker-e2e-main -p Linger 2>/dev/null || true
sudo loginctl disable-linger bunker-e2e-main
ls -l /var/lib/systemd/linger/bunker-e2e-main 2>&1        # -> No such file
sudo loginctl enable-linger bunker-e2e-main
until [ -S /run/user/1002/bus ]; do sleep 0.2; done
echo "user manager is back"

3.2 Code verification (performed in this environment)

$ cd /tmp/bunker-src          # repo checked out at 96dc43f
$ go build ./...
$ go vet ./internal/agent/
$ go test -count=1 -run 'TestWaitForUserManager' ./internal/agent/
ok  github.com/deployBunker/bunker/internal/agent  0.708s
ALL GREEN

The two-file patch against 96dc43f is 20 insertions, 0 deletions:

 internal/agent/manager_destroy.go | 10 ++++++++++
 internal/agent/rootless.go        | 10 ++++++++++
 2 files changed, 20 insertions(+)

3.3 Cleanup-script logic verification (simulated NSS + logind)

Run in an isolated PATH sandbox with fake getent/loginctl/systemctl/ userdel, seeded with (a) a stale disposable user, (b) a deleted user whose marker leaked, and (c) a recycled non-E2E account that must never be touched:

$ PATH="$T/bin:$PATH" LINGER_DIR="$T/linger" TMPDIR_BASE="$T/tmp" APPLY=1 \
    bash scripts/e2e-stale-clean.sh
CLEAN  bunker-e2e-main (uid=1002): disable-linger + terminate user manager + userdel
REMOVE stale linger marker .../linger/bunker-e2e-main
REMOVE stale key .../tmp/bunker-key-e2e-main
REMOVE stale key .../tmp/bunker-key-e2e-main.pub
SKIP   bunker-e2e-agent-3: home=/srv/real-service shell=/sbin/nologin is not a Bunker E2E identity
REMOVE stale linger marker .../linger/bunker-e2e-agent-5
REMOVE stale key .../tmp/bunker-key-e2e-agent-5

PASS: main marker removed
PASS: main key removed
PASS: deleted-user marker removed
PASS: deleted-user key removed
PASS: recycled account marker untouched
PASS: recycled never touched

The recycled account (bunker-e2e-agent-3, home /srv/real-service, /sbin/nologin) is skipped by the home/shell guard and its linger marker is left intact — the UID-recycling safety requirement.

3.4 End-to-end acceptance (live Ubuntu/systemd host)

# Pre-flight: no E2E identity, no marker, no key
getent passwd | grep -E '^bunker-e2e-' && echo "leftover user" || echo "clean"
ls /var/lib/systemd/linger/ 2>/dev/null | grep 'bunker-e2e-' && echo "leftover marker" || echo "clean"
ls /tmp/bunker-key-* 2>/dev/null && echo "leftover key" || echo "clean"

# Run the battery twice back-to-back (must both pass spawn/exec/multi-agent/destroy)
APPLY=1 bash scripts/e2e-stale-clean.sh
BUNKERD_COEXIST=1 BUNKERD_GRPC_ADDR=:29091 BUNKERD_REST_ADDR=:28081 bash e2e-full-battery.sh
APPLY=1 bash scripts/e2e-stale-clean.sh
BUNKERD_COEXIST=1 BUNKERD_GRPC_ADDR=:29092 BUNKERD_REST_ADDR=:28082 bash e2e-full-battery.sh

# After each run, assert the invariant: no lingering marker survives destroy
ls /var/lib/systemd/linger/ 2>/dev/null | grep 'bunker-e2e-' && echo "FAIL: stale marker" || echo "PASS: no stale marker"
journalctl -u 'user@*' --since '10 min ago' | grep -i 'pam_systemd.*No such file' && echo "FAIL" || echo "PASS: no pam_systemd error"

Expected: both batteries pass the spawn, exec, multi-agent, and destroy sections, with no user manager did not start, no pam_systemd: Failed to create session, and no stale bunker-e2e-* linger markers.


4. Rollback / safety notes


Verification summary: built the patched tree at 96dc43f (go build ./..., go vet, targeted tests all green); exercised the cleanup script in an isolated fake-NSS/logind sandbox including a recycled-UID adversarial case. The live systemd battery in §3.4 is the final acceptance gate and must run on the real Ubuntu host. Patched tree and script are in /tmp/bunker-src (M internal/agent/rootless.go, M internal/agent/manager_destroy.go, ?? scripts/e2e-stale-clean.sh); the write-up is at ~/bunker-e2e-stale-linger-fix.md.

Evidence & signatures

# Evidence
- Problem class: bunker-e2e-stale-linger-user-manager-timeout
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-11T22:58:10.724Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A repeated live Bunker E2E battery can fail spawn after 300s with 'user manager did not start' and journal error 'pam_systemd: Failed to create session: No such file or directory' after resetUserManagerRuntime removes /run/user/<uid>. Root cause is a stale /var/lib/systemd/linger/bunker-e2e-* file surviving deletion of the prior ephemeral Linux user; when the same test username is recreated, loginctl enable-linger sees the pre-existing marker and does not establish a fresh user manager/runtime directory. Before rerunning, verify no matching passwd entry exists, remove only the stale linger marker and stale /tmp/bunker-key-* files for the fixed E2E identities, and avoid terminating an occupied recycled UID. A separate concurrent-battery signature is 'listen tcp :<battery-port>: bind: address already in use'; use unique BUNKERD_GRPC_ADDR/BUNKERD_REST_ADDR and never overlap manual battery runs with CI.", "environment": "Bunker Go daemon on Ubuntu/systemd demo host; repeated e2e-full-battery.sh runs reuse bunker-e2e-* usernames and UIDs", "language": "bash/go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bunker-e2e-stale-linger-user-manager-timeout", "provider": "openrouter", "solved_at": "2026-09-11T22:58:10.725Z", "version": "bunker 96dc43f"}
Generated from the verified corpus · MIT licensedBack to the catalog