su-login-strips-session-bus-environment
I diagnosed the defect against the actual repo (github.com/deployBunker/bunker @ a499216), implemented the fix, and verified it.
runUserSessionCmd set XDG_RUNTIME_DIR / DBUS_SESSION_BUS_ADDRESS on cmd.Env of the su process, but the command is su - (login mode), which resets the child environment. The command su runs therefore gets neither variable, and systemctl --user reports exactly Failed to connect to bus: No medium found (the no-variable error, distinct from the missing-socket No such file or directory). The 30 s readiness retries failed because the client never had a bus address at all.
The old unit test was a phantom pass: the fake userSessionRunner received runtimeDir as a separate argument and never inspected the command string, so it passed even though real su forwards nothing.
userSessionEnv ([]string for cmd.Env) with one shared in-band builder userSessionEnvPrefix/userSessionScript + shellQuote, and route all three consumers through it:su - user -c "XDG=... DBUS=... systemctl --user daemon-reload"FORCE_ROOTLESS_INSTALL=1 SKIP_IPTABLES=1, shell-quoted pathuserSessionRunner now receives the final command string, so fakes inspect exactly what production delivers.New internal/agent/session_env_inband_test.go uses a stub su that performs a real login-style env reset (env -i sh -c) plus a reporting systemctl:
TestProbeUserManagerReachable_EnvSurvivesLoginSu → sees XDG=[/run/user/1002] DBUS=[unix:path=/run/user/1002/bus]TestRunUserSessionCmd_OldEnvOnlyApproachIsTheBug → the pre-fix cmd.Env-only form prints XDG=[] DBUS=[] (negative control pinning the mechanism)TestRunRootlessInstallerCmd_EnvIsInBand, hostile-path quoting test — passgo build ./... OK; all user-manager/installer tests pass; the only full-package failure is the pre-existing, environment-only TestApplyUserSliceLimits_NotRoot_Coverage (read-only /etc/systemd/system), which fails identically on unmodified a499216.
Deliverable: ~/solution-su-login-strips-session-bus-environment.md (title, root-cause analysis, exact patch, test code, and host-level + live acceptance verification steps).
# Evidence - Problem class: su-login-strips-session-bus-environment - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-17T17:12:09.492Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: every `bunker spawn` handled by a daemon that was not launched by systemd dies at stage rootless-install with `session probe: Failed to connect to bus: No medium found`, while the systemd-managed production daemon on the SAME host with the SAME binary spawns agents successfully. The probe is `su - <agent-user> -c \"systemctl --user daemon-reload\"` run through a package seam (`userSessionRunner` -> `runUserSessionCmd`), which sets `cmd.Env = userSessionEnv(runtimeDir)` = `os.Environ()` + `XDG_RUNTIME_DIR=<runtimeDir>` + `DBUS_SESSION_BUS_ADDRESS=unix:path=<runtimeDir>/bus`. The code comment claims those values are 'layered on top of the inherited environment' and that probe, daemon-reload retry and installer all 'speak to the same manager'.\n\nROOT CAUSE: `su -` (login mode) RESETS the child environment, so the two variables the daemon sets on the `su` PROCESS never reach the command `su` runs. The agent session then has neither variable, and systemctl answers exactly 'Failed to connect to bus: No medium found'. Proven by three controlled probes on the host, all as root:\n(1) env stripping: `XDG_RUNTIME_DIR=/run/user/1002 DBUS_SESSION_BUS_ADDRESS=unix:path=/run/user/1002/bus su - kara -c 'echo XDG=[$XDG_RUNTIME_DIR] DBUS=[$DBUS_SESSION_BUS_ADDRESS]'` prints `XDG=[] DBUS=[]`.\n(2) the error text is the no-variable case: `env -u DBUS_SESSION_BUS_ADDRESS -u XDG_RUNTIME_DIR systemctl --user daemon-reload` prints verbatim 'Failed to connect to bus: No medium found' (a missing socket PATH prints 'No such file or directory' instead, so the message discriminates the two conditions).\n(3) in-band injection works: `su - kara -c \"XDG_RUNTIME_DIR=/run/user/1002 DBUS_SESSION_BUS_ADDRESS=unix:path=/run/user/1002/bus sh -c 'echo $XDG_RUNTIME_DIR'\"` prints `/run/user/1002`.\n\nWHY IT LOOKED LIKE A TIMING BUG FIRST (and the measurement that disproved it): an earlier fix added a bounded readiness wait that repeats the probe inside a 30s package budget, on the theory that the bus was simply not answering yet. With that build deployed, the daemon log shows the retry line at 16:38:09.324Z, the budget-exhaustion attribution at 16:38:39.387Z (the FULL 30s), the one-shot manager teardown + bring-up, and then a SECOND 30s readiness wait (16:38:40.167Z -> 16:39:10.204Z) failing with the same error: ~60s of probing across a manager restart with zero successful answers. A not-yet-answering socket answers; this one never does.\n\nCONTRAST THAT SCOPES THE DEFECT: the same host's systemd-managed daemon (whose own environment carries NEITHER variable) spawns fine, and the host journal records no `pam_systemd` 'New session ... of user bunker-<id>' line for the failing su calls (only `pam_unix(su-l:session)` open/close), i.e. no login stack supplies the bus address in that context either. So the only reliable channel is the command the session actually runs.\n\nFIX DIRECTION: stop relying on the su process environment. Build the session environment INSIDE the command string (e.g. `su - <user> -c 'XDG_RUNTIME_DIR=<dir> DBUS_SESSION_BUS_ADDRESS=unix:path=<dir>/bus <script>'`, or an `env ... sh -c` wrapper) for ALL THREE consumers (reachability probe, daemon-reload retry, rootless installer), keeping ONE shared construction so the 'same manager for every user-session command' property becomes true rather than aspirational. Unit-test it with a runner that inspects the COMMAND STRING, not cmd.Env: a seam test that asserts cmd.Env alone is exactly the phantom pass that let this ship (the fake session runner read the env the production path never delivers).\n\nVERIFICATION OF THE FIX DIRECTION: probe (3) above proves the in-band form reaches the child. The acceptance proof is live: the nested regression suite on the demo host currently reports PASS=16 / FAIL=14 and every red cell is a rootless-install spawn death; with the bus address in-band the same suite must clear those cells.", "environment": "Linux + systemd 255 (Ubuntu 24.04); root daemon (bunker) provisioning per-agent users and rootless Docker through each user's systemd user manager; the daemon uses `su - <user> -c <script>` for every user-session command and sets cmd.Env = os.Environ() + XDG_RUNTIME_DIR + DBUS_SESSION_BUS_ADDRESS; failure observed when the daemon itself was NOT launched by systemd (started from an ssh shell for the standalone E2E gate), while the systemd-managed production daemon on the same host spawns fine", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "su-login-strips-session-bus-environment", "provider": "openrouter", "solved_at": "2026-09-17T17:12:09.493Z", "version": "bunker main a499216 (defect present since the INT-CI-009 fix; reported at INT-SPAWN-003/004)"}I diagnosed the defect against the actual repo (github.com/deployBunker/bunker @ a499216), implemented the fix, and verified it.
runUserSessionCmd set XDG_RUNTIME_DIR / DBUS_SESSION_BUS_ADDRESS on cmd.Env of the su process, but the command is su - (login mode), which resets the child environment. The command su runs therefore gets neither variable, and systemctl --user reports exactly Failed to connect to bus: No medium found (the no-variable error, distinct from the missing-socket No such file or directory). The 30 s readiness retries failed because the client never had a bus address at all.
The old unit test was a phantom pass: the fake userSessionRunner received runtimeDir as a separate argument and never inspected the command string, so it passed even though real su forwards nothing.
userSessionEnv ([]string for cmd.Env) with one shared in-band builder userSessionEnvPrefix/userSessionScript + shellQuote, and route all three consumers through it:su - user -c "XDG=... DBUS=... systemctl --user daemon-reload"FORCE_ROOTLESS_INSTALL=1 SKIP_IPTABLES=1, shell-quoted pathuserSessionRunner now receives the final command string, so fakes inspect exactly what production delivers.New internal/agent/session_env_inband_test.go uses a stub su that performs a real login-style env reset (env -i sh -c) plus a reporting systemctl:
TestProbeUserManagerReachable_EnvSurvivesLoginSu → sees XDG=[/run/user/1002] DBUS=[unix:path=/run/user/1002/bus]TestRunUserSessionCmd_OldEnvOnlyApproachIsTheBug → the pre-fix cmd.Env-only form prints XDG=[] DBUS=[] (negative control pinning the mechanism)TestRunRootlessInstallerCmd_EnvIsInBand, hostile-path quoting test — passgo build ./... OK; all user-manager/installer tests pass; the only full-package failure is the pre-existing, environment-only TestApplyUserSliceLimits_NotRoot_Coverage (read-only /etc/systemd/system), which fails identically on unmodified a499216.
Deliverable: ~/solution-su-login-strips-session-bus-environment.md (title, root-cause analysis, exact patch, test code, and host-level + live acceptance verification steps).
# Evidence - Problem class: su-login-strips-session-bus-environment - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-17T17:12:09.492Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: every `bunker spawn` handled by a daemon that was not launched by systemd dies at stage rootless-install with `session probe: Failed to connect to bus: No medium found`, while the systemd-managed production daemon on the SAME host with the SAME binary spawns agents successfully. The probe is `su - <agent-user> -c \"systemctl --user daemon-reload\"` run through a package seam (`userSessionRunner` -> `runUserSessionCmd`), which sets `cmd.Env = userSessionEnv(runtimeDir)` = `os.Environ()` + `XDG_RUNTIME_DIR=<runtimeDir>` + `DBUS_SESSION_BUS_ADDRESS=unix:path=<runtimeDir>/bus`. The code comment claims those values are 'layered on top of the inherited environment' and that probe, daemon-reload retry and installer all 'speak to the same manager'.\n\nROOT CAUSE: `su -` (login mode) RESETS the child environment, so the two variables the daemon sets on the `su` PROCESS never reach the command `su` runs. The agent session then has neither variable, and systemctl answers exactly 'Failed to connect to bus: No medium found'. Proven by three controlled probes on the host, all as root:\n(1) env stripping: `XDG_RUNTIME_DIR=/run/user/1002 DBUS_SESSION_BUS_ADDRESS=unix:path=/run/user/1002/bus su - kara -c 'echo XDG=[$XDG_RUNTIME_DIR] DBUS=[$DBUS_SESSION_BUS_ADDRESS]'` prints `XDG=[] DBUS=[]`.\n(2) the error text is the no-variable case: `env -u DBUS_SESSION_BUS_ADDRESS -u XDG_RUNTIME_DIR systemctl --user daemon-reload` prints verbatim 'Failed to connect to bus: No medium found' (a missing socket PATH prints 'No such file or directory' instead, so the message discriminates the two conditions).\n(3) in-band injection works: `su - kara -c \"XDG_RUNTIME_DIR=/run/user/1002 DBUS_SESSION_BUS_ADDRESS=unix:path=/run/user/1002/bus sh -c 'echo $XDG_RUNTIME_DIR'\"` prints `/run/user/1002`.\n\nWHY IT LOOKED LIKE A TIMING BUG FIRST (and the measurement that disproved it): an earlier fix added a bounded readiness wait that repeats the probe inside a 30s package budget, on the theory that the bus was simply not answering yet. With that build deployed, the daemon log shows the retry line at 16:38:09.324Z, the budget-exhaustion attribution at 16:38:39.387Z (the FULL 30s), the one-shot manager teardown + bring-up, and then a SECOND 30s readiness wait (16:38:40.167Z -> 16:39:10.204Z) failing with the same error: ~60s of probing across a manager restart with zero successful answers. A not-yet-answering socket answers; this one never does.\n\nCONTRAST THAT SCOPES THE DEFECT: the same host's systemd-managed daemon (whose own environment carries NEITHER variable) spawns fine, and the host journal records no `pam_systemd` 'New session ... of user bunker-<id>' line for the failing su calls (only `pam_unix(su-l:session)` open/close), i.e. no login stack supplies the bus address in that context either. So the only reliable channel is the command the session actually runs.\n\nFIX DIRECTION: stop relying on the su process environment. Build the session environment INSIDE the command string (e.g. `su - <user> -c 'XDG_RUNTIME_DIR=<dir> DBUS_SESSION_BUS_ADDRESS=unix:path=<dir>/bus <script>'`, or an `env ... sh -c` wrapper) for ALL THREE consumers (reachability probe, daemon-reload retry, rootless installer), keeping ONE shared construction so the 'same manager for every user-session command' property becomes true rather than aspirational. Unit-test it with a runner that inspects the COMMAND STRING, not cmd.Env: a seam test that asserts cmd.Env alone is exactly the phantom pass that let this ship (the fake session runner read the env the production path never delivers).\n\nVERIFICATION OF THE FIX DIRECTION: probe (3) above proves the in-band form reaches the child. The acceptance proof is live: the nested regression suite on the demo host currently reports PASS=16 / FAIL=14 and every red cell is a rootless-install spawn death; with the bus address in-band the same suite must clear those cells.", "environment": "Linux + systemd 255 (Ubuntu 24.04); root daemon (bunker) provisioning per-agent users and rootless Docker through each user's systemd user manager; the daemon uses `su - <user> -c <script>` for every user-session command and sets cmd.Env = os.Environ() + XDG_RUNTIME_DIR + DBUS_SESSION_BUS_ADDRESS; failure observed when the daemon itself was NOT launched by systemd (started from an ssh shell for the standalone E2E gate), while the systemd-managed production daemon on the same host spawns fine", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "su-login-strips-session-bus-environment", "provider": "openrouter", "solved_at": "2026-09-17T17:12:09.493Z", "version": "bunker main a499216 (defect present since the INT-CI-009 fix; reported at INT-SPAWN-003/004)"}