Host: cube-las-00 · rootless Docker, multi-lane (bunkerd)
I couldn't reach a live cube-las-00 from this sandbox (cube-root doesn't resolve; only ~ + the Pi agent exist), so the "verification" is delivered as concrete, runnable checks with expected results. The full runbook is at ~/eduos-voice-recovery.md. Here it is:
eduos-voice stack lost after retiring the bunker-eduos-agent (uid 1002) Docker laneHost: cube-las-00 · rootless Docker, multi-lane (bunkerd)
Symptom: voice-agent on <ip-address>:8099 → connection refused, while the app on :3002/:3003/:5433 is healthy.
Verdict: The voice compose project ran on the retired lane's rootless dockerd (/run/user/1002/docker.sock, uid 1002). Retiring that lane deleted the daemon and socket, taking all eduos-voice-* containers with it. The app survived only because it runs on the new lane's daemon (/run/user/1003/docker.sock, uid 1003) with restart policies.
| Project | Compose file | Network | Daemon | Status |
|---|---|---|---|---|
app (eduos) |
docker-compose.yml |
eduos_default |
uid 1003 rootless dockerd | healthy |
voice (eduos-voice) |
docker-compose.voice.yml (name: eduos-voice) |
joins eduos_default as external |
uid 1002 (retired) | gone |
Failure chain: lane user uid 1002 retired → logind stopped its session → rootless dockerd/rootlesskit exited → /run/user/1002/docker.sock removed → all 5 eduos-voice-* containers unreachable, :8099 no longer bound. The app is unaffected because its containers live in the uid-1003 daemon with restart: unless-stopped and auto-start at boot.
Why a naive restart fails on the new lane:
1. No images — eduos-voice-api/ml/tts were built into the dead uid-1002 content store.
2. No .env — env_file: ./services/voice-agent/.env is absent on the new home; up -d aborts.
3. No models — ${VOICE_MODELS_DIR:-~/voice-harness/models} points at the retired home; weights must be re-downloaded.
4. Stale DOCKER_HOST — may still export unix:///run/user/1002/docker.sock, masking the live app daemon.
ls -l /run/user/1002/docker.sock # expect: No such file or directory
ls -l /run/user/1003/docker.sock # expect: socket owned by the new lane user
ss -tlnp | grep -E ':(3002|3003|5433|8099)\b'
# :3002/:3003/:5433 -> rootlesskit (uid 1003) => app daemon alive
# :8099 -> nothing => voice gone
DOCKER_HOST=unix:///run/user/1003/docker.sock docker ps -a
grep -R "0002/docker.sock" ~/.bash_history ~/.zsh_history /home/*/.*history 2>/dev/null
env | grep DOCKER_HOST # the false-negative trap
Run as the new lane user (uid 1003):
NEW_USER=<uid-1003-lane-user>
REPO=/home/$NEW_USER/eduos
sudo -iu "$NEW_USER" # or machinectl shell
export XDG_RUNTIME_DIR=/run/user/1003
export DOCKER_HOST=unix:///run/user/1003/docker.sock
export DOCKER_HOST="unix:///run/user/1003/docker.sock"
grep -R "run/user/1002/docker.sock" /etc/profile /etc/profile.d ~/.bashrc ~/.profile 2>/dev/null
# fix/remove any hard-coded 1002 line. Optionally use `docker -H <sock> compose ...` everywhere.
export VOICE_MODELS_DIR=/home/$NEW_USER/voice-harness/models
mkdir -p "$VOICE_MODELS_DIR"
printf 'VOICE_MODELS_DIR=%s\n' "$VOICE_MODELS_DIR" > "$REPO/.env" # project-level, used for ${...}
Key subtlety: service
env_fileis not used for${...}interpolation; the project.env/ shell env is.
services/voice-agent/.env (voice-cloud, no GPU)cat > "$REPO/services/voice-agent/.env" <<EOF
VOICE_PROFILE=voice-cloud
VOICE_MODELS_DIR=$VOICE_MODELS_DIR
DEMO_USER=demo
DEMO_PASS=REPLACE_WITH_SECRET # required, or up -d/health fails
VOICE_API_KEY=REPLACE_WITH_SECRET
EOF
chmod 600 "$REPO/services/voice-agent/.env"
cd "$REPO"
docker compose -f docker-compose.voice.yml -p eduos-voice config >/tmp/voice.resolved.yml
grep "$VOICE_MODELS_DIR" /tmp/voice.resolved.yml # must NOT show bunker-eduos-agent
cd "$REPO/services/voice-agent"
python3 deploy/download_models.py --help
VOICE_MODELS_DIR="$VOICE_MODELS_DIR" python3 deploy/download_models.py --dest "$VOICE_MODELS_DIR"
du -sh "$VOICE_MODELS_DIR"
cd "$REPO"
docker compose -f docker-compose.voice.yml -p eduos-voice build
docker images --format '{{.Repository}}:{{.Tag}}' | grep '^eduos-voice-'
docker network ls | grep eduos_default
docker compose -f docker-compose.voice.yml -p eduos-voice up -d
# ensure each voice service has: restart: unless-stopped
sudo loginctl enable-linger "$NEW_USER"
sudo systemctl --user -M "$NEW_USER@" enable --now docker 2>/dev/null || true
Long-term: put models on a stable path (/srv/eduos-voice/models) and run voice under a system-level service account so lane retirement can never take it down again.
# Containers + port
docker compose -f docker-compose.voice.yml -p eduos-voice ps # all 5 healthy
ss -tlnp | grep ':8099' # <ip-address>:8099 rootlesskit (uid 1003)
# Health + end-to-end turn (DEMO_PASS from the env file)
curl -fsS -m 5 http://<ip-address>:8099/health
curl -fsS -m 30 -X POST http://<ip-address>:8099/turn \
-u "demo:${DEMO_PASS}" -H 'Content-Type: application/json' \
-d '{"session_id":"recovery-smoke","text":"hello"}' \
-o /tmp/turn.out -w 'HTTP %{http_code}\n'; cat /tmp/turn.out
# Cross-project network reachability over eduos_default
docker exec $(docker ps --filter name=voice-api -q | head -1) getent hosts eduos_default
# Survivability
sudo systemctl --user -M "$NEW_USER@" restart docker; sleep 10
ss -tlnp | grep ':8099' && curl -fsS http://<ip-address>:8099/health
# Close-out: no retired-lane references remain
grep -R "run/user/1002\|bunker-eduos-agent" "$REPO/docker-compose.voice.yml" "$REPO/.env" "$REPO/services/voice-agent/.env"
NEW_USER=<uid-1003-user>; REPO=/home/$NEW_USER/eduos
export XDG_RUNTIME_DIR=/run/user/1003 DOCKER_HOST=unix:///run/user/1003/docker.sock
export VOICE_MODELS_DIR=/home/$NEW_USER/voice-harness/models; mkdir -p "$VOICE_MODELS_DIR"
printf 'VOICE_MODELS_DIR=%s\n' "$VOICE_MODELS_DIR" > "$REPO/.env"
# recreate $REPO/services/voice-agent/.env (voice-cloud + DEMO_PASS + keys), then:
cd "$REPO/services/voice-agent" && VOICE_MODELS_DIR="$VOICE_MODELS_DIR" python3 deploy/download_models.py --dest "$VOICE_MODELS_DIR"
cd "$REPO" && docker compose -f docker-compose.voice.yml -p eduos-voice build && docker compose -f docker-compose.voice.yml -p eduos-voice up -d
ss -tlnp | grep ':8099' && curl -fsS http://<ip-address>:8099/health
The critical diagnosis: do not debug voice itself — the entire unit is on a dead daemon. Point DOCKER_HOST at uid 1003, rebuild images there, recreate .env + models under a uid-1003 path, then up -d; the external eduos_default network is already present because the app is healthy on that same daemon.
# Evidence - Problem class: eduos-voice-stack-lost-with-retired-lane-user-dockerd - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-21T01:41:33.540Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: voice-agent :8099 connection-refused while app :3002/:3003/:5433 healthy on the same host (cube-las-00, rootless docker multi-lane). No eduos-voice-* containers in docker ps -a on the lane's own rootless daemon; no modes service in docker-compose.yml (app unit) \u2014 the voice unit is a SEPARATE compose project docker-compose.voice.yml (name: eduos-voice, 5 services voice-api/modes/scorer/stt/tts, joins the app network eduos_default as external, publishes only <ip-address>:8099, env_file ./services/voice-agent/.env, weights mount ${VOICE_MODELS_DIR:-~/voice-harness/models}). ROOT CAUSE: the voice containers were created under the RETIRED bunker-eduos-agent lane user (uid 1002) rootless dockerd; its socket /run/user/1002/docker.sock no longer exists, so the whole eduos-voice unit disappeared when that lane was retired. The app stack survived because its containers live under the NEW lane user (uid 1003) daemon (/run/user/1003/docker.sock) with restart policies (auto-started at boot). COMPLICATIONS for recovery on the new lane: (1) voice images (eduos-voice-api/ml/tts) were never built on the uid-1003 daemon \u2014 a full `docker compose -p eduos-voice build` + models download (~GB whisper/scorer weights) + .env recreation (services/voice-agent/.env absent on the new lane home; compose env_file blocks up -d without it) are all required; (2) compose defaults reference the retired lane home for models. RECOVERY OUTLINE: recreate .env (voice-cloud profile; no local GPU on cube), set VOICE_MODELS_DIR to a uid-1003-owned dir, download models via services/voice-agent/deploy/download_models.py, build the three images, then `docker compose -p eduos-voice up -d` from the uid-1003 socket; verify /health on :8099 + POST /turn end-to-end (needs DEMO_PASS in the env_file). FORENSIC HINTS: check whose dockerd ran the containers via session history ('docker -H unix:///run/user/1002/docker.sock' invocations in wave lanes, t731 NOTES 'deployed scorer container'), and probe port ownership with ss -tlnp \u2014 a rootlesskit pid owning :3002/:3003/:5433 proves the app unit's dockerd is alive even when a stale DOCKER_HOST env variable says otherwise.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "eduos-voice-stack-lost-with-retired-lane-user-dockerd", "provider": "openrouter", "solved_at": "2026-09-21T01:41:33.565Z", "version": ""}I couldn't reach a live cube-las-00 from this sandbox (cube-root doesn't resolve; only ~ + the Pi agent exist), so the "verification" is delivered as concrete, runnable checks with expected results. The full runbook is at ~/eduos-voice-recovery.md. Here it is:
eduos-voice stack lost after retiring the bunker-eduos-agent (uid 1002) Docker laneHost: cube-las-00 · rootless Docker, multi-lane (bunkerd)
Symptom: voice-agent on <ip-address>:8099 → connection refused, while the app on :3002/:3003/:5433 is healthy.
Verdict: The voice compose project ran on the retired lane's rootless dockerd (/run/user/1002/docker.sock, uid 1002). Retiring that lane deleted the daemon and socket, taking all eduos-voice-* containers with it. The app survived only because it runs on the new lane's daemon (/run/user/1003/docker.sock, uid 1003) with restart policies.
| Project | Compose file | Network | Daemon | Status |
|---|---|---|---|---|
app (eduos) |
docker-compose.yml |
eduos_default |
uid 1003 rootless dockerd | healthy |
voice (eduos-voice) |
docker-compose.voice.yml (name: eduos-voice) |
joins eduos_default as external |
uid 1002 (retired) | gone |
Failure chain: lane user uid 1002 retired → logind stopped its session → rootless dockerd/rootlesskit exited → /run/user/1002/docker.sock removed → all 5 eduos-voice-* containers unreachable, :8099 no longer bound. The app is unaffected because its containers live in the uid-1003 daemon with restart: unless-stopped and auto-start at boot.
Why a naive restart fails on the new lane:
1. No images — eduos-voice-api/ml/tts were built into the dead uid-1002 content store.
2. No .env — env_file: ./services/voice-agent/.env is absent on the new home; up -d aborts.
3. No models — ${VOICE_MODELS_DIR:-~/voice-harness/models} points at the retired home; weights must be re-downloaded.
4. Stale DOCKER_HOST — may still export unix:///run/user/1002/docker.sock, masking the live app daemon.
ls -l /run/user/1002/docker.sock # expect: No such file or directory
ls -l /run/user/1003/docker.sock # expect: socket owned by the new lane user
ss -tlnp | grep -E ':(3002|3003|5433|8099)\b'
# :3002/:3003/:5433 -> rootlesskit (uid 1003) => app daemon alive
# :8099 -> nothing => voice gone
DOCKER_HOST=unix:///run/user/1003/docker.sock docker ps -a
grep -R "0002/docker.sock" ~/.bash_history ~/.zsh_history /home/*/.*history 2>/dev/null
env | grep DOCKER_HOST # the false-negative trap
Run as the new lane user (uid 1003):
NEW_USER=<uid-1003-lane-user>
REPO=/home/$NEW_USER/eduos
sudo -iu "$NEW_USER" # or machinectl shell
export XDG_RUNTIME_DIR=/run/user/1003
export DOCKER_HOST=unix:///run/user/1003/docker.sock
export DOCKER_HOST="unix:///run/user/1003/docker.sock"
grep -R "run/user/1002/docker.sock" /etc/profile /etc/profile.d ~/.bashrc ~/.profile 2>/dev/null
# fix/remove any hard-coded 1002 line. Optionally use `docker -H <sock> compose ...` everywhere.
export VOICE_MODELS_DIR=/home/$NEW_USER/voice-harness/models
mkdir -p "$VOICE_MODELS_DIR"
printf 'VOICE_MODELS_DIR=%s\n' "$VOICE_MODELS_DIR" > "$REPO/.env" # project-level, used for ${...}
Key subtlety: service
env_fileis not used for${...}interpolation; the project.env/ shell env is.
services/voice-agent/.env (voice-cloud, no GPU)cat > "$REPO/services/voice-agent/.env" <<EOF
VOICE_PROFILE=voice-cloud
VOICE_MODELS_DIR=$VOICE_MODELS_DIR
DEMO_USER=demo
DEMO_PASS=REPLACE_WITH_SECRET # required, or up -d/health fails
VOICE_API_KEY=REPLACE_WITH_SECRET
EOF
chmod 600 "$REPO/services/voice-agent/.env"
cd "$REPO"
docker compose -f docker-compose.voice.yml -p eduos-voice config >/tmp/voice.resolved.yml
grep "$VOICE_MODELS_DIR" /tmp/voice.resolved.yml # must NOT show bunker-eduos-agent
cd "$REPO/services/voice-agent"
python3 deploy/download_models.py --help
VOICE_MODELS_DIR="$VOICE_MODELS_DIR" python3 deploy/download_models.py --dest "$VOICE_MODELS_DIR"
du -sh "$VOICE_MODELS_DIR"
cd "$REPO"
docker compose -f docker-compose.voice.yml -p eduos-voice build
docker images --format '{{.Repository}}:{{.Tag}}' | grep '^eduos-voice-'
docker network ls | grep eduos_default
docker compose -f docker-compose.voice.yml -p eduos-voice up -d
# ensure each voice service has: restart: unless-stopped
sudo loginctl enable-linger "$NEW_USER"
sudo systemctl --user -M "$NEW_USER@" enable --now docker 2>/dev/null || true
Long-term: put models on a stable path (/srv/eduos-voice/models) and run voice under a system-level service account so lane retirement can never take it down again.
# Containers + port
docker compose -f docker-compose.voice.yml -p eduos-voice ps # all 5 healthy
ss -tlnp | grep ':8099' # <ip-address>:8099 rootlesskit (uid 1003)
# Health + end-to-end turn (DEMO_PASS from the env file)
curl -fsS -m 5 http://<ip-address>:8099/health
curl -fsS -m 30 -X POST http://<ip-address>:8099/turn \
-u "demo:${DEMO_PASS}" -H 'Content-Type: application/json' \
-d '{"session_id":"recovery-smoke","text":"hello"}' \
-o /tmp/turn.out -w 'HTTP %{http_code}\n'; cat /tmp/turn.out
# Cross-project network reachability over eduos_default
docker exec $(docker ps --filter name=voice-api -q | head -1) getent hosts eduos_default
# Survivability
sudo systemctl --user -M "$NEW_USER@" restart docker; sleep 10
ss -tlnp | grep ':8099' && curl -fsS http://<ip-address>:8099/health
# Close-out: no retired-lane references remain
grep -R "run/user/1002\|bunker-eduos-agent" "$REPO/docker-compose.voice.yml" "$REPO/.env" "$REPO/services/voice-agent/.env"
NEW_USER=<uid-1003-user>; REPO=/home/$NEW_USER/eduos
export XDG_RUNTIME_DIR=/run/user/1003 DOCKER_HOST=unix:///run/user/1003/docker.sock
export VOICE_MODELS_DIR=/home/$NEW_USER/voice-harness/models; mkdir -p "$VOICE_MODELS_DIR"
printf 'VOICE_MODELS_DIR=%s\n' "$VOICE_MODELS_DIR" > "$REPO/.env"
# recreate $REPO/services/voice-agent/.env (voice-cloud + DEMO_PASS + keys), then:
cd "$REPO/services/voice-agent" && VOICE_MODELS_DIR="$VOICE_MODELS_DIR" python3 deploy/download_models.py --dest "$VOICE_MODELS_DIR"
cd "$REPO" && docker compose -f docker-compose.voice.yml -p eduos-voice build && docker compose -f docker-compose.voice.yml -p eduos-voice up -d
ss -tlnp | grep ':8099' && curl -fsS http://<ip-address>:8099/health
The critical diagnosis: do not debug voice itself — the entire unit is on a dead daemon. Point DOCKER_HOST at uid 1003, rebuild images there, recreate .env + models under a uid-1003 path, then up -d; the external eduos_default network is already present because the app is healthy on that same daemon.
# Evidence - Problem class: eduos-voice-stack-lost-with-retired-lane-user-dockerd - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-21T01:41:33.540Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: voice-agent :8099 connection-refused while app :3002/:3003/:5433 healthy on the same host (cube-las-00, rootless docker multi-lane). No eduos-voice-* containers in docker ps -a on the lane's own rootless daemon; no modes service in docker-compose.yml (app unit) \u2014 the voice unit is a SEPARATE compose project docker-compose.voice.yml (name: eduos-voice, 5 services voice-api/modes/scorer/stt/tts, joins the app network eduos_default as external, publishes only <ip-address>:8099, env_file ./services/voice-agent/.env, weights mount ${VOICE_MODELS_DIR:-~/voice-harness/models}). ROOT CAUSE: the voice containers were created under the RETIRED bunker-eduos-agent lane user (uid 1002) rootless dockerd; its socket /run/user/1002/docker.sock no longer exists, so the whole eduos-voice unit disappeared when that lane was retired. The app stack survived because its containers live under the NEW lane user (uid 1003) daemon (/run/user/1003/docker.sock) with restart policies (auto-started at boot). COMPLICATIONS for recovery on the new lane: (1) voice images (eduos-voice-api/ml/tts) were never built on the uid-1003 daemon \u2014 a full `docker compose -p eduos-voice build` + models download (~GB whisper/scorer weights) + .env recreation (services/voice-agent/.env absent on the new lane home; compose env_file blocks up -d without it) are all required; (2) compose defaults reference the retired lane home for models. RECOVERY OUTLINE: recreate .env (voice-cloud profile; no local GPU on cube), set VOICE_MODELS_DIR to a uid-1003-owned dir, download models via services/voice-agent/deploy/download_models.py, build the three images, then `docker compose -p eduos-voice up -d` from the uid-1003 socket; verify /health on :8099 + POST /turn end-to-end (needs DEMO_PASS in the env_file). FORENSIC HINTS: check whose dockerd ran the containers via session history ('docker -H unix:///run/user/1002/docker.sock' invocations in wave lanes, t731 NOTES 'deployed scorer container'), and probe port ownership with ss -tlnp \u2014 a rootlesskit pid owning :3002/:3003/:5433 proves the app unit's dockerd is alive even when a stale DOCKER_HOST env variable says otherwise.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "eduos-voice-stack-lost-with-retired-lane-user-dockerd", "provider": "openrouter", "solved_at": "2026-09-21T01:41:33.565Z", "version": ""}