userns-uid-mapping-apparmor-setuid-denied
I reproduced the failure on this host and wrote the full self-contained solution to:
~/TJ-DF-015-solution.md — the requested markdown (title, root cause, exact fix, verification)~/scripts/fs-isolation-probe.py — runnable proof probe (exit 0 = FULL, 1 = DEGRADED)Layer 1 — empty uid_map is not an identity change (reproduced as unprivileged kara):
$ unshare --user bash -c 'id -u; grep ^Uid: /proc/self/status; cat ~/secret600.txt'
65534
Uid: 65534 65534 65534 65534
hello-600-secret # caller's mode-600 file still readable
A file created inside is kata(1000)/kara(1000)-owned on the host. /proc/self/uid_map is empty — only the display overflows to 65534.
Layer 2 — AppArmor denies capable() after userns creation:
$ grep -n 'deny capability' /etc/apparmor.d/unprivileged_userns
8: audit deny capability, # but: allow userns
allow userns lets the namespace exist; deny capability makes every later capable() (map write, setuid, setgid) fail with EPERM — including the nested PID namespace, so unshare --user --pid --fork fails outright:
$ unshare --user --map-root-user id -> write failed /proc/self/uid_map: Operation not permitted
$ unshare --user --pid --fork --kill-child=SIGKILL true -> unshare failed: Operation not permitted
The two layers are independent, so --keep-caps + setpriv still fails, and an intermediate exec clears namespace caps.
Probe result on this host:
RESULT: FS isolation DEGRADED - identity is illusory, caller files are reachable
EXIT=1
--map-users=65534:<subuid_start>:1 --map-groups=65534:<subgid_start>:1 -S 65534 -G 65534 --pid --fork true) — unshare --user true succeeding proves nothing.no filesystem isolation, export TERMINAL_JAIL_FS_ISOLATION=degraded.$HOME return EACCES; never infer from id -u or CapEff.sudo unshare --user --map-users=65534:100000:1 --map-groups=65534:100000:1 -S 65534 -G 65534 --pid --fork bash -c '<payload>', requiring setuid newuidmap + /etc/subuid/subgid range; -S/-G must be unshare's own options with no intermediate exec.--user "runs as nobody" as filesystem isolation.Note: the named repo paths (standalone/terminal-jail, …/interruptor/decider.py, README/docs/specs/skill) are not present in this container (only a deployed copy under /usr/local/lib/terminal-jail, which is root-owned/read-only), so the doc supplies exact drop-in patches for each rather than editing them in place.
# Evidence - Problem class: userns-uid-mapping-apparmor-setuid-denied - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-15T11:44:59.092Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM. `unshare --user` (no uid mapping) gives an ILLUSORY identity change, not isolation: inside the namespace `id -u` prints 65534 and /proc/self/status shows Uid: 65534 65534 65534 65534, yet the process still reads the caller's mode-600 files and creates files in the caller's HOME owned by the caller. Control: real `sudo -u nobody` on the same 600 file returns EACCES errno 13. Root cause of the illusion: with an EMPTY uid_map the kernel keeps the process's global kuid (the caller's, e.g. 1000) for every DAC check; only the *display* value overflows to 65534. A namespace without a mapping is not an identity change.\n\nSECOND SYMPTOM (the trap when you try to fix it). Adding `--map-users` does NOT fix it either; the process's credentials are unchanged unless something performs setuid INSIDE the namespace. Three spellings that all fail on an Ubuntu host with AppArmor unprivileged-userns restrictions (kernel.apparmor_restrict_unprivileged_userns=1), even though /proc/self/status inside the namespace shows FULL CapEff (000001ffffffffff): (a) `unshare --user --map-users=65534:100000:1 --map-groups=65534:100000:1 -S 65534 -G 65534 ...` -> `unshare: setgroups failed: Operation not permitted`; (b) same without -G -> `unshare: setuid() failed: Operation not permitted`; (c) `--keep-caps` + `setpriv --reuid 65534` (or python os.setresgid/setresuid) -> `setresuid failed: Operation not permitted`; also `--map-root-user` -> `write failed /proc/self/uid_map: Operation not permitted`. Diagnostic authority is `sudo dmesg | grep -i apparmor`, which shows the real reason: `operation=\"userns_create\" ... info=\"Userns create - transitioning profile\" profile=\"unconfined\" requested=\"userns_create\" target=\"unprivileged_userns\"` immediately followed by `apparmor=\"DENIED\" operation=\"capable\" class=\"cap\" profile=\"unprivileged_userns\" capability=7 capname=\"setuid\"` and `capability=6 capname=\"setgid\"`. Ubuntu's stock /etc/apparmor.d/unprivileged_userns profile denies CAP_SETUID/CAP_SETGID inside an unprivileged user namespace, so the capability is present in the cap set but every capable() check is denied. /proc/self/status is therefore NOT evidence of usable privilege.\n\nWORKING RECIPE (proven). Create the namespace as root (or on a host without the AppArmor restriction) and let util-linux do the setuid before any exec:\n sudo unshare --user --map-users=65534:100000:1 --map-groups=65534:100000:1 -S 65534 -G 65534 --pid --fork bash -c '<payload>'\n-> `id` = uid=65534(nobody) gid=65534(nogroup) and reading the caller's 600 file returns EACCES. Requirements: newuidmap/newgidmap setuid-root, an /etc/subuid + /etc/subgid range for the calling user (e.g. `kara:100000:65536`) \u2014 otherwise the map write fails with `newuidmap: uid range [65534-65535) -> [100000-100001) not allowed`, rc=1. Ordering matters: `-S/-G` must be unshare's own options; an intermediate `setpriv`/`python` exec clears the namespace capability set (exec of a non-root-in-init-ns process recomputes caps), which is why shim-based setuid attempts fail. Note the intermediate-exec rule and the AppArmor cap denial are independent: with `--keep-caps` the caps survive the exec and setresuid STILL fails under the AppArmor profile.\n\nFIX PATTERN for a jail/sandbox CLI. (1) Attempt the mapped launch (`--map-users=<nobody>:<subuid_start>:1 --map-groups=<nobody_gid>:<subgid_start>:1 -S <nobody> -G <nobody>`) and PREFLIGHT it with the same flags (`unshare <flags> true`, rc==0) \u2014 a preflight is required because a mapping-less namespace can be created successfully while the mapping itself is denied. (2) On preflight failure, fall back to the legacy mapping-less mode but print a LOUD warning containing a grep-able phrase such as `no filesystem isolation`, naming the likely cause, and export a machine-readable marker (e.g. `TERMINAL_JAIL_FS_ISOLATION=degraded`) so probes/tests can assert the degradation instead of assuming protection. (3) Ship a host-capability probe that actually attempts both negative probes inside the namespace (read a caller-owned 600 file, create a file in the caller's HOME) and reports FULL only when BOTH fail; never infer isolation from `id -u` output or from CapEff.\n\nHOST REMEDIATION (root, not code): `sysctl -w kernel.apparmor_restrict_unprivileged_userns=0`, or grant an AppArmor exception for the sandboxing binary, or ship newuidmap + subuid/subgid ranges. Without one of those, an unprivileged caller cannot obtain a real uid mapping on Ubuntu 24.04-era hosts, and the honest behaviour is loud degradation, not silent pretend-isolation.\n\nVERIFICATION. 4 probe generations on the affected host (unprivileged, privileged control, dmesg root-cause, and post-fix expectation) plus a docs audit: any document that says a `--user`/userns mode 'runs as nobody' must be read as PID/lifecycle containment unless a uid mapping is actually established.", "environment": "linux, Ubuntu with AppArmor unprivileged_userns profile, util-linux 2.41.3, kernel.apparmor_restrict_unprivileged_userns=1, /etc/subuid entry present, newuidmap setuid-root", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "userns-uid-mapping-apparmor-setuid-denied", "provider": "openrouter", "solved_at": "2026-09-15T11:44:59.092Z", "version": ""}I reproduced the failure on this host and wrote the full self-contained solution to:
~/TJ-DF-015-solution.md — the requested markdown (title, root cause, exact fix, verification)~/scripts/fs-isolation-probe.py — runnable proof probe (exit 0 = FULL, 1 = DEGRADED)Layer 1 — empty uid_map is not an identity change (reproduced as unprivileged kara):
$ unshare --user bash -c 'id -u; grep ^Uid: /proc/self/status; cat ~/secret600.txt'
65534
Uid: 65534 65534 65534 65534
hello-600-secret # caller's mode-600 file still readable
A file created inside is kata(1000)/kara(1000)-owned on the host. /proc/self/uid_map is empty — only the display overflows to 65534.
Layer 2 — AppArmor denies capable() after userns creation:
$ grep -n 'deny capability' /etc/apparmor.d/unprivileged_userns
8: audit deny capability, # but: allow userns
allow userns lets the namespace exist; deny capability makes every later capable() (map write, setuid, setgid) fail with EPERM — including the nested PID namespace, so unshare --user --pid --fork fails outright:
$ unshare --user --map-root-user id -> write failed /proc/self/uid_map: Operation not permitted
$ unshare --user --pid --fork --kill-child=SIGKILL true -> unshare failed: Operation not permitted
The two layers are independent, so --keep-caps + setpriv still fails, and an intermediate exec clears namespace caps.
Probe result on this host:
RESULT: FS isolation DEGRADED - identity is illusory, caller files are reachable
EXIT=1
--map-users=65534:<subuid_start>:1 --map-groups=65534:<subgid_start>:1 -S 65534 -G 65534 --pid --fork true) — unshare --user true succeeding proves nothing.no filesystem isolation, export TERMINAL_JAIL_FS_ISOLATION=degraded.$HOME return EACCES; never infer from id -u or CapEff.sudo unshare --user --map-users=65534:100000:1 --map-groups=65534:100000:1 -S 65534 -G 65534 --pid --fork bash -c '<payload>', requiring setuid newuidmap + /etc/subuid/subgid range; -S/-G must be unshare's own options with no intermediate exec.--user "runs as nobody" as filesystem isolation.Note: the named repo paths (standalone/terminal-jail, …/interruptor/decider.py, README/docs/specs/skill) are not present in this container (only a deployed copy under /usr/local/lib/terminal-jail, which is root-owned/read-only), so the doc supplies exact drop-in patches for each rather than editing them in place.
# Evidence - Problem class: userns-uid-mapping-apparmor-setuid-denied - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-15T11:44:59.092Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM. `unshare --user` (no uid mapping) gives an ILLUSORY identity change, not isolation: inside the namespace `id -u` prints 65534 and /proc/self/status shows Uid: 65534 65534 65534 65534, yet the process still reads the caller's mode-600 files and creates files in the caller's HOME owned by the caller. Control: real `sudo -u nobody` on the same 600 file returns EACCES errno 13. Root cause of the illusion: with an EMPTY uid_map the kernel keeps the process's global kuid (the caller's, e.g. 1000) for every DAC check; only the *display* value overflows to 65534. A namespace without a mapping is not an identity change.\n\nSECOND SYMPTOM (the trap when you try to fix it). Adding `--map-users` does NOT fix it either; the process's credentials are unchanged unless something performs setuid INSIDE the namespace. Three spellings that all fail on an Ubuntu host with AppArmor unprivileged-userns restrictions (kernel.apparmor_restrict_unprivileged_userns=1), even though /proc/self/status inside the namespace shows FULL CapEff (000001ffffffffff): (a) `unshare --user --map-users=65534:100000:1 --map-groups=65534:100000:1 -S 65534 -G 65534 ...` -> `unshare: setgroups failed: Operation not permitted`; (b) same without -G -> `unshare: setuid() failed: Operation not permitted`; (c) `--keep-caps` + `setpriv --reuid 65534` (or python os.setresgid/setresuid) -> `setresuid failed: Operation not permitted`; also `--map-root-user` -> `write failed /proc/self/uid_map: Operation not permitted`. Diagnostic authority is `sudo dmesg | grep -i apparmor`, which shows the real reason: `operation=\"userns_create\" ... info=\"Userns create - transitioning profile\" profile=\"unconfined\" requested=\"userns_create\" target=\"unprivileged_userns\"` immediately followed by `apparmor=\"DENIED\" operation=\"capable\" class=\"cap\" profile=\"unprivileged_userns\" capability=7 capname=\"setuid\"` and `capability=6 capname=\"setgid\"`. Ubuntu's stock /etc/apparmor.d/unprivileged_userns profile denies CAP_SETUID/CAP_SETGID inside an unprivileged user namespace, so the capability is present in the cap set but every capable() check is denied. /proc/self/status is therefore NOT evidence of usable privilege.\n\nWORKING RECIPE (proven). Create the namespace as root (or on a host without the AppArmor restriction) and let util-linux do the setuid before any exec:\n sudo unshare --user --map-users=65534:100000:1 --map-groups=65534:100000:1 -S 65534 -G 65534 --pid --fork bash -c '<payload>'\n-> `id` = uid=65534(nobody) gid=65534(nogroup) and reading the caller's 600 file returns EACCES. Requirements: newuidmap/newgidmap setuid-root, an /etc/subuid + /etc/subgid range for the calling user (e.g. `kara:100000:65536`) \u2014 otherwise the map write fails with `newuidmap: uid range [65534-65535) -> [100000-100001) not allowed`, rc=1. Ordering matters: `-S/-G` must be unshare's own options; an intermediate `setpriv`/`python` exec clears the namespace capability set (exec of a non-root-in-init-ns process recomputes caps), which is why shim-based setuid attempts fail. Note the intermediate-exec rule and the AppArmor cap denial are independent: with `--keep-caps` the caps survive the exec and setresuid STILL fails under the AppArmor profile.\n\nFIX PATTERN for a jail/sandbox CLI. (1) Attempt the mapped launch (`--map-users=<nobody>:<subuid_start>:1 --map-groups=<nobody_gid>:<subgid_start>:1 -S <nobody> -G <nobody>`) and PREFLIGHT it with the same flags (`unshare <flags> true`, rc==0) \u2014 a preflight is required because a mapping-less namespace can be created successfully while the mapping itself is denied. (2) On preflight failure, fall back to the legacy mapping-less mode but print a LOUD warning containing a grep-able phrase such as `no filesystem isolation`, naming the likely cause, and export a machine-readable marker (e.g. `TERMINAL_JAIL_FS_ISOLATION=degraded`) so probes/tests can assert the degradation instead of assuming protection. (3) Ship a host-capability probe that actually attempts both negative probes inside the namespace (read a caller-owned 600 file, create a file in the caller's HOME) and reports FULL only when BOTH fail; never infer isolation from `id -u` output or from CapEff.\n\nHOST REMEDIATION (root, not code): `sysctl -w kernel.apparmor_restrict_unprivileged_userns=0`, or grant an AppArmor exception for the sandboxing binary, or ship newuidmap + subuid/subgid ranges. Without one of those, an unprivileged caller cannot obtain a real uid mapping on Ubuntu 24.04-era hosts, and the honest behaviour is loud degradation, not silent pretend-isolation.\n\nVERIFICATION. 4 probe generations on the affected host (unprivileged, privileged control, dmesg root-cause, and post-fix expectation) plus a docs audit: any document that says a `--user`/userns mode 'runs as nobody' must be read as PID/lifecycle containment unless a uid mapping is actually established.", "environment": "linux, Ubuntu with AppArmor unprivileged_userns profile, util-linux 2.41.3, kernel.apparmor_restrict_unprivileged_userns=1, /etc/subuid entry present, newuidmap setuid-root", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "userns-uid-mapping-apparmor-setuid-denied", "provider": "openrouter", "solved_at": "2026-09-15T11:44:59.092Z", "version": ""}