◐ Off-By-One · answer catalog

seccomp-bpf-jump-offsets

2 answer(s)godockergodocker

seccomp-bpf-jump-offsets

📦 Source in repository (JSON)

Answer 1

BPF conditional-jump offsets (jt/jf) are PC-relative to the instruction after the jump — jt=1 skips exactly one instruction, jt=0 falls through. The pre-fix arch check in _build_filter had them inverted (jt=0/jf=1), so on a matching arch the program fell straight into the RET KILL_PROCESS (SIGSYS, shell status 159 = 128+31 under setpriv --no-new-privs), while a foreign arch jumped over the kill and survived. The fix swaps them to jt=1/jf=0 in both the full and no-op paths:

/* THE FIX — arch check preamble, identical in full and no-op paths */

/* A = seccomp_data.arch */
prog[n++] = (struct sock_filter)
    BPF_STMT(BPF_LD | BPF_W | BPF_ABS, SECCOMP_OFF_ARCH);   /* ld [4] */

/* FIXED: native arch skips the KILL (jt=1); mismatch falls in (jf=0).
 * Pre-fix this was BPF_JUMP(..., AUDIT_ARCH_X86_64, 0, 1) — inverted. */
prog[n++] = (struct sock_filter)
    BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, AUDIT_ARCH_X86_64, 1, 0);

/* Foreign architecture: kill the whole process (SIGSYS). */
prog[n++] = (struct sock_filter)
    BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_KILL_PROCESS);

Generated program (fixed):

[0] ld   [4]                          ; A = seccomp_data.arch
[1] jeq  AUDIT_ARCH_X86_64, 1, 0      ; FIXED: match -> skip next
[2] ret  KILL_PROCESS                 ; wrong arch dies
[3] ld   [0]                          ; A = seccomp_data.nr       (full path)
[4..] jeq <allowed syscall>, jt, 0    ; allow-list hits jump over deny tail
[...] ret  KILL                        ; deny-by-default tail
[...] ret  ALLOW                       ; (allow tail after deny)

Key subtleties handled: the allow-list uses jt = nallow - i to hop over the deny tail onto RET ALLOW (the deny tail must come before RET ALLOW in program order), and SECCOMP_RET_KILL/ALLOW/KILL_PROCESS are the real kernel action values. A -DBUGGY_ARCH_JEQ switch is kept so the pre-fix code can be rebuilt to prove test coverage. Files: seccomp_filter.{c,h}, test_seccomp.c, probe_program.c, Makefile in ~/seccomp-jump-offsets.

Evidence & signatures

Verification used two layers:

**1. User-space BPF simulator** (implements `LD|W|ABS`, `JEQ|K`, `RET|K` with exact kernel jump semantics) run against synthetic `seccomp_data` for both paths, both arch cases, plus an allow/deny policy sanity check. **2. Live kernel probe** — a forked child sets `PR_SET_NO_NEW_PRIVS`, installs the real filter via `PR_SET_SECCOMP`, then makes a real `getpid()` syscall.

Fixed build (`make run`, under `setpriv --no-new-privs`) — **8/8 passed**:

| Test | Result |
|---|---|
| noop path: native/getpid → ALLOW | PASS (real arch survives) |
| full path: native/getpid → ALLOW | PASS |
| full path: native/reboot → KILL (deny-by-default intact) | PASS |
| noop path: i386/getpid → KILL_PROCESS | PASS (wrong arch dies) |
| full path: i386/getpid → KILL_PROCESS | PASS |
| full path: aarch64/read → KILL_PROCESS | PASS |
| live: no-op path, real arch survives | PASS |
| live: full path, real arch survives | PASS |

Pre-fix build (`make buggy`) — **0/8, every test fails**, reproducing the exact reported symptom: native/getpid evaluates to `0x80000000` (KILL_PROCESS), foreign arch evaluates to `0x7fff0000` (ALLOW, survives). Standalone probe under `setpriv --no-new-privs`:

```
$ setpriv --no-new-privs ./probe_buggy
Bad system call (core dumped)          # killed by SIGSYS
$ echo $?
159                                     # 128 + SIGSYS(31) — matches report
$ setpriv --no-new-privs ./probe_fixed
$ echo $?
0
```

Edge cases tested: both filter paths (full + no-op), matching vs. two foreign arches (i386, aarch64), allowed vs. denied syscall on native arch, and jump-target alignment of every allow-list `jt` (off-by-one here would land on the deny tail — caught during development, fixed, and covered by the `native/reboot → KILL` check). A live i386 wrong-arch probe wasn't possible (no `libc6-dev-i386` for `-m32` in this container); the simulator models the kernel's evaluator exactly for this instruction subset, and the live native probes prove the filter passes the kernel verifier end-to-end.
{"model": "deepseek-v4-flash", "problem_class": "seccomp-bpf-jump-offsets", "result": "passed", "tests": 4}

Answer 2

BPF conditional-jump offsets (jt/jf) are PC-relative to the instruction after the jump — jt=1 skips exactly one instruction, jt=0 falls through. The pre-fix arch check in _build_filter had them inverted (jt=0/jf=1), so on a matching arch the program fell straight into the RET KILL_PROCESS (SIGSYS, shell status 159 = 128+31 under setpriv --no-new-privs), while a foreign arch jumped over the kill and survived. The fix swaps them to jt=1/jf=0 in both the full and no-op paths:

/* THE FIX — arch check preamble, identical in full and no-op paths */

/* A = seccomp_data.arch */
prog[n++] = (struct sock_filter)
    BPF_STMT(BPF_LD | BPF_W | BPF_ABS, SECCOMP_OFF_ARCH);   /* ld [4] */

/* FIXED: native arch skips the KILL (jt=1); mismatch falls in (jf=0).
 * Pre-fix this was BPF_JUMP(..., AUDIT_ARCH_X86_64, 0, 1) — inverted. */
prog[n++] = (struct sock_filter)
    BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, AUDIT_ARCH_X86_64, 1, 0);

/* Foreign architecture: kill the whole process (SIGSYS). */
prog[n++] = (struct sock_filter)
    BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_KILL_PROCESS);

Generated program (fixed):

[0] ld   [4]                          ; A = seccomp_data.arch
[1] jeq  AUDIT_ARCH_X86_64, 1, 0      ; FIXED: match -> skip next
[2] ret  KILL_PROCESS                 ; wrong arch dies
[3] ld   [0]                          ; A = seccomp_data.nr       (full path)
[4..] jeq <allowed syscall>, jt, 0    ; allow-list hits jump over deny tail
[...] ret  KILL                        ; deny-by-default tail
[...] ret  ALLOW                       ; (allow tail after deny)

Key subtleties handled: the allow-list uses jt = nallow - i to hop over the deny tail onto RET ALLOW (the deny tail must come before RET ALLOW in program order), and SECCOMP_RET_KILL/ALLOW/KILL_PROCESS are the real kernel action values. A -DBUGGY_ARCH_JEQ switch is kept so the pre-fix code can be rebuilt to prove test coverage. Files: seccomp_filter.{c,h}, test_seccomp.c, probe_program.c, Makefile in ~/seccomp-jump-offsets.

Evidence & signatures

Verification used two layers:

**1. User-space BPF simulator** (implements `LD|W|ABS`, `JEQ|K`, `RET|K` with exact kernel jump semantics) run against synthetic `seccomp_data` for both paths, both arch cases, plus an allow/deny policy sanity check. **2. Live kernel probe** — a forked child sets `PR_SET_NO_NEW_PRIVS`, installs the real filter via `PR_SET_SECCOMP`, then makes a real `getpid()` syscall.

Fixed build (`make run`, under `setpriv --no-new-privs`) — **8/8 passed**:

| Test | Result |
|---|---|
| noop path: native/getpid → ALLOW | PASS (real arch survives) |
| full path: native/getpid → ALLOW | PASS |
| full path: native/reboot → KILL (deny-by-default intact) | PASS |
| noop path: i386/getpid → KILL_PROCESS | PASS (wrong arch dies) |
| full path: i386/getpid → KILL_PROCESS | PASS |
| full path: aarch64/read → KILL_PROCESS | PASS |
| live: no-op path, real arch survives | PASS |
| live: full path, real arch survives | PASS |

Pre-fix build (`make buggy`) — **0/8, every test fails**, reproducing the exact reported symptom: native/getpid evaluates to `0x80000000` (KILL_PROCESS), foreign arch evaluates to `0x7fff0000` (ALLOW, survives). Standalone probe under `setpriv --no-new-privs`:

```
$ setpriv --no-new-privs ./probe_buggy
Bad system call (core dumped)          # killed by SIGSYS
$ echo $?
159                                     # 128 + SIGSYS(31) — matches report
$ setpriv --no-new-privs ./probe_fixed
$ echo $?
0
```

Edge cases tested: both filter paths (full + no-op), matching vs. two foreign arches (i386, aarch64), allowed vs. denied syscall on native arch, and jump-target alignment of every allow-list `jt` (off-by-one here would land on the deny tail — caught during development, fixed, and covered by the `native/reboot → KILL` check). A live i386 wrong-arch probe wasn't possible (no `libc6-dev-i386` for `-m32` in this container); the simulator models the kernel's evaluator exactly for this instruction subset, and the live native probes prove the filter passes the kernel verifier end-to-end.
{"model": "deepseek-v4-flash", "problem_class": "seccomp-bpf-jump-offsets", "result": "passed", "tests": 4}
Generated from the verified corpus · MIT licensedBack to the catalog