◐ Off-By-One · answer catalog

sh-flock-fd-inheritance-hang

1 answer(s)shellbash

runprotectedjob # may spawn sleep 300 &, daemons, etc.

📦 Source in repository (JSON)

Answer

Root cause. flock(2) locks are attached to the open file description, not to the guarding process. A descriptor opened by exec 9>"$LOCK" carries no FD_CLOEXEC, so every &-backgrounded child (and its descendants) inherits fd 9 through fork. The kernel therefore keeps the lock held until every inherited descriptor is closed — long after the guard script exits. The next cron invocation then blocks (or, with flock -n, spuriously reports "already running").

Fix principle. Guarantee that no process outliving the job can hold the lock descriptor. The only holder must be the guard itself, so the lock's lifetime becomes exactly the guard's lifetime. In pure bash this is done by running the job in a subshell that closes the inherited fd before any child can be spawned:

#!/bin/bash
LOCK=/var/lock/cron-job.lock

exec 9>"$LOCK" || exit 1                      # open lock file on fd 9
flock -n 9 || { echo "job already running, skipping"; exec 9>&-; exit 0; }
trap 'exec 9>&-' EXIT                          # belt-and-suspenders release

# Run the protected job in a subshell whose copy of fd 9 is closed first.
# Any background children, daemons, or grandchildren the job spawns are
# forked from that subshell, so they can never inherit the lock descriptor.
(
    exec 9>&-
    run_protected_job                          # may spawn sleep 300 &, daemons, etc.
) &
wait                                           # keep the lock until the job completes

Why it works: the subshell's exec 9>&- closes only its own fd 9 copy (the parent's stays open, so the lock is held throughout). The job then forks children from a process with fd 9 already closed, so the lock is released the moment the guard exits — wait returns, the EXIT trap runs, fd 9 closes, kernel drops the lock. This is also SIGKILL-safe: even a kill -9 of the guard releases the lock, because no other process holds the descriptor.

Alternatives (verified):

Evidence & signatures

All tests run under `bash 5.3.9` + `flock from util-linux 2.41.3`. Job = 1 s foreground work + an 8 s background child designed to outlive the guard.

| # | Test | Result |
|---|------|--------|
| 1 | Buggy guard exits, background child's `/proc/<pid>/fd/9` still → `sflock.test` | **fd held** (bug confirmed) |
| 2 | Buggy guard, second invocation (`timeout 2`) | **hung, exit 124** (bug confirmed) |
| 3 | Fixed guard exits, child's fd table contains no lock file | **no fd** |
| 4 | Fixed guard, rerun while old 8 s child still alive | **exit 0, elapsed 1.11 s** (immediate, full job ran) |
| 5 | Overlap protection intact: `-n` run while a job is genuinely active | **`ALREADY_RUNNING`, exit 0** |
| 6 | Grandchild (child of the job's child) inherits nothing | **no fd leak** |
| 7 | Daemonized child (`setsid nohup … &`); rerun while daemon alive | **no fd leak; rerun exit 0, 1.11 s** |
| 8 | `flock -u 9` variant: child still holds fd 9 but rerun proceeds | **immediate (1.11 s), works** |
| 9 | `flock -c '… & …'` variant | **leaks fd 3 to child — documented as unsafe here** |
| 10 | Independent probe (`flock -n lock -c …` from a third process) after fixed guard exits | **`LOCK_ACQUIRED_BY_PROBE`** (lock truly free) |
| 11 | Guard `kill -9`'d mid-job | **`LOCK_FREE_AFTER_SIGKILL`** |

No stray test children left behind; lock files cleaned up. Edge cases covered: background child outliving the guard, grandchild chains, setsid/nohup daemons, overlapping invocations, non-blocking vs. blocking acquisition, explicit-unlock variant, `flock -c` version caveat, and guard killed with SIGKILL.
{"model": "deepseek-v4-flash", "problem_class": "sh-flock-fd-inheritance-hang", "result": "passed", "tests": 11}
Generated from the verified corpus · MIT licensedBack to the catalog