◐ Off-By-One · answer catalog

bash-flock-cron-overlap-state-corruption

2 answer(s)shellbashshellbash

bash-flock-cron-overlap-state-corruption

📦 Source in repository (JSON)

Answer 1

Done. The solution is verified and delivered at ~/lockfix/SOLUTION.md, with a broken reproduction and a fixed script alongside.

Summary

Root cause — two independent bugs combine to corrupt the manifest:

  1. Subshell fd lifetime (Bug 1): flock locks a file descriptor. When fd 9 is opened inside a subshell, every fd closes when that subshell exits — so the lock is released before the critical section in the parent runs. A second overlapping cron run then acquires the lock "successfully" too. (Empirically proven: rc=0 when re-acquiring after the subshell exits.)

  2. Masked exit code (Bug 2): flock -n 9 | cat is a pipeline, whose status is the last command's (cat), not flock's. flock's busy exit code 1 is swallowed, so the fallback believes the lock was acquired and runs the critical section unlocked. (Proven: piped rc=0, direct rc=1 while another process holds the lock.)

The fix (fixed_backup.sh): - Open fd 9 once in the parent with exec 9>>"$LOCKFILE" — no subshell parentheses — so the lock is held for the entire critical section and released only afterwards (exec 9>&-). - Test flock -n 9 directly inside if ! flock -n 9; then skip (exit 0); fi, so a busy lock takes the clean-skip path instead of running unlocked.

Verification highlights: 3 overlapping fixed runs produced one runner (pid 133) with a contiguous 5-row block and two clean skip (exit 0) outcomes; the buggy script produced 10 interleaved rows from two concurrent writers. Sequential runs both succeed, confirming the lock releases after each run. Skip semantics exit 0, so cron doesn't page on expected overlap.

Evidence & signatures

# Evidence
- Problem class: bash-flock-cron-overlap-state-corruption
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-18T16:36:41.249Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A cron-deployed backup script guards its critical section with flock -n 9 on an fd opened inside a subshell, so the lock fd is closed when that subshell exits and a second overlapping cron run proceeds concurrently, corrupting the shared manifest file. The script also pipes flock -n through a pipeline, so its exit status 1 (lock busy) is masked and the fallback path treats a failed acquisition as success and runs unlocked. Fix so the lock is held for the entire critical section in the parent shell and the busy exit code is checked explicitly so overlapping runs skip cleanly instead of corrupting state.", "environment": "bash", "language": "shell", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "bash-flock-cron-overlap-state-corruption", "provider": "openrouter", "solved_at": "2026-08-18T16:36:41.249Z", "version": "5.2"}

Answer 2

Done. The solution is verified and delivered at ~/lockfix/SOLUTION.md, with a broken reproduction and a fixed script alongside.

Summary

Root cause — two independent bugs combine to corrupt the manifest:

  1. Subshell fd lifetime (Bug 1): flock locks a file descriptor. When fd 9 is opened inside a subshell, every fd closes when that subshell exits — so the lock is released before the critical section in the parent runs. A second overlapping cron run then acquires the lock "successfully" too. (Empirically proven: rc=0 when re-acquiring after the subshell exits.)

  2. Masked exit code (Bug 2): flock -n 9 | cat is a pipeline, whose status is the last command's (cat), not flock's. flock's busy exit code 1 is swallowed, so the fallback believes the lock was acquired and runs the critical section unlocked. (Proven: piped rc=0, direct rc=1 while another process holds the lock.)

The fix (fixed_backup.sh): - Open fd 9 once in the parent with exec 9>>"$LOCKFILE" — no subshell parentheses — so the lock is held for the entire critical section and released only afterwards (exec 9>&-). - Test flock -n 9 directly inside if ! flock -n 9; then skip (exit 0); fi, so a busy lock takes the clean-skip path instead of running unlocked.

Verification highlights: 3 overlapping fixed runs produced one runner (pid 133) with a contiguous 5-row block and two clean skip (exit 0) outcomes; the buggy script produced 10 interleaved rows from two concurrent writers. Sequential runs both succeed, confirming the lock releases after each run. Skip semantics exit 0, so cron doesn't page on expected overlap.

Evidence & signatures

# Evidence
- Problem class: bash-flock-cron-overlap-state-corruption
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-18T16:36:41.249Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A cron-deployed backup script guards its critical section with flock -n 9 on an fd opened inside a subshell, so the lock fd is closed when that subshell exits and a second overlapping cron run proceeds concurrently, corrupting the shared manifest file. The script also pipes flock -n through a pipeline, so its exit status 1 (lock busy) is masked and the fallback path treats a failed acquisition as success and runs unlocked. Fix so the lock is held for the entire critical section in the parent shell and the busy exit code is checked explicitly so overlapping runs skip cleanly instead of corrupting state.", "environment": "bash", "language": "shell", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "bash-flock-cron-overlap-state-corruption", "provider": "openrouter", "solved_at": "2026-08-18T16:36:41.249Z", "version": "5.2"}
Generated from the verified corpus · MIT licensedBack to the catalog