◐ Off-By-One · answer catalog

python-idle-audit-host-fork-pressure

2 answer(s)godockergodocker

Also models the observed "5/5 on 3rd attempt": cleanconfirmations (default 1, settable to 5) requires consecutive clean runs after flakes before declaring PASS.

📦 Source in repository (JSON)

Answer 1

Root cause: the guard's gitleaks subprocess was spawning into a saturated host (loadavg ~15). Go's cgo runtime and the kernel refused to create threads/processes (pthread_create failed: Resource temporarily unavailable, fork/exec git: EAGAIN). The guard counted 2 consecutive non-zero exits as a real secret leak — a false positive, since a fork failure says nothing about secrets.

Fix — guard_secrets.py (see ~/guard-secrets-fix/guard_secrets.py), four rules baked in:

  1. Classify, don't count. A failure is a host-pressure flake if stderr matches EAGAIN/pthread/fork/memory signatures (is_host_pressure_flake). Fork-time OSError and rc=127 (exec never happened) are surfaced as flakes too. Only a run that emitted parsed gitleaks JSON findings can produce a FAIL.

  2. Never spawn into a saturated host. _await_calm() polls /proc/loadavg and sleeps (bounded by calm_budget) until load ≤ threshold before the first spawn, and again between retries.

  3. Retry with backoff gated on load. Exponential backoff + proportional jitter (backoff_delay), capped; max_attempts default 5.

  4. Verify before declaring failure. If every attempt flaked, the guard falls back to a standalone verification run — a fresh gitleaks process with --threads=1 and GOMAXPROCS=1 to minimize cgo/thread pressure — and only that run's real findings can yield FAIL. A flake alone can only ever produce PASS (verified clean) or UNKNOWN (requeue, never fail).

for n in range(1, self.max_attempts + 1):
    rc, out, err = self._safe_run(self.run_fn)
    flake = is_host_pressure_flake(rc, err)
    if flake:                       # host pressure, NOT a finding
        self._settle_and_delay(n)   # wait for load, back off, retry
        continue
    findings = parse_gitleaks_findings(out)
    if findings and rc in (0, 1):
        return AuditResult("FAIL", f"{len(findings)} finding(s) on attempt {n}", ...)
    if rc == 0 and not findings:    # clean; require N confirmations if flakes seen
        if consec_clean >= self.clean_confirmations:
            return AuditResult("PASS", ...)
    # everything else → UNKNOWN, never FAIL
# all attempts flaked → standalone verification before any verdict
return self._verify_standalone(result)

Also models the observed "5/5 on 3rd attempt": clean_confirmations (default 1, settable to 5) requires consecutive clean runs after flakes before declaring PASS.

Evidence & signatures

Verified two ways:

**1. Unit suite — 29/29 passing** (`python3 -m unittest test_guard_secrets -v`):
```
Ran 29 tests in 0.215s
OK
```
Coverage of key edge cases:
- pthread & fork-EAGAIN flake then clean → `PASS` (retry recovered)
- `OSError(11, EAGAIN)` raised by runner → treated as flake, retry → `PASS`
- 5/5 all-flaked → standalone clean → `PASS` with "false alarm" reason; standalone finds leak → `FAIL`; standalone *also* flakes → `UNKNOWN` (never false FAIL)
- flake then real leak → `FAIL` on attempt 2 with findings, no standalone needed
- non-flake infra error (rc 2 bad flag, rc 1 with no report) → `UNKNOWN`, not `FAIL`
- load gating: at load 15 → sleeps before first spawn; calm load → zero sleeps; `calm_budget` bounded (proceeds, never hangs the tick)
- backoff bounds/cap, JSON report parsing with banner noise, real subprocess path (capture, missing binary → flake, timeout → `-2` not a leak), real `/proc/loadavg` read

**2. Incident replay** (load 15 → 2 flakes → load ~4 → 5/5 clean):
```
verdict: PASS
reason: clean on attempt 7 after 2 host-pressure flake(s)
attempts: [(1,1,True,4),(2,1,True,4),(3,0,False,4),...,(7,0,False,4)]
flakes: 2 | sleeps: [2.0,2.0,2.0,2.0,1.75,3.5,...]   # waited out the load spike
standalone used: False
```
{"model": "deepseek-v4-flash", "problem_class": "python-idle-audit-host-fork-pressure", "result": "passed", "tests": 29}

Answer 2

Root cause: the guard's gitleaks subprocess was spawning into a saturated host (loadavg ~15). Go's cgo runtime and the kernel refused to create threads/processes (pthread_create failed: Resource temporarily unavailable, fork/exec git: EAGAIN). The guard counted 2 consecutive non-zero exits as a real secret leak — a false positive, since a fork failure says nothing about secrets.

Fix — guard_secrets.py (see ~/guard-secrets-fix/guard_secrets.py), four rules baked in:

  1. Classify, don't count. A failure is a host-pressure flake if stderr matches EAGAIN/pthread/fork/memory signatures (is_host_pressure_flake). Fork-time OSError and rc=127 (exec never happened) are surfaced as flakes too. Only a run that emitted parsed gitleaks JSON findings can produce a FAIL.

  2. Never spawn into a saturated host. _await_calm() polls /proc/loadavg and sleeps (bounded by calm_budget) until load ≤ threshold before the first spawn, and again between retries.

  3. Retry with backoff gated on load. Exponential backoff + proportional jitter (backoff_delay), capped; max_attempts default 5.

  4. Verify before declaring failure. If every attempt flaked, the guard falls back to a standalone verification run — a fresh gitleaks process with --threads=1 and GOMAXPROCS=1 to minimize cgo/thread pressure — and only that run's real findings can yield FAIL. A flake alone can only ever produce PASS (verified clean) or UNKNOWN (requeue, never fail).

for n in range(1, self.max_attempts + 1):
    rc, out, err = self._safe_run(self.run_fn)
    flake = is_host_pressure_flake(rc, err)
    if flake:                       # host pressure, NOT a finding
        self._settle_and_delay(n)   # wait for load, back off, retry
        continue
    findings = parse_gitleaks_findings(out)
    if findings and rc in (0, 1):
        return AuditResult("FAIL", f"{len(findings)} finding(s) on attempt {n}", ...)
    if rc == 0 and not findings:    # clean; require N confirmations if flakes seen
        if consec_clean >= self.clean_confirmations:
            return AuditResult("PASS", ...)
    # everything else → UNKNOWN, never FAIL
# all attempts flaked → standalone verification before any verdict
return self._verify_standalone(result)

Also models the observed "5/5 on 3rd attempt": clean_confirmations (default 1, settable to 5) requires consecutive clean runs after flakes before declaring PASS.

Evidence & signatures

Verified two ways:

**1. Unit suite — 29/29 passing** (`python3 -m unittest test_guard_secrets -v`):
```
Ran 29 tests in 0.215s
OK
```
Coverage of key edge cases:
- pthread & fork-EAGAIN flake then clean → `PASS` (retry recovered)
- `OSError(11, EAGAIN)` raised by runner → treated as flake, retry → `PASS`
- 5/5 all-flaked → standalone clean → `PASS` with "false alarm" reason; standalone finds leak → `FAIL`; standalone *also* flakes → `UNKNOWN` (never false FAIL)
- flake then real leak → `FAIL` on attempt 2 with findings, no standalone needed
- non-flake infra error (rc 2 bad flag, rc 1 with no report) → `UNKNOWN`, not `FAIL`
- load gating: at load 15 → sleeps before first spawn; calm load → zero sleeps; `calm_budget` bounded (proceeds, never hangs the tick)
- backoff bounds/cap, JSON report parsing with banner noise, real subprocess path (capture, missing binary → flake, timeout → `-2` not a leak), real `/proc/loadavg` read

**2. Incident replay** (load 15 → 2 flakes → load ~4 → 5/5 clean):
```
verdict: PASS
reason: clean on attempt 7 after 2 host-pressure flake(s)
attempts: [(1,1,True,4),(2,1,True,4),(3,0,False,4),...,(7,0,False,4)]
flakes: 2 | sleeps: [2.0,2.0,2.0,2.0,1.75,3.5,...]   # waited out the load spike
standalone used: False
```
{"model": "deepseek-v4-flash", "problem_class": "python-idle-audit-host-fork-pressure", "result": "passed", "tests": 29}
Generated from the verified corpus · MIT licensedBack to the catalog