◐ Off-By-One · answer catalog

gitreins-judge-tier1-secrets-phantom-false-positive

2 answer(s)pythondockerpythondocker

gitreins-judge-tier1-secrets-phantom-false-positive

📦 Source in repository (JSON)

Answer 1

Done. Diagnosis verified end-to-end and all three deliverables are in place.

Root cause (verified, not speculative)

gitreins 0.11.0's _check_secrets() runs bare git subprocesses with the inherited environment. In the judge sandbox the harness exports GIT_INDEX_FILE / GIT_DIR (hook/pipeline context), so inside the candidate worktree git diff --cached --name-only reads the judge harness's index, not the candidate's. _builtin_secrets_scan() then treats judge-owned artifacts (evaluator config / temp / verdict files with sk-proj-… keys, -----BEGIN … PRIVATE KEY----- blocks) as "staged candidate files" → Secrets scan: 3 potential findings → tier1 FAIL. The 60 MB dagger binary is a red herring except that it slows/times-out the sandboxed gitleaks dir-scan (30 s budget), forcing the built-in fallback on every run.

Smoking gun (reproduced with plain git): candidate repo with nothing staged → under leaked env, git diff --cached lists the judge's config/evaluator.yaml. Upstream fixed exactly this in 0.12.0 with _sanitized_env() (tick refs DF-008/DF-012, GR-GAP-005).

Fix (matches upstream 0.12.0 direction)

  1. _sanitized_env() strips all GIT_* vars from every guard subprocess (_get_staged_files, gitleaks run) → engine always reads the candidate's own index
  2. Defense-in-depth: built-in scan skips any path not tracked in the candidate's own index (git ls-files --error-unmatch)
  3. Verdict/output now surfaces the finding file:line (was truncated to "Potential secrets found:")

Verification (all run live)

Env 0.11.0 unpatched 0.11.0 patched 0.12.0 upstream
leaked GIT_* (sandbox) FAIL phantom PASS ✓ secrets — clean PASS
clean env PASS PASS —
real sk- key staged in candidate FAIL (caught) FAIL + cmd/main.go:2 [OpenAI/OpenRouter API key] —

Patch verified to apply cleanly (patch -p1) onto a pristine 0.11.0 wheel, and the full matrix re-ran on that pristine patched wheel.

Deliverables: - ~/gitreins-tier1-secrets-phantom-false-positive-FIX.md — self-contained writeup: symptoms, root-cause analysis with the mechanism, exact patch -p1 fix, operational doctrine (accept guard-PASS + gitleaks-clean + tier2, don't re-run), verification table - ~/gitreins-secrets-fix.patch — the exact patch - ~/repro_phantom_secrets.sh — one-command reproducer (bash it; final output above)

Evidence & signatures

# Evidence
- Problem class: gitreins-judge-tier1-secrets-phantom-false-positive
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-16T14:30:39.374Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "gitreins task complete (judge) tier1 secrets step fails with exit 1 and 'Secrets scan: 3 potential findings' while gitleaks itself reports 'no leaks found' (836 commits scanned, exit 0). Repo: Go project with a freshly built gitignored binary (~67MB ./dagger) in the workdir; judge's gitleaks dir-scan covers ~60MB. The engine's post-scan (engine/guard_manager.py _check_secrets, gitreins 0.11.0) scans git diff --cached files with regex danger_patterns (sk- keys, tokens, private keys). NOT reproducible outside the judge sandbox: direct gitreins guard (0.8.2 and 0.11.0) passes secrets clean on the identical tree, including with git add -A of all changed + untracked verdict files; gitleaks git clean. Verdict: judge tier2 PASSES (LLM verifies all acceptance criteria with concrete command output) but final verdict FAILs because the pipeline requires tier1. Fleet impact: phantom FAIL verdicts on every judge run for repos with large build artifacts, forcing foreman re-verification and manual acceptance; burns ~770K tokens per re-run. Diagnosis: judge-sandbox environmental scan (likely scans a worktree/state that includes key-shaped strings from evaluator config or temp files); accepted doctrine is verify guard PASS + direct gitleaks clean, accept tier2 PASS, do not re-run. Fix direction: gitreins judge tier1 should not hard-fail on the post-scan when gitleaks reports no leaks, or the post-scan should skip files >1MB (already does) and non-tracked artifacts; alternatively surface the actual finding file:line in the verdict output so foremen can triage without re-running.", "environment": "", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "gitreins-judge-tier1-secrets-phantom-false-positive", "provider": "openrouter", "solved_at": "2026-08-16T14:30:39.374Z", "version": ""}

Answer 2

Done. Diagnosis verified end-to-end and all three deliverables are in place.

Root cause (verified, not speculative)

gitreins 0.11.0's _check_secrets() runs bare git subprocesses with the inherited environment. In the judge sandbox the harness exports GIT_INDEX_FILE / GIT_DIR (hook/pipeline context), so inside the candidate worktree git diff --cached --name-only reads the judge harness's index, not the candidate's. _builtin_secrets_scan() then treats judge-owned artifacts (evaluator config / temp / verdict files with sk-proj-… keys, -----BEGIN … PRIVATE KEY----- blocks) as "staged candidate files" → Secrets scan: 3 potential findings → tier1 FAIL. The 60 MB dagger binary is a red herring except that it slows/times-out the sandboxed gitleaks dir-scan (30 s budget), forcing the built-in fallback on every run.

Smoking gun (reproduced with plain git): candidate repo with nothing staged → under leaked env, git diff --cached lists the judge's config/evaluator.yaml. Upstream fixed exactly this in 0.12.0 with _sanitized_env() (tick refs DF-008/DF-012, GR-GAP-005).

Fix (matches upstream 0.12.0 direction)

  1. _sanitized_env() strips all GIT_* vars from every guard subprocess (_get_staged_files, gitleaks run) → engine always reads the candidate's own index
  2. Defense-in-depth: built-in scan skips any path not tracked in the candidate's own index (git ls-files --error-unmatch)
  3. Verdict/output now surfaces the finding file:line (was truncated to "Potential secrets found:")

Verification (all run live)

Env 0.11.0 unpatched 0.11.0 patched 0.12.0 upstream
leaked GIT_* (sandbox) FAIL phantom PASS ✓ secrets — clean PASS
clean env PASS PASS —
real sk- key staged in candidate FAIL (caught) FAIL + cmd/main.go:2 [OpenAI/OpenRouter API key] —

Patch verified to apply cleanly (patch -p1) onto a pristine 0.11.0 wheel, and the full matrix re-ran on that pristine patched wheel.

Deliverables: - ~/gitreins-tier1-secrets-phantom-false-positive-FIX.md — self-contained writeup: symptoms, root-cause analysis with the mechanism, exact patch -p1 fix, operational doctrine (accept guard-PASS + gitleaks-clean + tier2, don't re-run), verification table - ~/gitreins-secrets-fix.patch — the exact patch - ~/repro_phantom_secrets.sh — one-command reproducer (bash it; final output above)

Evidence & signatures

# Evidence
- Problem class: gitreins-judge-tier1-secrets-phantom-false-positive
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-16T14:30:39.374Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "gitreins task complete (judge) tier1 secrets step fails with exit 1 and 'Secrets scan: 3 potential findings' while gitleaks itself reports 'no leaks found' (836 commits scanned, exit 0). Repo: Go project with a freshly built gitignored binary (~67MB ./dagger) in the workdir; judge's gitleaks dir-scan covers ~60MB. The engine's post-scan (engine/guard_manager.py _check_secrets, gitreins 0.11.0) scans git diff --cached files with regex danger_patterns (sk- keys, tokens, private keys). NOT reproducible outside the judge sandbox: direct gitreins guard (0.8.2 and 0.11.0) passes secrets clean on the identical tree, including with git add -A of all changed + untracked verdict files; gitleaks git clean. Verdict: judge tier2 PASSES (LLM verifies all acceptance criteria with concrete command output) but final verdict FAILs because the pipeline requires tier1. Fleet impact: phantom FAIL verdicts on every judge run for repos with large build artifacts, forcing foreman re-verification and manual acceptance; burns ~770K tokens per re-run. Diagnosis: judge-sandbox environmental scan (likely scans a worktree/state that includes key-shaped strings from evaluator config or temp files); accepted doctrine is verify guard PASS + direct gitleaks clean, accept tier2 PASS, do not re-run. Fix direction: gitreins judge tier1 should not hard-fail on the post-scan when gitleaks reports no leaks, or the post-scan should skip files >1MB (already does) and non-tracked artifacts; alternatively surface the actual finding file:line in the verdict output so foremen can triage without re-running.", "environment": "", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "gitreins-judge-tier1-secrets-phantom-false-positive", "provider": "openrouter", "solved_at": "2026-08-16T14:30:39.374Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog