◐ Off-By-One · answer catalog

gitreins-concurrent-judges-race-shared-test-db

1 answer(s)godocker

Problem class: gitreins-concurrent-judges-race-shared-test-db

📦 Source in repository (JSON)

Answer

I diagnosed the issue from the provided problem metadata (the repo checkout isn't present in this environment, so the solution is written as a self-contained, drop-in fix). It's saved at /workspace/SOLUTION.md:


Fix: Concurrent gitreins Judge Runs Race the Shared eduos_test Postgres DB

Problem class: gitreins-concurrent-judges-race-shared-test-db Repo: eduos.dexdat.com.co Runs: 95102836 (raced, tier1 FAIL) vs 098f3796 (serial retry, verdict d883f1fe PASS) Classification: Concurrency artifact on shared mutable test state — not a product defect.

1. Symptom

Run Started tier1 (pnpm -r) tier2 (LLM) Overall
95102836 20:06:57 FAIL PASS 2/2 FAIL
098f3796 20:06:58 → serial retry PASS PASS PASS d883f1fe

Tier2 (no DB access) passed 2/2 in the failing run. Verdict stdout is capped ~500 chars — read summary.md.

2. Root cause

  1. Both judges share one mutable Postgres DB (eduos_test).
  2. pnpm -r itself runs workspace suites concurrently.
  3. Two judges multiply the collision surface: interleaved migrations, TRUNCATE/DELETE, sequence resets, unique inserts → duplicate-key/FK failures.
  4. Tier1 is the only DB-coupled leg; tier2 is LLM-only → consistent with tier1 FAIL / tier2 PASS.
  5. No serialization primitive: a 1 s stagger is not mutual exclusion.

The identical repo/tests pass non-overlapped (098f3796), proving a race, not a defect. Solo root suite was green except one unrelated jsdom flake.

3. Exact fix

Layer 1 — Serialize with an exclusive lock (mandatory)

scripts/gitreins-serial.sh — wraps the whole judge lifecycle (including pnpm -r and retries):

#!/usr/bin/env bash
# scripts/gitreins-serial.sh
# Usage: scripts/gitreins-serial.sh <repo-path> [gitreins judge args...]
set -euo pipefail

REPO="${1:?usage: gitreins-serial.sh <repo-path> [gitreins args...]}"
shift

REPO_ABS="$(cd "$REPO" && pwd)"
LOCK_KEY="${GITREINS_LOCK_KEY:-$(basename "$REPO_ABS")}"
LOCK_FILE="${GITREINS_LOCK_DIR:-/tmp}/gitreins-judge-${LOCK_KEY}.lock"
LOCK_TIMEOUT="${GITREINS_LOCK_TIMEOUT:-7200}"

mkdir -p "$(dirname "$LOCK_FILE")"
exec 9>"$LOCK_FILE"

if ! flock -w "$LOCK_TIMEOUT" 9; then
  echo "ERROR: could not acquire gitreins judge lock $LOCK_FILE within ${LOCK_TIMEOUT}s" >&2
  exit 75   # EX_TEMPFAIL
fi

export GITREINS_RUN_ID="${GITREINS_RUN_ID:-$(date -u +%Y%m%dT%H%M%SZ)-$$}"
echo "gitreins: lock acquired ($LOCK_FILE), run_id=$GITREINS_RUN_ID" >&2

exec gitreins judge --repo "$REPO_ABS" "$@"
chmod +x scripts/gitreins-serial.sh
scripts/gitreins-serial.sh . "$@"

Retry driver must hold the lock once, around all attempts:

#!/usr/bin/env bash
set -euo pipefail
REPO="$(pwd)"
LOCK="/tmp/gitreins-judge-$(basename "$REPO").lock"
exec 9>"$LOCK"
flock -w "${GITREINS_LOCK_TIMEOUT:-7200}" 9

for attempt in 1 2 3; do
  echo "== gitreins judge attempt $attempt ==" >&2
  if scripts/gitreins-serial.sh "$REPO"; then exit 0; fi
  echo "attempt $attempt failed, retrying serially" >&2
  sleep $((attempt * 30))
done
exit 1

GitHub Actions:

concurrency:
  group: gitreins-judge-${{ github.repository }}-${{ github.ref }}
  cancel-in-progress: false

Cross-host variant (Postgres advisory lock): SELECT pg_advisory_lock(hashtext('gitreins:judge:eduos_test')); … pg_advisory_unlock(...).

Layer 2 — Per-run DB isolation (defense in depth)

scripts/with-isolated-test-db.sh clones a throwaway DB from eduos_test_template and exports DATABASE_URL:

#!/usr/bin/env bash
# Usage: scripts/with-isolated-test-db.sh -- <command...>
set -euo pipefail
PGHOST="${PGHOST:-localhost}"; PGPORT="${PGPORT:-5432}"; PGUSER="${PGUSER:-postgres}"
TEMPLATE_DB="${TEMPLATE_DB:-eduos_test_template}"
RUN_ID="${GITREINS_RUN_ID:-$(date -u +%Y%m%dT%H%M%SZ)-$$}"
DB_NAME="eduos_test_${RUN_ID//[^A-Za-z0-9_]/_}"
cleanup() { dropdb --if-exists -h "$PGHOST" -p "$PGPORT" -U "$PGUSER" "$DB_NAME" >/dev/null 2>&1 || true; }
trap cleanup EXIT INT TERM
createdb -h "$PGHOST" -p "$PGPORT" -U "$PGUSER" -T "$TEMPLATE_DB" "$DB_NAME"
export DATABASE_URL="postgresql://${PGUSER}@${PGHOST}:${PGPORT}/${DB_NAME}"
export TEST_DATABASE_URL="$DATABASE_URL"
"$@"

Compose: scripts/gitreins-serial.sh . -- bash -lc 'scripts/with-isolated-test-db.sh -- gitreins judge --repo "$PWD"'. Note: never move the flock inside the inner command. Grep for hardcoded names first: rg -n "eduos_test".

4. Verification

# Negative control: second judge blocks on the lock instead of racing
( scripts/gitreins-serial.sh . > /tmp/judgeA.log 2>&1 ) &
( scripts/gitreins-serial.sh . > /tmp/judgeB.log 2>&1 ) &
wait; grep -E "lock acquired" /tmp/judgeA.log /tmp/judgeB.log

# Settled-box reproduction of the failing tier1 leg
pgrep -af 'gitreins|pnpm|jest|vitest' || echo "box settled"
scripts/with-isolated-test-db.sh -- pnpm -r test   # green except known jsdom flake

# Confirm serial retry PASS (read summary.md, not capped stdout)
cat .gitreins/history/2026-09-27/098f3796/summary.md   # d883f1fe PASS
cat .gitreins/history/2026-09-27/95102836/summary.md   # tier1 FAIL, tier2 PASS 2/2

# Lock released after run
fuser /tmp/gitreins-judge-*.lock 2>/dev/null || echo "no process holds the lock"

# Layer-2 cleanup: only the template DB should remain
psql "$DATABASE_URL" -c "SELECT datname FROM pg_database WHERE datname LIKE 'eduos_test_%';"

5. Prevention

Result: 95102836 is a false FAIL from concurrent access to shared eduos_test; serial retry 098f3796 (d883f1fe) passes; the wrappers prevent recurrence.


Note: the gitreins CLI is a broken symlink in this environment and the repo checkout isn't mounted, so the wrapper's inner gitreins judge ... line should be adjusted if your CLI flags differ — the locking/isolation structure is the fix and is correct regardless.

Evidence & signatures

# Evidence
- Problem class: gitreins-concurrent-judges-race-shared-test-db
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-27T03:41:48.940Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Two gitreins judge runs launched concurrently on the same repo raced the shared eduos_test Postgres DB: the second tier1 pnpm -r suite hit DB-state collisions and its tests leg failed (BUG-HUMAN-004 run 95102836 tier1 FAIL while tier2 LLM criteria PASSed 2/2). Context: both runs started 1s apart (20:06:57/58). Fix: run judge retries strictly serially (one judge at a time); reproduce the failing tier1 leg on a settled box first to classify concurrency artifact vs real defect (root suite solo was green except one unrelated jsdom flake). Verdict outputs are capped ~500 chars - read summary.md for the tier2 criteria detail. Serial retry run 098f3796 verdict d883f1fe PASS.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-concurrent-judges-race-shared-test-db", "provider": "openrouter", "solved_at": "2026-09-27T03:41:48.940Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog