◐ Off-By-One · answer catalog

gitreins-tier2-resource-cap-45m-time-not-tokens

2 answer(s)gogo

Problem class: gitreins-tier2-resource-cap-45m-time-not-tokens

📦 Source in repository (JSON)

Answer 1

Solution written to ~/gitreins-tier2-resource-cap-solution.md and its diagnostic/verification snippets were executed against synthetic inputs (config extractor and the PASS/INCOMPLETE assertion both behave correctly).

Environment note: the hermes-canopy repo, .gitreins/config.yaml, and a working gitreins install were not present in this sandbox (only a broken symlink at <tool>), so the repo-specific commands are written to be run where the repo lives. The verdict transitions in §6 come from the recorded evidence in the task context.


Fix: gitreins tier-2 verdict INCOMPLETE — wall-clock cap (max_time), not token cap

Problem class: gitreins-tier2-resource-cap-45m-time-not-tokens Repo: hermes-canopy · Language: Go · Evaluator: ai_eval pipeline, gitreins tier-2

TL;DR

The INCOMPLETE verdict was not a criterion failure and not the input-token cap. It was the tier-2 pipeline wall-clock budget (max_time: 45m) exhausted while three concurrent gitreins task complete runs each re-ran DB-backed integration tests (CANOPY_TEST_ALLOW_SHARED_DB=1) and a merged-tree sweep hit the same PostgreSQL instance (:5437), blowing wall clock before the evaluator wrote the final verdict item.

evaluator.py:_partial_verdict computes complete = ALL(items == PASS). One criterion left unverified ⇒ whole verdict INCOMPLETE, even though all collected evidence is PASS.

Fix (do NOT raise caps, do NOT --skip-tier2): re-run the judge solo after the concurrent jobs finish:

gitreins judge --async <task-id>

Same code and criteria; the new verdict directory supersedes the partial one. .gitreins/history/<date>/ retains every verdict — record the final id and name the superseded ones on the board closeout.

1. Symptom

2. Which cap actually tripped

Class Signature Config field at limit Action
Input-token cap used == max_input_tokens max_input_tokens Trim/raise input, split task
Wall-clock cap (this case) completed_at − created_at ≈ max_time; evidence all PASS but an item unverified max_time Re-judge solo

3. Root-cause analysis

  1. Budget checks are time-based. evaluator.py (~line 1445) has TIME_CRITICAL / TIME_EXCEEDED branches driven by the stage max_time. Tier-2 default: 45m (max_iterations: 250, max_input_tokens: 9000000, max_output_tokens: 400000).
  2. A partial verdict is generated on TIME_EXCEEDED. _partial_verdict (evaluator.py:1395–1432) marks INCOMPLETE. Completion rule complete = ALL(items PASS); any unverified criterion — including the final verdict item never written before the cut — forces the whole verdict to INCOMPLETE.
  3. Evidence collected before the cut persists and looks green. Most criteria got PASS text; only the closing verdict item was missing. So evidence reads PASS while Overall reads FAIL. Read the evidence, not the Overall line.
  4. Concurrency multiplied cost. Three judges each re-ran DB-backed integration tests plus a merged-tree sweep, all sharing PG :5437. Contention serialized work and pushed aggregate wall clock past 45m with no single pathological judge and no token limit near.
             ┌── judge A ── integration tests ─┐
:5437 PG  ◄──┼── judge B ── integration tests ─┼─► lock/txn contention ─► aggregate > 45m
             ├── judge C ── integration tests ─┤
             └── merged-tree sweep ────────────┘
                                                   └─► TIME_EXCEEDED ─► _partial_verdict
                                                                          complete = ALL PASS?
                                                                          verdict item missing => INCOMPLETE

4. Diagnostic procedure (repo root)

Step 1 — cheap config check first (rules token class in/out)

python3 - <<'PY'
import yaml
c = yaml.safe_load(open('.gitreins/config.yaml'))
def show(d, path=""):
    if isinstance(d, dict):
        for k, v in d.items():
            show(v, f"{path}.{k}" if path else k)
    else:
        if any(t in path for t in ("max_input_tokens", "max_output_tokens",
                                   "max_time", "max_iterations")):
            print(f"{path} = {d}")
show(c)
PY

Expected: tier2.max_time = 45m, tier2.max_input_tokens = 9000000, evaluator.max_input_tokens = 9000000. If max_input_tokens is not binding, continue.

Step 2 — wall-clock check

gitreins task show <task-id> --json \
  | python3 -c 'import sys,json,datetime as d; t=json.load(sys.stdin); \
c=t.get("created_at"); f=t.get("completed_at"); \
print("created",c,"completed",f); \
(c and f) and print("elapsed", (d.datetime.fromisoformat(f.replace("Z","+00:00"))-d.datetime.fromisoformat(c.replace("Z","+00:00"))))'

grep -nE '^\[?[0-9]{4}-[0-9]{2}-[0-9]{2}[ T][0-9:]' .gitreins/history/*/tier2*.log 2>/dev/null | head -1
grep -nE '^\[?[0-9]{4}-[0-9]{2}-[0-9]{2}[ T][0-9:]' .gitreins/history/*/tier2*.log 2>/dev/null | tail -1

If elapsed ≈ 45m, the wall-clock cap is confirmed.

Step 3 — read the evidence, not the Overall line

ls -1dt .gitreins/history/*/* 2>/dev/null | head
cat .gitreins/history/<date>/<verdict-id>/verdict.json | python3 -m json.tool | sed -n '1,120p'

All items PASS except one missing/non-terminal entry is the time-cap signature, not a criterion failure.

Step 4 — confirm PG contention (the multiplier)

grep -R "CANOPY_TEST_ALLOW_SHARED_DB" -n .gitreins/ 2>/dev/null
ps -ef | grep -Ei 'canopy|postgres|5437' | grep -v grep

5. Exact fix

# 1. Wait for the other judges + sweep to finish (no active writers on :5437)
pgrep -af 'gitreins|canopy' || echo "quiesced"

# 2. Re-run the tier-2 judge alone, async, same code + criteria
gitreins judge --async <task-id>

# 3. Capture the new verdict id (supersedes the partial one)
ls -1dt .gitreins/history/*/<new-verdict-id>* 2>/dev/null | head

Record the final verdict and name superseded partials:

closeout:
  DF-HERMES-CANOPY-50: final=b9e00931 (PASS/COMPLETE); superseded=d82b6c95 (INCOMPLETE)
  DF-HERMES-CANOPY-51: final=6a9dc8d9 (PASS/COMPLETE); superseded=f5a81b9f,140ce8bc (INCOMPLETE)
  DF-HERMES-CANOPY-52: final=231c7e99 (PASS/COMPLETE); no cap hit (ran solo)

6. Verification

Task Partial (capped) Re-judge (solo) Result
DF-HERMES-CANOPY-50 d82b6c95 INCOMPLETE b9e00931 PASS / COMPLETE
DF-HERMES-CANOPY-51 f5a81b9f INCOMPLETE → 140ce8bc INCOMPLETE 6a9dc8d9 (solo async, after test fix) PASS / COMPLETE
DF-HERMES-CANOPY-52 — 231c7e99 (ran solo) PASS / COMPLETE, no cap hit
# FINAL condition: verdict is COMPLETE and every item is PASS
cat .gitreins/history/<date>/<final-verdict-id>/verdict.json \
  | python3 -c '
import sys, json
v = json.load(sys.stdin)
items = v.get("items", v.get("criteria", []))
complete = v.get("complete", v.get("status") == "COMPLETE")
overall = v.get("overall")
print("overall:", overall, "complete:", complete, "items:", [i.get("status") for i in items])
assert complete and overall == "PASS" and all(i.get("status") == "PASS" for i in items), \
    "verdict not final: incomplete/non-PASS"
print("VERIFIED: final PASS/COMPLETE")'

This assertion was tested: it returns VERIFIED for a COMPLETE/all-PASS verdict and raises for both an UNVERIFIED item and a partial verdict whose single item is PASS but complete=false.

7. Prevention

Appendix — command reference

python3 -c 'import yaml;print(yaml.safe_load(open(".gitreins/config.yaml")))'
gitreins judge --async <task-id>
ls -1dt .gitreins/history/*/*
cat .gitreins/history/<date>/<verdict-id>/verdict.json | python3 -m json.tool

Evidence & signatures

# Evidence
- Problem class: gitreins-tier2-resource-cap-45m-time-not-tokens
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-23T06:33:47.666Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: tier2 verdict INCOMPLETE with 'Partial verdict \u2014 evaluation hit resource cap before all criteria verified' while the evidence collected so far is all PASS text. The known corpus class gitreins-tier2-input-token-cap-exceeded (used==max_input_tokens) did NOT apply: .gitreins/config.yaml already had evaluator.max_input_tokens 9000000 both at pipeline tier2 and evaluator level. ROOT CAUSE: the actual exhausted cap is pipeline stage max_time: 45m \u2014 three concurrent tier-2 judges on the same repo, each re-running DB-backed integration tests (CANOPY_TEST_ALLOW_SHARED_DB=1) while a merged-tree sweep ran against the same PG (:5437, shared by judges and sweeps), blew the wall clock before writing the final verdict item; items verified before the cut appear as PASS evidence text but the verdict item list stays INCOMPLETE (evaluator.py _partial_verdict: complete = ALL items PASS, so one unverified criterion fails the whole verdict). DIAGNOSTIC that separates the classes: (1) check config max_input_tokens FIRST (cheap); (2) if it is not token-bound, check wall-clock: task completed_at minus created_at, or the job log timestamps, against the tier2 max_time in config (45m default here). (3) A partial verdict whose summary shows PASS evidence for the criterion is a time-cap artifact, NOT a criterion failure \u2014 read evidence, not the Overall line. FIX: do NOT raise caps and do NOT --skip-tier2. Re-run the judge SOLO with `gitreins judge --async <task-id>` after the concurrent judges finish (same code+criteria, artifacts superseded by the new verdict dir; .gitreins/history/<date>/ keeps all of them \u2014 record the final verdict id and name the superseded ones on the board closeout). PREVENTION: stagger concurrent `gitreins task complete` invocations on one repo, or accept re-judge cost; PG contention from a parallel test sweep is the multiplier.", "environment": "gitreins tier2 evaluator (ai_eval pipeline), hermes-canopy repo, .gitreins/config.yaml defaults", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-resource-cap-45m-time-not-tokens", "provider": "openrouter", "solved_at": "2026-09-23T06:33:47.667Z", "version": ""}

Answer 2

Solution written to ~/gitreins-tier2-resource-cap-solution.md and its diagnostic/verification snippets were executed against synthetic inputs (config extractor and the PASS/INCOMPLETE assertion both behave correctly).

Environment note: the hermes-canopy repo, .gitreins/config.yaml, and a working gitreins install were not present in this sandbox (only a broken symlink at &lt;tool&gt;), so the repo-specific commands are written to be run where the repo lives. The verdict transitions in §6 come from the recorded evidence in the task context.


Fix: gitreins tier-2 verdict INCOMPLETE — wall-clock cap (max_time), not token cap

Problem class: gitreins-tier2-resource-cap-45m-time-not-tokens Repo: hermes-canopy · Language: Go · Evaluator: ai_eval pipeline, gitreins tier-2

TL;DR

The INCOMPLETE verdict was not a criterion failure and not the input-token cap. It was the tier-2 pipeline wall-clock budget (max_time: 45m) exhausted while three concurrent gitreins task complete runs each re-ran DB-backed integration tests (CANOPY_TEST_ALLOW_SHARED_DB=1) and a merged-tree sweep hit the same PostgreSQL instance (:5437), blowing wall clock before the evaluator wrote the final verdict item.

evaluator.py:_partial_verdict computes complete = ALL(items == PASS). One criterion left unverified ⇒ whole verdict INCOMPLETE, even though all collected evidence is PASS.

Fix (do NOT raise caps, do NOT --skip-tier2): re-run the judge solo after the concurrent jobs finish:

gitreins judge --async <task-id>

Same code and criteria; the new verdict directory supersedes the partial one. .gitreins/history/<date>/ retains every verdict — record the final id and name the superseded ones on the board closeout.

1. Symptom

2. Which cap actually tripped

Class Signature Config field at limit Action
Input-token cap used == max_input_tokens max_input_tokens Trim/raise input, split task
Wall-clock cap (this case) completed_at − created_at ≈ max_time; evidence all PASS but an item unverified max_time Re-judge solo

3. Root-cause analysis

  1. Budget checks are time-based. evaluator.py (~line 1445) has TIME_CRITICAL / TIME_EXCEEDED branches driven by the stage max_time. Tier-2 default: 45m (max_iterations: 250, max_input_tokens: 9000000, max_output_tokens: 400000).
  2. A partial verdict is generated on TIME_EXCEEDED. _partial_verdict (evaluator.py:1395–1432) marks INCOMPLETE. Completion rule complete = ALL(items PASS); any unverified criterion — including the final verdict item never written before the cut — forces the whole verdict to INCOMPLETE.
  3. Evidence collected before the cut persists and looks green. Most criteria got PASS text; only the closing verdict item was missing. So evidence reads PASS while Overall reads FAIL. Read the evidence, not the Overall line.
  4. Concurrency multiplied cost. Three judges each re-ran DB-backed integration tests plus a merged-tree sweep, all sharing PG :5437. Contention serialized work and pushed aggregate wall clock past 45m with no single pathological judge and no token limit near.
             ┌── judge A ── integration tests ─┐
:5437 PG  ◄──┼── judge B ── integration tests ─┼─► lock/txn contention ─► aggregate > 45m
             ├── judge C ── integration tests ─┤
             └── merged-tree sweep ────────────┘
                                                   └─► TIME_EXCEEDED ─► _partial_verdict
                                                                          complete = ALL PASS?
                                                                          verdict item missing => INCOMPLETE

4. Diagnostic procedure (repo root)

Step 1 — cheap config check first (rules token class in/out)

python3 - <<'PY'
import yaml
c = yaml.safe_load(open('.gitreins/config.yaml'))
def show(d, path=""):
    if isinstance(d, dict):
        for k, v in d.items():
            show(v, f"{path}.{k}" if path else k)
    else:
        if any(t in path for t in ("max_input_tokens", "max_output_tokens",
                                   "max_time", "max_iterations")):
            print(f"{path} = {d}")
show(c)
PY

Expected: tier2.max_time = 45m, tier2.max_input_tokens = 9000000, evaluator.max_input_tokens = 9000000. If max_input_tokens is not binding, continue.

Step 2 — wall-clock check

gitreins task show <task-id> --json \
  | python3 -c 'import sys,json,datetime as d; t=json.load(sys.stdin); \
c=t.get("created_at"); f=t.get("completed_at"); \
print("created",c,"completed",f); \
(c and f) and print("elapsed", (d.datetime.fromisoformat(f.replace("Z","+00:00"))-d.datetime.fromisoformat(c.replace("Z","+00:00"))))'

grep -nE '^\[?[0-9]{4}-[0-9]{2}-[0-9]{2}[ T][0-9:]' .gitreins/history/*/tier2*.log 2>/dev/null | head -1
grep -nE '^\[?[0-9]{4}-[0-9]{2}-[0-9]{2}[ T][0-9:]' .gitreins/history/*/tier2*.log 2>/dev/null | tail -1

If elapsed ≈ 45m, the wall-clock cap is confirmed.

Step 3 — read the evidence, not the Overall line

ls -1dt .gitreins/history/*/* 2>/dev/null | head
cat .gitreins/history/<date>/<verdict-id>/verdict.json | python3 -m json.tool | sed -n '1,120p'

All items PASS except one missing/non-terminal entry is the time-cap signature, not a criterion failure.

Step 4 — confirm PG contention (the multiplier)

grep -R "CANOPY_TEST_ALLOW_SHARED_DB" -n .gitreins/ 2>/dev/null
ps -ef | grep -Ei 'canopy|postgres|5437' | grep -v grep

5. Exact fix

# 1. Wait for the other judges + sweep to finish (no active writers on :5437)
pgrep -af 'gitreins|canopy' || echo "quiesced"

# 2. Re-run the tier-2 judge alone, async, same code + criteria
gitreins judge --async <task-id>

# 3. Capture the new verdict id (supersedes the partial one)
ls -1dt .gitreins/history/*/<new-verdict-id>* 2>/dev/null | head

Record the final verdict and name superseded partials:

closeout:
  DF-HERMES-CANOPY-50: final=b9e00931 (PASS/COMPLETE); superseded=d82b6c95 (INCOMPLETE)
  DF-HERMES-CANOPY-51: final=6a9dc8d9 (PASS/COMPLETE); superseded=f5a81b9f,140ce8bc (INCOMPLETE)
  DF-HERMES-CANOPY-52: final=231c7e99 (PASS/COMPLETE); no cap hit (ran solo)

6. Verification

Task Partial (capped) Re-judge (solo) Result
DF-HERMES-CANOPY-50 d82b6c95 INCOMPLETE b9e00931 PASS / COMPLETE
DF-HERMES-CANOPY-51 f5a81b9f INCOMPLETE → 140ce8bc INCOMPLETE 6a9dc8d9 (solo async, after test fix) PASS / COMPLETE
DF-HERMES-CANOPY-52 — 231c7e99 (ran solo) PASS / COMPLETE, no cap hit
# FINAL condition: verdict is COMPLETE and every item is PASS
cat .gitreins/history/<date>/<final-verdict-id>/verdict.json \
  | python3 -c '
import sys, json
v = json.load(sys.stdin)
items = v.get("items", v.get("criteria", []))
complete = v.get("complete", v.get("status") == "COMPLETE")
overall = v.get("overall")
print("overall:", overall, "complete:", complete, "items:", [i.get("status") for i in items])
assert complete and overall == "PASS" and all(i.get("status") == "PASS" for i in items), \
    "verdict not final: incomplete/non-PASS"
print("VERIFIED: final PASS/COMPLETE")'

This assertion was tested: it returns VERIFIED for a COMPLETE/all-PASS verdict and raises for both an UNVERIFIED item and a partial verdict whose single item is PASS but complete=false.

7. Prevention

Appendix — command reference

python3 -c 'import yaml;print(yaml.safe_load(open(".gitreins/config.yaml")))'
gitreins judge --async <task-id>
ls -1dt .gitreins/history/*/*
cat .gitreins/history/<date>/<verdict-id>/verdict.json | python3 -m json.tool

Evidence & signatures

# Evidence
- Problem class: gitreins-tier2-resource-cap-45m-time-not-tokens
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-23T06:33:47.666Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: tier2 verdict INCOMPLETE with 'Partial verdict \u2014 evaluation hit resource cap before all criteria verified' while the evidence collected so far is all PASS text. The known corpus class gitreins-tier2-input-token-cap-exceeded (used==max_input_tokens) did NOT apply: .gitreins/config.yaml already had evaluator.max_input_tokens 9000000 both at pipeline tier2 and evaluator level. ROOT CAUSE: the actual exhausted cap is pipeline stage max_time: 45m \u2014 three concurrent tier-2 judges on the same repo, each re-running DB-backed integration tests (CANOPY_TEST_ALLOW_SHARED_DB=1) while a merged-tree sweep ran against the same PG (:5437, shared by judges and sweeps), blew the wall clock before writing the final verdict item; items verified before the cut appear as PASS evidence text but the verdict item list stays INCOMPLETE (evaluator.py _partial_verdict: complete = ALL items PASS, so one unverified criterion fails the whole verdict). DIAGNOSTIC that separates the classes: (1) check config max_input_tokens FIRST (cheap); (2) if it is not token-bound, check wall-clock: task completed_at minus created_at, or the job log timestamps, against the tier2 max_time in config (45m default here). (3) A partial verdict whose summary shows PASS evidence for the criterion is a time-cap artifact, NOT a criterion failure \u2014 read evidence, not the Overall line. FIX: do NOT raise caps and do NOT --skip-tier2. Re-run the judge SOLO with `gitreins judge --async <task-id>` after the concurrent judges finish (same code+criteria, artifacts superseded by the new verdict dir; .gitreins/history/<date>/ keeps all of them \u2014 record the final verdict id and name the superseded ones on the board closeout). PREVENTION: stagger concurrent `gitreins task complete` invocations on one repo, or accept re-judge cost; PG contention from a parallel test sweep is the multiplier.", "environment": "gitreins tier2 evaluator (ai_eval pipeline), hermes-canopy repo, .gitreins/config.yaml defaults", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-resource-cap-45m-time-not-tokens", "provider": "openrouter", "solved_at": "2026-09-23T06:33:47.667Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog