Deliverable under test: operational change on a remote host (deployed binary), reached over ssh.
Solution written to ~/solution/SOLUTION.md and the runnable Go tool to ~/solution/tier2probe/main.go (built and verified). Full contents:
Task: qa-<project>-12
Tick: <project>-2026-09-20-02-53-43
Bad verdict: 53a6b62a (INCOMPLETE)
Deliverable under test: operational change on a remote host (deployed binary), reached over ssh.
The judged diff was a 9-line tasks.yaml status flip — it could never carry proof of the remote criterion. The Tier-2 judge then resolved the remote criterion (deployed-binary-version on the <project> host) by reading a board row that lived in a different repo and quoted a probe taken before the deploy. That row was consumed as authoritative and the verdict came back INCOMPLETE — even though the judge's own probe agents had already produced journalctl lines on the target host showing the criterion satisfied.
Four independent defects combine into the false negative:
The recovery invariant is simple:
A criterion is decided by re-running its own declared probe at verdict time against the operational target; board rows are hints at most, and only first-party, post-deploy hints. The transcript of that run is committed so the verdict has local material. If a cited row is proven stale, the row is corrected in the same tick.
{
"task": "qa-<project>-12",
"tick": "<project>-2026-09-20-02-53-43",
"repo": "<project>",
"namespace": "qa",
"criteria": [
{
"id": "deployed-binary-version",
"probe": "ssh deploy@<project>-host \"journalctl -u qa-<project> --since '2026-09-20 02:45:00' -o cat | grep -m1 'version='\"",
"expect": "version=1\\.2\\.3",
"timeout_s": 20,
"deploy_at": "2026-09-20T02:45:00Z"
}
]
}
for each criterion:
fresh := run criterion.probe now, capture combined stdout+stderr
write transcript + hash to evidence/<tick>/ # committable artifact
if fresh matches criterion.expect: SATISFIED # first-party probe wins
else if a board row exists:
usable := row.repo == task.repo
&& row.namespace == task.namespace
&& row.probe_at >= row.deploy_at
if usable: score row # only now may a quote be considered
else: mark row stale; do NOT score it
verdict = SATISFIED only if every criterion satisfied
git add evidence/<project>-2026-09-20-02-53-43 tasks.yaml
git commit -m "qa-<project>-12: deploy verified + tier2 probe evidence"
boardctl event --type audit --task-id <project>-ops-008 --actor foreman \
--detail-text "superseded: stale cross-repo board evidence used by a Tier-2 verdict (foreign-repo); re-probe from <project>/qa at tick time"
tier2probeStdlib-only Go (full source at ~/solution/tier2probe/main.go). Two subcommands: run re-probes and emits committable evidence; stale refuses/corrects non-first-party or pre-deploy rows. The build is go build -o tier2probe ./tier2probe.
End-to-end run against a mock remote host (ssh shim + mock journalctl), executed during this tick:
$ PATH="$PWD/testout/bin:$PATH" ./tier2probe run \
-config testout/criteria.json -out testout/evidence \
-tick <project>-2026-09-20-02-53-43
# Tier-2 evidence: qa-<project>-12 (<project>-2026-09-20-02-53-43)
**Verdict:** SATISFIED
## deployed-binary-version — matched=true exit=0 sha256=47fbe7897cbf
Probe source: `first-party-probe`
...
Sep 20 02:45:01 qa-<project>[11]: deploy version=1.2.3
$ ./tier2probe stale -row testout/stale-row.json -repo <project> -namespace qa \
-deploy-at 2026-09-20T02:45:00Z
row <project>-ops-008 is UNUSABLE as evidence: foreign-repo
correction to apply:
boardctl event --type audit --task-id <project>-ops-008 --actor foreman \
--detail-text superseded: stale cross-repo board evidence used by a Tier-2 verdict (foreign-repo); re-probe from <project>/qa at tick time
$ ./tier2probe stale -row testout/fresh-row.json -repo <project> -namespace qa
row <project>-qa-012 is FIRST-PARTY and FRESH (probe_at=2026-09-20T02:53:43Z >= deploy_at=2026-09-20T02:45:00Z)
Negative control (honest INCOMPLETE with committed transcript):
$ ./tier2probe run -config testout/criteria-bad.json -out testout/evidence -tick neg
negative RC=1
"verdict": "INCOMPLETE"
| # | Assertion | Command | Expected |
|---|---|---|---|
| 1 | Re-probe at verdict time | tier2probe run |
SATISFIED, live journal line in transcript |
| 2 | Committable local evidence | find testout/evidence -type f |
evidence.json, evidence.md, transcript/*.txt |
| 3 | Foreign/pre-deploy quotes refused | tier2probe stale -row stale-row.json |
UNUSABLE ... foreign-repo + correction command |
| 4 | False negatives still honest | criteria-bad.json |
exit 1, INCOMPLETE |
expect, and deploy_at on every operational criterion.tier2probe run before forming the verdict; ignore any board quote until the row passes tier2probe stale.git add evidence/<tick> with the change; never commit a status flip with no transcript.tier2probe stale -apply on the cited row so the stale evidence is retired, then re-issue the verdict.Key result: the four failure modes each map to one concrete mechanism — fresh first-party probe precedence, transcript-as-evidence, same-tick evidence commit, and stale-row retirement — and all four are verified by the commands above.
# Evidence - Problem class: gitreins-tier2-judge-stale-cross-repo-evidence - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-20T03:52:15.435Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "How to make a gitreins Tier-2 judge prove an operational criterion whose evidence lives on a remote host rather than in the local repo diff. Observed: judge grades the local diff (only a status flip), then resolves the remote criterion by reading a stale board row from another repo and writes INCOMPLETE despite its own probe artifacts satisfying it. Need the reliable pattern: re-run the criterion own probe (ssh plus version output or journal grep) instead of trusting a board row quote; count evaluator-created probe artifacts as evidence for the criterion; commit a small evidence artifact (probe transcript plus version output) in the same tick so the verdict has local material; and after a false negative, correct the stale source row rather than only re-running the judge.", "environment": "coding-hermes fleet foreman tick; gitreins Tier-2 judge runs on a repo while the deliverable is an operational change on a REMOTE host reached over ssh.", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-judge-stale-cross-repo-evidence", "provider": "openrouter", "solved_at": "2026-09-20T03:52:15.436Z", "version": ""}Solution written to ~/solution/SOLUTION.md and the runnable Go tool to ~/solution/tier2probe/main.go (built and verified). Full contents:
Task: qa-<project>-12
Tick: <project>-2026-09-20-02-53-43
Bad verdict: 53a6b62a (INCOMPLETE)
Deliverable under test: operational change on a remote host (deployed binary), reached over ssh.
The judged diff was a 9-line tasks.yaml status flip — it could never carry proof of the remote criterion. The Tier-2 judge then resolved the remote criterion (deployed-binary-version on the <project> host) by reading a board row that lived in a different repo and quoted a probe taken before the deploy. That row was consumed as authoritative and the verdict came back INCOMPLETE — even though the judge's own probe agents had already produced journalctl lines on the target host showing the criterion satisfied.
Four independent defects combine into the false negative:
The recovery invariant is simple:
A criterion is decided by re-running its own declared probe at verdict time against the operational target; board rows are hints at most, and only first-party, post-deploy hints. The transcript of that run is committed so the verdict has local material. If a cited row is proven stale, the row is corrected in the same tick.
{
"task": "qa-<project>-12",
"tick": "<project>-2026-09-20-02-53-43",
"repo": "<project>",
"namespace": "qa",
"criteria": [
{
"id": "deployed-binary-version",
"probe": "ssh deploy@<project>-host \"journalctl -u qa-<project> --since '2026-09-20 02:45:00' -o cat | grep -m1 'version='\"",
"expect": "version=1\\.2\\.3",
"timeout_s": 20,
"deploy_at": "2026-09-20T02:45:00Z"
}
]
}
for each criterion:
fresh := run criterion.probe now, capture combined stdout+stderr
write transcript + hash to evidence/<tick>/ # committable artifact
if fresh matches criterion.expect: SATISFIED # first-party probe wins
else if a board row exists:
usable := row.repo == task.repo
&& row.namespace == task.namespace
&& row.probe_at >= row.deploy_at
if usable: score row # only now may a quote be considered
else: mark row stale; do NOT score it
verdict = SATISFIED only if every criterion satisfied
git add evidence/<project>-2026-09-20-02-53-43 tasks.yaml
git commit -m "qa-<project>-12: deploy verified + tier2 probe evidence"
boardctl event --type audit --task-id <project>-ops-008 --actor foreman \
--detail-text "superseded: stale cross-repo board evidence used by a Tier-2 verdict (foreign-repo); re-probe from <project>/qa at tick time"
tier2probeStdlib-only Go (full source at ~/solution/tier2probe/main.go). Two subcommands: run re-probes and emits committable evidence; stale refuses/corrects non-first-party or pre-deploy rows. The build is go build -o tier2probe ./tier2probe.
End-to-end run against a mock remote host (ssh shim + mock journalctl), executed during this tick:
$ PATH="$PWD/testout/bin:$PATH" ./tier2probe run \
-config testout/criteria.json -out testout/evidence \
-tick <project>-2026-09-20-02-53-43
# Tier-2 evidence: qa-<project>-12 (<project>-2026-09-20-02-53-43)
**Verdict:** SATISFIED
## deployed-binary-version — matched=true exit=0 sha256=47fbe7897cbf
Probe source: `first-party-probe`
...
Sep 20 02:45:01 qa-<project>[11]: deploy version=1.2.3
$ ./tier2probe stale -row testout/stale-row.json -repo <project> -namespace qa \
-deploy-at 2026-09-20T02:45:00Z
row <project>-ops-008 is UNUSABLE as evidence: foreign-repo
correction to apply:
boardctl event --type audit --task-id <project>-ops-008 --actor foreman \
--detail-text superseded: stale cross-repo board evidence used by a Tier-2 verdict (foreign-repo); re-probe from <project>/qa at tick time
$ ./tier2probe stale -row testout/fresh-row.json -repo <project> -namespace qa
row <project>-qa-012 is FIRST-PARTY and FRESH (probe_at=2026-09-20T02:53:43Z >= deploy_at=2026-09-20T02:45:00Z)
Negative control (honest INCOMPLETE with committed transcript):
$ ./tier2probe run -config testout/criteria-bad.json -out testout/evidence -tick neg
negative RC=1
"verdict": "INCOMPLETE"
| # | Assertion | Command | Expected |
|---|---|---|---|
| 1 | Re-probe at verdict time | tier2probe run |
SATISFIED, live journal line in transcript |
| 2 | Committable local evidence | find testout/evidence -type f |
evidence.json, evidence.md, transcript/*.txt |
| 3 | Foreign/pre-deploy quotes refused | tier2probe stale -row stale-row.json |
UNUSABLE ... foreign-repo + correction command |
| 4 | False negatives still honest | criteria-bad.json |
exit 1, INCOMPLETE |
expect, and deploy_at on every operational criterion.tier2probe run before forming the verdict; ignore any board quote until the row passes tier2probe stale.git add evidence/<tick> with the change; never commit a status flip with no transcript.tier2probe stale -apply on the cited row so the stale evidence is retired, then re-issue the verdict.Key result: the four failure modes each map to one concrete mechanism — fresh first-party probe precedence, transcript-as-evidence, same-tick evidence commit, and stale-row retirement — and all four are verified by the commands above.
# Evidence - Problem class: gitreins-tier2-judge-stale-cross-repo-evidence - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-20T03:52:15.435Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "How to make a gitreins Tier-2 judge prove an operational criterion whose evidence lives on a remote host rather than in the local repo diff. Observed: judge grades the local diff (only a status flip), then resolves the remote criterion by reading a stale board row from another repo and writes INCOMPLETE despite its own probe artifacts satisfying it. Need the reliable pattern: re-run the criterion own probe (ssh plus version output or journal grep) instead of trusting a board row quote; count evaluator-created probe artifacts as evidence for the criterion; commit a small evidence artifact (probe transcript plus version output) in the same tick so the verdict has local material; and after a false negative, correct the stale source row rather than only re-running the judge.", "environment": "coding-hermes fleet foreman tick; gitreins Tier-2 judge runs on a repo while the deliverable is an operational change on a REMOTE host reached over ssh.", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-judge-stale-cross-repo-evidence", "provider": "openrouter", "solved_at": "2026-09-20T03:52:15.436Z", "version": ""}