jev-question-gate-reproposes-just-answered-question
Done. The verified solution is written to /workspace/solution.md (19 KB, self-contained).
Root cause. auger ask had two gates against storing a duplicate question, and the live bug needed both to fail — they did:
already_answered noul is computed in cmd_ask against evidence recalled from the project seed, not from the proposed question text. A verbatim repeat of a just-answered question scored 0.24–0.30, below T_ANSWERED = 0.55, so JEV called it "genuinely new". It is a relevance judge over embeddings, not a duplicate detector.store_proposed_question queried question … status=eq.open. Once answer --question-id Q-000001 flipped Q-000001 to answered, it dropped out of that query, so Q-000002 (identical text) was inserted.Fix. Added question_already_recorded(ns, project_id, text) which reads the whole question ledger (all statuses) and refuses the store, naming the prior row:
- open → already open as Q-000001
- answered/linked → already answered as Q-000001
- moot/budget_thin → already recorded as Q-000001 (…) (a question needed again is reopened by the graph, not cloned)
store_proposed_question now calls it instead of the open-only query, so the deterministic gate wins regardless of the model's noul. Fail-closed on a ledger read error.
Verification.
- Offline before/after harness (verify_fix.py, embedded in the doc): old guard stores Q-000002, new guard refuses with already answered as Q-000001; first-ask and open-duplicate paths preserved. ALL CHECKS PASS.
- py_compile: OK · ruff check: clean · pytest tests/test_auger.py: 13 passed, 118 loudly skipped (no DuckBrain token in this sandbox; live e2e command included).
- Added regression test test_ask_does_not_repropose_a_question_the_last_answer_closed.
Live DuckBrain at <ip-address>:3000 requires a token not present here, so the live arm could not run in-sandbox; the exact command to confirm it on a token-bearing host is in section 5.3.
# Evidence - Problem class: jev-question-gate-reproposes-just-answered-question - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T03:44:49.076Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: in a decision-drilling CLI backed by an LLM gate (auger 'ask' verb uses JEV scoring), the loop reproposed the IDENTICAL question text ('What test proves this works?') in the round immediately after the user answered exactly that question. Root cause: the ask pipeline computes jev_already_answered (noul) for the PROPOSED question text against stored evidence, but the just-recorded decision's embedded evidence was written with a chosen text that differs from the option/label wording (see sibling class: chosen/option mismatch); the noul check scored 0.24-0.3 (< threshold ~0.45+) so the gate judged the already-answered question 'genuinely new' and stored Q-000002 duplicating Q-000001. After the duplicate answer D-002 landed, the next ask returned 'nothing left worth asking' (JEV subject below threshold) \u2014 the loop self-heals only after wasting one answer. Fix direction: the already-answered gate should compare against the just-closed question id (Q-000001 was closed by D-001) or embed the closed-question id in the evidence key and exclude recently-closed questions from askable proposals for one round. Diagnosed live 2026-09-24 namespace df5-real: ask -> Q-000001; answer --question-id Q-000001; ask -> stored Q-000002 with same text; answer; ask -> 'nothing left worth asking'.", "environment": "auger v0.1 CLI + JEV gate (typesafe/jev-1.13-20260917) via 9router, DuckBrain substrate localhost", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "jev-question-gate-reproposes-just-answered-question", "provider": "openrouter", "solved_at": "2026-09-24T03:44:49.077Z", "version": ""}Done. The verified solution is written to /workspace/solution.md (19 KB, self-contained).
Root cause. auger ask had two gates against storing a duplicate question, and the live bug needed both to fail — they did:
already_answered noul is computed in cmd_ask against evidence recalled from the project seed, not from the proposed question text. A verbatim repeat of a just-answered question scored 0.24–0.30, below T_ANSWERED = 0.55, so JEV called it "genuinely new". It is a relevance judge over embeddings, not a duplicate detector.store_proposed_question queried question … status=eq.open. Once answer --question-id Q-000001 flipped Q-000001 to answered, it dropped out of that query, so Q-000002 (identical text) was inserted.Fix. Added question_already_recorded(ns, project_id, text) which reads the whole question ledger (all statuses) and refuses the store, naming the prior row:
- open → already open as Q-000001
- answered/linked → already answered as Q-000001
- moot/budget_thin → already recorded as Q-000001 (…) (a question needed again is reopened by the graph, not cloned)
store_proposed_question now calls it instead of the open-only query, so the deterministic gate wins regardless of the model's noul. Fail-closed on a ledger read error.
Verification.
- Offline before/after harness (verify_fix.py, embedded in the doc): old guard stores Q-000002, new guard refuses with already answered as Q-000001; first-ask and open-duplicate paths preserved. ALL CHECKS PASS.
- py_compile: OK · ruff check: clean · pytest tests/test_auger.py: 13 passed, 118 loudly skipped (no DuckBrain token in this sandbox; live e2e command included).
- Added regression test test_ask_does_not_repropose_a_question_the_last_answer_closed.
Live DuckBrain at <ip-address>:3000 requires a token not present here, so the live arm could not run in-sandbox; the exact command to confirm it on a token-bearing host is in section 5.3.
# Evidence - Problem class: jev-question-gate-reproposes-just-answered-question - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T03:44:49.076Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: in a decision-drilling CLI backed by an LLM gate (auger 'ask' verb uses JEV scoring), the loop reproposed the IDENTICAL question text ('What test proves this works?') in the round immediately after the user answered exactly that question. Root cause: the ask pipeline computes jev_already_answered (noul) for the PROPOSED question text against stored evidence, but the just-recorded decision's embedded evidence was written with a chosen text that differs from the option/label wording (see sibling class: chosen/option mismatch); the noul check scored 0.24-0.3 (< threshold ~0.45+) so the gate judged the already-answered question 'genuinely new' and stored Q-000002 duplicating Q-000001. After the duplicate answer D-002 landed, the next ask returned 'nothing left worth asking' (JEV subject below threshold) \u2014 the loop self-heals only after wasting one answer. Fix direction: the already-answered gate should compare against the just-closed question id (Q-000001 was closed by D-001) or embed the closed-question id in the evidence key and exclude recently-closed questions from askable proposals for one round. Diagnosed live 2026-09-24 namespace df5-real: ask -> Q-000001; answer --question-id Q-000001; ask -> stored Q-000002 with same text; answer; ask -> 'nothing left worth asking'.", "environment": "auger v0.1 CLI + JEV gate (typesafe/jev-1.13-20260917) via 9router, DuckBrain substrate localhost", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "jev-question-gate-reproposes-just-answered-question", "provider": "openrouter", "solved_at": "2026-09-24T03:44:49.077Z", "version": ""}