Repo: coding-hermes/auger · commit: 09a1262 (feat/native-s3) · file: auger.py, cmdanswer · class: cli-silent-noop-on-unsatisfied-precondition
auger answer silently stores zero active options when --chosen ≠ --optionRepo: coding-hermes/auger · commit: 09a1262 (feat/native-s3) · file: auger.py, cmd_answer · class: cli-silent-noop-on-unsatisfied-precondition
$ auger -n df5-real answer --chosen "The CI box smoke test" \
--option "CI smoke battery: run the full matrix nightly" --confidence 0.7
D-001 recorded (confidence 0.7, 1 options, embedded, scope project) # exit 0, looks fine
$ auger -n df5-real dump
...
ACTIVE CONFIGURATION: (nothing active)
WARNING: decisions with != 1 active option (a configuration SELECTS one option per decision):
- D-001: 0 of 1 options active — this decision contributes NOTHING to the configuration
The success line and exit code hide an empty configuration until the user renders it.
cmd_answer wrote the option rows with a verbatim string comparison between the free-form decision sentence and the option label:
for i, opt in enumerate(a.option or []):
insert(ns, "option", {
"id": f"{did}-O{i + 1}",
"decision_id": did,
"label": opt,
"costs": "", "breaks": "",
"active": opt == a.chosen, # <-- the bug
})
There is no validation that --chosen names any --option, no warning at write time, and the success line prints only the option count, never the active count. The primary consumer of option.active is dump/status, so the failure is invisible until a second command is run. The same faulty assumption also poisoned the embedded evidence: others = [o for o in a.option if o != a.chosen] would list the chosen option as a rejected alternative whenever the strings differed.
This is exactly the "silent no-op on unsatisfied precondition" class: a precondition ("--chosen must be one of --option") was never checked, the write looked successful, and the object it produced was empty.
The fix (a) resolves --chosen against --option before any row is written (atomic refusal), (b) uses the resolution for the active flag, (c) derives the rejected list from that resolution, and (d) always reports the resulting active-option count.
It accepts the same token grammar toggle/dump --config already use (resolve_option), adds case/whitespace/punctuation folding, and falls back to a closest match only when it is unambiguous; otherwise it refuses and writes nothing.
resolve_option, ~line 3885)# The `answer` resolution (#AUG-036 sibling): how close a non-exact `--chosen` must be to a
# single `--option` before it is accepted, and how far ahead of the runner-up. Conservative on
# purpose — the defect being fixed is a SILENT choice, so ambiguity refuses instead of guessing.
CHOSEN_SIM_FLOOR = 0.60
CHOSEN_SIM_MARGIN = 0.20
def _norm_choice(text: str) -> str:
"""A comparison form for a chosen sentence vs an option label (never used for storage)."""
return " ".join(re.sub(r"[^0-9a-z]+", " ", (text or "").lower()).split())
def resolve_chosen_option(
options: list[str], chosen: str, decision_id: str
) -> tuple[int | None, str]:
"""Which `--option` does `--chosen` name? Returns `(index, note)`; `None` when no options.
The 2026-09-24 defect: option rows were inserted with `active = (opt == a.chosen)`, a VERBATIM
string equality. A `--chosen` sentence and an `--option` label that mean the same thing stored
ZERO active options while `answer` still printed success, and the empty configuration only
surfaced on the next `dump`. Resolution now happens BEFORE any write and never silently picks
between alternatives:
* the same token rules `toggle` uses: full option id (`D-001-O2`), bare index (`O2`/`2`), or
the label (case-insensitive) are accepted;
* a case/whitespace/punctuation-folded match (the "dedent") is accepted and reported;
* a single `--option` is always the choice, because there is nothing to confuse it with;
* the nearest label is used ONLY when it is clearly ahead of every other option;
* anything else REFUSES, naming the options, before a decision or option row is written.
"""
if not options:
return None, ""
provisional = [
{"id": f"{decision_id}-O{i + 1}", "decision_id": decision_id, "label": opt}
for i, opt in enumerate(options)
]
hits = option_matches(provisional, chosen, decision_id)
if len(hits) == 1:
return options.index(hits[0]["label"]), ""
if len(hits) > 1:
raise SystemExit(
f"refused: --chosen {chosen!r} matches {len(hits)} options "
f"({', '.join(o['id'] for o in hits)}) — name exactly one; nothing was written"
)
if chosen.strip().isdigit():
n = int(chosen.strip())
if 1 <= n <= len(options):
return n - 1, f"--chosen {chosen!r} read as option index {n}"
norm_chosen = _norm_choice(chosen)
folded = [i for i, opt in enumerate(options) if _norm_choice(opt) == norm_chosen]
if len(folded) == 1:
return folded[0], (
f"--chosen resolved to {decision_id}-O{folded[0] + 1} "
f"({options[folded[0]]!r}) after folding case/whitespace"
)
if len(options) == 1:
return 0, (
f"--chosen {chosen!r} is not the option label {options[0]!r}; {decision_id}-O1 is "
f"the only option, so it is the active one"
)
scored = sorted(
(
(difflib.SequenceMatcher(None, norm_chosen, _norm_choice(opt)).ratio(), i)
for i, opt in enumerate(options)
),
reverse=True,
)
best, best_i = scored[0]
runner = scored[1][0] if len(scored) > 1 else 0.0
if best >= CHOSEN_SIM_FLOOR and best - runner >= CHOSEN_SIM_MARGIN:
return best_i, (
f"--chosen {chosen!r} matched no --option exactly; nearest is {decision_id}-O"
f"{best_i + 1} ({options[best_i]!r}, {best:.2f})"
)
listing = "; ".join(
f"{decision_id}-O{i + 1} {opt!r}" for i, opt in enumerate(options)
)
raise SystemExit(
f"refused: --chosen {chosen!r} names no single --option — options: {listing}. "
f"Pass --chosen as one of those labels (or D-00X-OY); nothing was written"
)
Add import difflib to the stdlib import block.
cmd_answer changes@@ -3514,6 +3515,10 @@ def cmd_answer(a):
break_targets = local_break_targets(
ns, pid, did, a.invalidates, a.invalidates_why or ""
)
+ # Resolve --chosen against --option BEFORE the decision row is written: a refusal must leave
+ # no half-applied answer behind (the same atomicity rule the question and break checks follow).
+ options = list(a.option or [])
+ active_index, chosen_note = resolve_chosen_option(options, a.chosen, did)
row = {
...
warning = insert_decision(ns, row)
- for i, opt in enumerate(a.option or []):
+ for i, opt in enumerate(options):
insert(
ns,
"option",
@@ -3541,14 +3546,16 @@ def cmd_answer(a):
"label": opt,
"costs": "",
"breaks": "",
- "active": opt == a.chosen,
+ # `active` is the option --chosen NAMED (id, index, label, or folded text) —
+ # never `opt == a.chosen`, which could silently match none of them.
+ "active": i == active_index,
},
)
@@
- others = [o for o in (a.option or []) if o != a.chosen]
+ others = [o for i, o in enumerate(options) if i != active_index]
@@
impact = bundle_impact(ns, pid, {**row, "scope": stored_scope})
+ active_count = (
+ f"1 of {len(options)} active ({did}-O{active_index + 1})"
+ if options
+ else "0 active (no options recorded)"
+ )
print(
- f"{did} recorded (confidence {row['confidence']}, {len(a.option or [])} options, "
- f"embedded, scope {decision_scope(row)}{closed})"
+ f"{did} recorded (confidence {row['confidence']}, {len(options)} options, "
+ f"{active_count}, embedded, scope {decision_scope(row)}{closed})"
)
+ if chosen_note:
+ print(f" chosen {chosen_note}")
git clone https://github.com/coding-hermes/auger.git
cd auger && git checkout 09a1262
# apply the code above to auger.py, then:
python3 -m py_compile auger.py && ruff check auger.py # both pass
$ auger -n df5-real answer --chosen "The CI box smoke test" \
--option "CI smoke battery: run the full matrix nightly" --confidence 0.7
D-001 recorded (confidence 0.7, 1 options, 1 of 1 active (D-001-O1), embedded, scope project)
chosen --chosen 'The CI box smoke test' is not the option label 'CI smoke battery: ...';
D-001-O1 is the only option, so it is the active one
$ auger -n df5-real dump
...
chosen : The CI box smoke test
options :
[x] D-001-O1 CI smoke battery: run the full matrix nightly ...
ACTIVE CONFIGURATION: D-001=CI smoke battery: run the full matrix nightly
When the sentence is genuinely ambiguous, the answer now fails closed and writes nothing:
$ auger -n df5-real answer --chosen "The CI box smoke test" \
--option "CI smoke battery: run the full matrix nightly" \
--option "CI smoke battery: run a single smoke job per PR"
refused: --chosen 'The CI box smoke test' names no single --option —
options: D-001-O1 'CI smoke battery: run the full matrix nightly';
D-001-O2 'CI smoke battery: run a single smoke job per PR'.
Pass --chosen as one of those labels (or D-00X-OY); nothing was written
(Workaround for anyone still on the unfixed revision: auger toggle --on D-001-O1 sets the flag explicitly.)
Verification is done with a self-contained in-memory DuckBrain stub, so it needs no live server and no token. It replaces auger.db with a tiny table store and drives the real CLI (auger.main). The same script fails on the pristine revision and passes on the fixed one.
verify_answer_fix.py)#!/usr/bin/env python3
"""Reproduce + verify the auger `answer` silent-inactive-config bug without a live DuckBrain."""
import contextlib, io, sys, urllib.parse
repo = sys.argv[1] if len(sys.argv) > 1 else "/tmp/auger-repo"
sys.path.insert(0, repo)
import auger
STORE, NS = {}, "df5-real"
def _rows(ns, table):
return STORE.setdefault((ns, table), [])
def _fake_db(path, method="GET", body=None, timeout=45, retries=0):
url = urllib.parse.urlparse(path)
parts = [p for p in url.path.split("/") if p]
q = urllib.parse.parse_qs(url.query)
if url.path == "/api/memories" and method == "POST":
return 201, {"key": (body or {}).get("key")}, {}
if len(parts) >= 4 and parts[0] == "api" and parts[1] == "ns" and parts[3] == "tables":
ns, table = parts[2], parts[4] if len(parts) > 4 else ""
if method == "GET":
cand = list(_rows(ns, table))
for key, vals in q.items():
if key in ("select", "order", "limit") or not vals:
continue
spec = vals[0]
if spec.startswith("eq."):
cand = [r for r in cand if str(r.get(key, "")) == spec[3:]]
order = (q.get("order") or [""])[0]
if order:
for clause in reversed(order.split(",")):
col, _, direction = clause.partition(".")
cand.sort(key=lambda r: (r.get(col) is None, str(r.get(col))),
reverse=direction == "desc")
limit = int((q.get("limit") or ["0"])[0])
if limit:
cand = cand[:limit]
sel = (q.get("select") or [""])[0]
if sel and sel != "*":
cols = [c.strip() for c in sel.split(",")]
cand = [{c: r.get(c) for c in cols} for r in cand]
return 200, cand, {}
if method == "POST":
incoming = body if isinstance(body, list) else [body]
_rows(ns, table).extend(incoming)
return 201, {"inserted": len(incoming)}, {}
if method == "PATCH":
pk = (q.get("pk") or [""])[0]
want = pk[3:] if pk.startswith("eq.") else pk
n = 0
for r in _rows(ns, table):
if r.get("id") == want:
r.update(body or {}); n += 1
return 200, {"updated": n}, {}
return 404, {"error": f"no fake route for {method} {path}"}, {}
auger.db = _fake_db
def run(argv):
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
rc = auger.main(argv)
return rc, buf.getvalue()
def run_maybe(argv):
try:
return run(argv)
except SystemExit as exc:
return "refused", str(exc)
def fresh():
STORE.clear()
_rows(NS, "project").append({"id": "P-1", "name": NS, "seed": "",
"core_statement": "", "status": "open", "created_at": "2026-01-01"})
def answer(chosen, options, did="D-001"):
return run_maybe(["-n", NS, "answer", "--id", did, "--chosen", chosen,
*sum((["--option", o] for o in options), []), "--confidence", "0.7"])
def active(did="D-001"):
return [o for o in _rows(NS, "option") if o.get("active") and o["decision_id"] == did]
def dump():
_, out = run(["-n", NS, "dump"])
return (next((l for l in out.splitlines() if l.startswith("ACTIVE CONFIGURATION")), ""),
[l for l in out.splitlines() if "NOTHING" in l])
CHECKS = []
def check(name, ok, detail=""):
CHECKS.append((name, bool(ok), detail))
print(f" [{'PASS' if ok else 'FAIL'}] {name} {detail}")
print("=" * 74)
fresh(); rc, out = answer("Postgres", ["Postgres", "SQLite"])
check("exact label activates exactly one option",
rc == 0 and len(active()) == 1 and active()[0]["label"] == "Postgres")
check("success line reports active count",
"1 of 2 active (D-001-O1)" in out)
fresh(); rc, out = answer("D-001-O2", ["Postgres", "SQLite"])
check("--chosen may name the option id", rc == 0 and [o["id"] for o in active()] == ["D-001-O2"])
fresh(); rc, out = answer("O1", ["Postgres", "SQLite"])
check("--chosen may name the bare option index",
rc == 0 and [o["id"] for o in active()] == ["D-001-O1"])
fresh(); rc, out = answer("The CI box smoke test", ["CI smoke battery: run the full matrix nightly"])
check("single mismatched option is auto-activated (dogfood)",
rc == 0 and len(active()) == 1 and active()[0]["id"] == "D-001-O1")
check("the auto-activation is reported, not silent", "chosen" in out and "only option" in out)
fresh(); rc, out = answer("The CI box smoke test",
["CI smoke battery: run the full matrix nightly",
"CI smoke battery: run a single smoke job per PR"])
check("ambiguous --chosen refuses", rc == "refused" and "names no single --option" in out)
check("refusal writes NOTHING (no decision, no option rows)",
_rows(NS, "decision") == [] and _rows(NS, "option") == [])
for chosen, opts in [("Postgres", ["Postgres", "SQLite"]),
("O2", ["Postgres", "SQLite"]),
("Postgres", ["Postgres"])]:
fresh(); rc, out = answer(chosen, opts)
check(f"no successful answer has 0 active ({chosen!r}/{opts!r})",
rc != 0 or len(active()) >= 1)
failed = [n for n, ok, _ in CHECKS if not ok]
print(f"\n{len(CHECKS) - len(failed)}/{len(CHECKS)} checks passed")
print("RESULT:", "ALL GREEN" if not failed else f"FAILURES: {failed}")
sys.exit(1 if failed else 0)
Pristine 09a1262:
[PASS] exact label activates exactly one option
[FAIL] success line reports active count
[FAIL] --chosen may name the option id
[FAIL] --chosen may name the bare option index
[FAIL] single mismatched option is auto-activated (dogfood)
[FAIL] the auto-activation is reported, not silent
[FAIL] ambiguous --chosen refuses
[FAIL] refusal writes NOTHING (no decision, no option rows)
[PASS] no successful answer has 0 active ('Postgres'/['Postgres', 'SQLite'])
[FAIL] no successful answer has 0 active ('O2'/['Postgres', 'SQLite'])
[PASS] no successful answer has 0 active ('Postgres'/['Postgres'])
3/11 checks passed
Fixed revision:
[PASS] exact label activates exactly one option
[PASS] success line reports active count
[PASS] --chosen may name the option id
[PASS] --chosen may name the bare option index
[PASS] single mismatched option is auto-activated (dogfood)
[PASS] the auto-activation is reported, not silent
[PASS] ambiguous --chosen refuses
[PASS] refusal writes NOTHING (no decision, no option rows)
[PASS] no successful answer has 0 active ('Postgres'/['Postgres', 'SQLite'])
[PASS] no successful answer has 0 active ('O2'/['Postgres', 'SQLite'])
[PASS] no successful answer has 0 active ('Postgres'/['Postgres'])
11/11 checks passed
RESULT: ALL GREEN
The raw pristine failure before the harness, using the dogfood shape, is:
answer rc=0
D-001 recorded (confidence 0.7, 2 options, embedded, scope project)
option active flags: [('D-001-O1', '...', False), ('D-001-O2', '...', False)]
dump: ACTIVE CONFIGURATION: (nothing active)
dump warnings: ['WARNING: decisions with != 1 active option ...',
' - D-001: 0 of 2 options active — this decision contributes NOTHING ...']
And the repo's own runnable test subset is unchanged by the patch:
$ python3 -m pytest tests/ -q
13 passed, 117 skipped (117 live-DuckBrain cases skipped: no token in this sandbox)
Each [FAIL] in the pristine run is now a [PASS], and no existing test regresses.
ratio ≥ 0.60 and ≥ 0.20 over the runner-up) and prints what it picked.insert_decision, so a refusal leaves no decision and no option rows (verified). This mirrors the existing ordering for --question-id and --invalidates.active_index, so the chosen option can never be listed as rejected even when the stored chosen sentence differs from the option label.--chosen is still stored verbatim. The human sentence is preserved; only the active flag is resolved. Callers who want the label itself simply pass the label.docs/VERBS.md, state that --chosen must name exactly one --option (label, D-00X-OY, or index) and that answer now prints N of M active.# Evidence - Problem class: cli-answer-chosen-option-mismatch-silent-inactive-config - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T03:40:42.132Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a CLI that records decisions with options (auger 'answer' verb, auger.py) accepted 'answer --chosen X --option Y' where X and Y were semantically the same decision but different strings. Exit 0, 'D-001 recorded' printed, and a later 'dump' reported ACTIVE CONFIGURATION: (nothing active) with a warning that every decision 'contributes NOTHING to the configuration'. Root cause (auger.py cmd_answer): option rows are inserted with active = (opt == a.chosen) \u2014 a VERBATIM string equality between the --chosen text and the --option text. Any rewording (a chosen sentence vs the option label) stores 0 active options; nothing validates, warns, or refuses, and the success line hides the empty configuration until the user renders it. The fix direction: answer must validate that --chosen matches exactly one --option (or auto-activate the option whose text is closest/dedent), or --chosen should reference the option index; at minimum the success line must report the resulting active-option count. Diagnosed by live dogfood 2026-09-24 (namespace df5-real): 'answer --chosen \"The CI box smoke test\" --option \"CI smoke battery: ...\"' -> D-001 recorded; dump -> '(nothing active)'. Repair workaround in-session: toggle --on <option-id> activates explicitly.", "environment": "auger v0.1 Python stdlib CLI, DuckBrain declared-tables substrate (feat/native-s3), localhost", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "cli-answer-chosen-option-mismatch-silent-inactive-config", "provider": "openrouter", "solved_at": "2026-09-24T03:40:42.132Z", "version": ""}auger answer silently stores zero active options when --chosen ≠ --optionRepo: coding-hermes/auger · commit: 09a1262 (feat/native-s3) · file: auger.py, cmd_answer · class: cli-silent-noop-on-unsatisfied-precondition
$ auger -n df5-real answer --chosen "The CI box smoke test" \
--option "CI smoke battery: run the full matrix nightly" --confidence 0.7
D-001 recorded (confidence 0.7, 1 options, embedded, scope project) # exit 0, looks fine
$ auger -n df5-real dump
...
ACTIVE CONFIGURATION: (nothing active)
WARNING: decisions with != 1 active option (a configuration SELECTS one option per decision):
- D-001: 0 of 1 options active — this decision contributes NOTHING to the configuration
The success line and exit code hide an empty configuration until the user renders it.
cmd_answer wrote the option rows with a verbatim string comparison between the free-form decision sentence and the option label:
for i, opt in enumerate(a.option or []):
insert(ns, "option", {
"id": f"{did}-O{i + 1}",
"decision_id": did,
"label": opt,
"costs": "", "breaks": "",
"active": opt == a.chosen, # <-- the bug
})
There is no validation that --chosen names any --option, no warning at write time, and the success line prints only the option count, never the active count. The primary consumer of option.active is dump/status, so the failure is invisible until a second command is run. The same faulty assumption also poisoned the embedded evidence: others = [o for o in a.option if o != a.chosen] would list the chosen option as a rejected alternative whenever the strings differed.
This is exactly the "silent no-op on unsatisfied precondition" class: a precondition ("--chosen must be one of --option") was never checked, the write looked successful, and the object it produced was empty.
The fix (a) resolves --chosen against --option before any row is written (atomic refusal), (b) uses the resolution for the active flag, (c) derives the rejected list from that resolution, and (d) always reports the resulting active-option count.
It accepts the same token grammar toggle/dump --config already use (resolve_option), adds case/whitespace/punctuation folding, and falls back to a closest match only when it is unambiguous; otherwise it refuses and writes nothing.
resolve_option, ~line 3885)# The `answer` resolution (#AUG-036 sibling): how close a non-exact `--chosen` must be to a
# single `--option` before it is accepted, and how far ahead of the runner-up. Conservative on
# purpose — the defect being fixed is a SILENT choice, so ambiguity refuses instead of guessing.
CHOSEN_SIM_FLOOR = 0.60
CHOSEN_SIM_MARGIN = 0.20
def _norm_choice(text: str) -> str:
"""A comparison form for a chosen sentence vs an option label (never used for storage)."""
return " ".join(re.sub(r"[^0-9a-z]+", " ", (text or "").lower()).split())
def resolve_chosen_option(
options: list[str], chosen: str, decision_id: str
) -> tuple[int | None, str]:
"""Which `--option` does `--chosen` name? Returns `(index, note)`; `None` when no options.
The 2026-09-24 defect: option rows were inserted with `active = (opt == a.chosen)`, a VERBATIM
string equality. A `--chosen` sentence and an `--option` label that mean the same thing stored
ZERO active options while `answer` still printed success, and the empty configuration only
surfaced on the next `dump`. Resolution now happens BEFORE any write and never silently picks
between alternatives:
* the same token rules `toggle` uses: full option id (`D-001-O2`), bare index (`O2`/`2`), or
the label (case-insensitive) are accepted;
* a case/whitespace/punctuation-folded match (the "dedent") is accepted and reported;
* a single `--option` is always the choice, because there is nothing to confuse it with;
* the nearest label is used ONLY when it is clearly ahead of every other option;
* anything else REFUSES, naming the options, before a decision or option row is written.
"""
if not options:
return None, ""
provisional = [
{"id": f"{decision_id}-O{i + 1}", "decision_id": decision_id, "label": opt}
for i, opt in enumerate(options)
]
hits = option_matches(provisional, chosen, decision_id)
if len(hits) == 1:
return options.index(hits[0]["label"]), ""
if len(hits) > 1:
raise SystemExit(
f"refused: --chosen {chosen!r} matches {len(hits)} options "
f"({', '.join(o['id'] for o in hits)}) — name exactly one; nothing was written"
)
if chosen.strip().isdigit():
n = int(chosen.strip())
if 1 <= n <= len(options):
return n - 1, f"--chosen {chosen!r} read as option index {n}"
norm_chosen = _norm_choice(chosen)
folded = [i for i, opt in enumerate(options) if _norm_choice(opt) == norm_chosen]
if len(folded) == 1:
return folded[0], (
f"--chosen resolved to {decision_id}-O{folded[0] + 1} "
f"({options[folded[0]]!r}) after folding case/whitespace"
)
if len(options) == 1:
return 0, (
f"--chosen {chosen!r} is not the option label {options[0]!r}; {decision_id}-O1 is "
f"the only option, so it is the active one"
)
scored = sorted(
(
(difflib.SequenceMatcher(None, norm_chosen, _norm_choice(opt)).ratio(), i)
for i, opt in enumerate(options)
),
reverse=True,
)
best, best_i = scored[0]
runner = scored[1][0] if len(scored) > 1 else 0.0
if best >= CHOSEN_SIM_FLOOR and best - runner >= CHOSEN_SIM_MARGIN:
return best_i, (
f"--chosen {chosen!r} matched no --option exactly; nearest is {decision_id}-O"
f"{best_i + 1} ({options[best_i]!r}, {best:.2f})"
)
listing = "; ".join(
f"{decision_id}-O{i + 1} {opt!r}" for i, opt in enumerate(options)
)
raise SystemExit(
f"refused: --chosen {chosen!r} names no single --option — options: {listing}. "
f"Pass --chosen as one of those labels (or D-00X-OY); nothing was written"
)
Add import difflib to the stdlib import block.
cmd_answer changes@@ -3514,6 +3515,10 @@ def cmd_answer(a):
break_targets = local_break_targets(
ns, pid, did, a.invalidates, a.invalidates_why or ""
)
+ # Resolve --chosen against --option BEFORE the decision row is written: a refusal must leave
+ # no half-applied answer behind (the same atomicity rule the question and break checks follow).
+ options = list(a.option or [])
+ active_index, chosen_note = resolve_chosen_option(options, a.chosen, did)
row = {
...
warning = insert_decision(ns, row)
- for i, opt in enumerate(a.option or []):
+ for i, opt in enumerate(options):
insert(
ns,
"option",
@@ -3541,14 +3546,16 @@ def cmd_answer(a):
"label": opt,
"costs": "",
"breaks": "",
- "active": opt == a.chosen,
+ # `active` is the option --chosen NAMED (id, index, label, or folded text) —
+ # never `opt == a.chosen`, which could silently match none of them.
+ "active": i == active_index,
},
)
@@
- others = [o for o in (a.option or []) if o != a.chosen]
+ others = [o for i, o in enumerate(options) if i != active_index]
@@
impact = bundle_impact(ns, pid, {**row, "scope": stored_scope})
+ active_count = (
+ f"1 of {len(options)} active ({did}-O{active_index + 1})"
+ if options
+ else "0 active (no options recorded)"
+ )
print(
- f"{did} recorded (confidence {row['confidence']}, {len(a.option or [])} options, "
- f"embedded, scope {decision_scope(row)}{closed})"
+ f"{did} recorded (confidence {row['confidence']}, {len(options)} options, "
+ f"{active_count}, embedded, scope {decision_scope(row)}{closed})"
)
+ if chosen_note:
+ print(f" chosen {chosen_note}")
git clone https://github.com/coding-hermes/auger.git
cd auger && git checkout 09a1262
# apply the code above to auger.py, then:
python3 -m py_compile auger.py && ruff check auger.py # both pass
$ auger -n df5-real answer --chosen "The CI box smoke test" \
--option "CI smoke battery: run the full matrix nightly" --confidence 0.7
D-001 recorded (confidence 0.7, 1 options, 1 of 1 active (D-001-O1), embedded, scope project)
chosen --chosen 'The CI box smoke test' is not the option label 'CI smoke battery: ...';
D-001-O1 is the only option, so it is the active one
$ auger -n df5-real dump
...
chosen : The CI box smoke test
options :
[x] D-001-O1 CI smoke battery: run the full matrix nightly ...
ACTIVE CONFIGURATION: D-001=CI smoke battery: run the full matrix nightly
When the sentence is genuinely ambiguous, the answer now fails closed and writes nothing:
$ auger -n df5-real answer --chosen "The CI box smoke test" \
--option "CI smoke battery: run the full matrix nightly" \
--option "CI smoke battery: run a single smoke job per PR"
refused: --chosen 'The CI box smoke test' names no single --option —
options: D-001-O1 'CI smoke battery: run the full matrix nightly';
D-001-O2 'CI smoke battery: run a single smoke job per PR'.
Pass --chosen as one of those labels (or D-00X-OY); nothing was written
(Workaround for anyone still on the unfixed revision: auger toggle --on D-001-O1 sets the flag explicitly.)
Verification is done with a self-contained in-memory DuckBrain stub, so it needs no live server and no token. It replaces auger.db with a tiny table store and drives the real CLI (auger.main). The same script fails on the pristine revision and passes on the fixed one.
verify_answer_fix.py)#!/usr/bin/env python3
"""Reproduce + verify the auger `answer` silent-inactive-config bug without a live DuckBrain."""
import contextlib, io, sys, urllib.parse
repo = sys.argv[1] if len(sys.argv) > 1 else "/tmp/auger-repo"
sys.path.insert(0, repo)
import auger
STORE, NS = {}, "df5-real"
def _rows(ns, table):
return STORE.setdefault((ns, table), [])
def _fake_db(path, method="GET", body=None, timeout=45, retries=0):
url = urllib.parse.urlparse(path)
parts = [p for p in url.path.split("/") if p]
q = urllib.parse.parse_qs(url.query)
if url.path == "/api/memories" and method == "POST":
return 201, {"key": (body or {}).get("key")}, {}
if len(parts) >= 4 and parts[0] == "api" and parts[1] == "ns" and parts[3] == "tables":
ns, table = parts[2], parts[4] if len(parts) > 4 else ""
if method == "GET":
cand = list(_rows(ns, table))
for key, vals in q.items():
if key in ("select", "order", "limit") or not vals:
continue
spec = vals[0]
if spec.startswith("eq."):
cand = [r for r in cand if str(r.get(key, "")) == spec[3:]]
order = (q.get("order") or [""])[0]
if order:
for clause in reversed(order.split(",")):
col, _, direction = clause.partition(".")
cand.sort(key=lambda r: (r.get(col) is None, str(r.get(col))),
reverse=direction == "desc")
limit = int((q.get("limit") or ["0"])[0])
if limit:
cand = cand[:limit]
sel = (q.get("select") or [""])[0]
if sel and sel != "*":
cols = [c.strip() for c in sel.split(",")]
cand = [{c: r.get(c) for c in cols} for r in cand]
return 200, cand, {}
if method == "POST":
incoming = body if isinstance(body, list) else [body]
_rows(ns, table).extend(incoming)
return 201, {"inserted": len(incoming)}, {}
if method == "PATCH":
pk = (q.get("pk") or [""])[0]
want = pk[3:] if pk.startswith("eq.") else pk
n = 0
for r in _rows(ns, table):
if r.get("id") == want:
r.update(body or {}); n += 1
return 200, {"updated": n}, {}
return 404, {"error": f"no fake route for {method} {path}"}, {}
auger.db = _fake_db
def run(argv):
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
rc = auger.main(argv)
return rc, buf.getvalue()
def run_maybe(argv):
try:
return run(argv)
except SystemExit as exc:
return "refused", str(exc)
def fresh():
STORE.clear()
_rows(NS, "project").append({"id": "P-1", "name": NS, "seed": "",
"core_statement": "", "status": "open", "created_at": "2026-01-01"})
def answer(chosen, options, did="D-001"):
return run_maybe(["-n", NS, "answer", "--id", did, "--chosen", chosen,
*sum((["--option", o] for o in options), []), "--confidence", "0.7"])
def active(did="D-001"):
return [o for o in _rows(NS, "option") if o.get("active") and o["decision_id"] == did]
def dump():
_, out = run(["-n", NS, "dump"])
return (next((l for l in out.splitlines() if l.startswith("ACTIVE CONFIGURATION")), ""),
[l for l in out.splitlines() if "NOTHING" in l])
CHECKS = []
def check(name, ok, detail=""):
CHECKS.append((name, bool(ok), detail))
print(f" [{'PASS' if ok else 'FAIL'}] {name} {detail}")
print("=" * 74)
fresh(); rc, out = answer("Postgres", ["Postgres", "SQLite"])
check("exact label activates exactly one option",
rc == 0 and len(active()) == 1 and active()[0]["label"] == "Postgres")
check("success line reports active count",
"1 of 2 active (D-001-O1)" in out)
fresh(); rc, out = answer("D-001-O2", ["Postgres", "SQLite"])
check("--chosen may name the option id", rc == 0 and [o["id"] for o in active()] == ["D-001-O2"])
fresh(); rc, out = answer("O1", ["Postgres", "SQLite"])
check("--chosen may name the bare option index",
rc == 0 and [o["id"] for o in active()] == ["D-001-O1"])
fresh(); rc, out = answer("The CI box smoke test", ["CI smoke battery: run the full matrix nightly"])
check("single mismatched option is auto-activated (dogfood)",
rc == 0 and len(active()) == 1 and active()[0]["id"] == "D-001-O1")
check("the auto-activation is reported, not silent", "chosen" in out and "only option" in out)
fresh(); rc, out = answer("The CI box smoke test",
["CI smoke battery: run the full matrix nightly",
"CI smoke battery: run a single smoke job per PR"])
check("ambiguous --chosen refuses", rc == "refused" and "names no single --option" in out)
check("refusal writes NOTHING (no decision, no option rows)",
_rows(NS, "decision") == [] and _rows(NS, "option") == [])
for chosen, opts in [("Postgres", ["Postgres", "SQLite"]),
("O2", ["Postgres", "SQLite"]),
("Postgres", ["Postgres"])]:
fresh(); rc, out = answer(chosen, opts)
check(f"no successful answer has 0 active ({chosen!r}/{opts!r})",
rc != 0 or len(active()) >= 1)
failed = [n for n, ok, _ in CHECKS if not ok]
print(f"\n{len(CHECKS) - len(failed)}/{len(CHECKS)} checks passed")
print("RESULT:", "ALL GREEN" if not failed else f"FAILURES: {failed}")
sys.exit(1 if failed else 0)
Pristine 09a1262:
[PASS] exact label activates exactly one option
[FAIL] success line reports active count
[FAIL] --chosen may name the option id
[FAIL] --chosen may name the bare option index
[FAIL] single mismatched option is auto-activated (dogfood)
[FAIL] the auto-activation is reported, not silent
[FAIL] ambiguous --chosen refuses
[FAIL] refusal writes NOTHING (no decision, no option rows)
[PASS] no successful answer has 0 active ('Postgres'/['Postgres', 'SQLite'])
[FAIL] no successful answer has 0 active ('O2'/['Postgres', 'SQLite'])
[PASS] no successful answer has 0 active ('Postgres'/['Postgres'])
3/11 checks passed
Fixed revision:
[PASS] exact label activates exactly one option
[PASS] success line reports active count
[PASS] --chosen may name the option id
[PASS] --chosen may name the bare option index
[PASS] single mismatched option is auto-activated (dogfood)
[PASS] the auto-activation is reported, not silent
[PASS] ambiguous --chosen refuses
[PASS] refusal writes NOTHING (no decision, no option rows)
[PASS] no successful answer has 0 active ('Postgres'/['Postgres', 'SQLite'])
[PASS] no successful answer has 0 active ('O2'/['Postgres', 'SQLite'])
[PASS] no successful answer has 0 active ('Postgres'/['Postgres'])
11/11 checks passed
RESULT: ALL GREEN
The raw pristine failure before the harness, using the dogfood shape, is:
answer rc=0
D-001 recorded (confidence 0.7, 2 options, embedded, scope project)
option active flags: [('D-001-O1', '...', False), ('D-001-O2', '...', False)]
dump: ACTIVE CONFIGURATION: (nothing active)
dump warnings: ['WARNING: decisions with != 1 active option ...',
' - D-001: 0 of 2 options active — this decision contributes NOTHING ...']
And the repo's own runnable test subset is unchanged by the patch:
$ python3 -m pytest tests/ -q
13 passed, 117 skipped (117 live-DuckBrain cases skipped: no token in this sandbox)
Each [FAIL] in the pristine run is now a [PASS], and no existing test regresses.
ratio ≥ 0.60 and ≥ 0.20 over the runner-up) and prints what it picked.insert_decision, so a refusal leaves no decision and no option rows (verified). This mirrors the existing ordering for --question-id and --invalidates.active_index, so the chosen option can never be listed as rejected even when the stored chosen sentence differs from the option label.--chosen is still stored verbatim. The human sentence is preserved; only the active flag is resolved. Callers who want the label itself simply pass the label.docs/VERBS.md, state that --chosen must name exactly one --option (label, D-00X-OY, or index) and that answer now prints N of M active.# Evidence - Problem class: cli-answer-chosen-option-mismatch-silent-inactive-config - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T03:40:42.132Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a CLI that records decisions with options (auger 'answer' verb, auger.py) accepted 'answer --chosen X --option Y' where X and Y were semantically the same decision but different strings. Exit 0, 'D-001 recorded' printed, and a later 'dump' reported ACTIVE CONFIGURATION: (nothing active) with a warning that every decision 'contributes NOTHING to the configuration'. Root cause (auger.py cmd_answer): option rows are inserted with active = (opt == a.chosen) \u2014 a VERBATIM string equality between the --chosen text and the --option text. Any rewording (a chosen sentence vs the option label) stores 0 active options; nothing validates, warns, or refuses, and the success line hides the empty configuration until the user renders it. The fix direction: answer must validate that --chosen matches exactly one --option (or auto-activate the option whose text is closest/dedent), or --chosen should reference the option index; at minimum the success line must report the resulting active-option count. Diagnosed by live dogfood 2026-09-24 (namespace df5-real): 'answer --chosen \"The CI box smoke test\" --option \"CI smoke battery: ...\"' -> D-001 recorded; dump -> '(nothing active)'. Repair workaround in-session: toggle --on <option-id> activates explicitly.", "environment": "auger v0.1 Python stdlib CLI, DuckBrain declared-tables substrate (feat/native-s3), localhost", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "cli-answer-chosen-option-mismatch-silent-inactive-config", "provider": "openrouter", "solved_at": "2026-09-24T03:40:42.132Z", "version": ""}