Problem class: confidence-sentinel-leaks-into-user-facing-render
Environment note: no application repository was present (/workspace contained only problem.json), so I built a self-contained reference implementation matching the described bug shape, verified it, and wrote the deliverable to /workspace/solution.md (alongside confidence.py and test_confidence.py).
Verification: 7 passed in 0.01s.
Below is the solution.
Problem class: confidence-sentinel-leaks-into-user-facing-render
Confidence is stored with a sentinel value -1.0 meaning "asserted directly, never scored". The store is allowed to keep that sentinel, but four reader sites treated -1.0 as an ordinary number, so unmeasured decisions were rendered and judged as if they had a real score. The fix is a single render guard (_conf_render) plus explicit sentinel handling in the aggregate, feedback, thin-decision, and priority paths. The sentinel itself is not changed in storage.
The store-level sentinel is a storage contract, but nothing enforced it at the read boundary. Each reader re-implemented its own (wrong) handling:
| Site | Buggy behavior | Why it happened |
|---|---|---|
dump |
printed conf=-1.00 |
formatted the raw float |
| answer echo | repeated conf -1.00 |
formatted the raw float |
feedback |
used 0 <= c < T to find low scores; -1.0 fails that test, so it fell into the "confident" bucket |
sentinel is negative, so it is neither in [0, T) nor recognized as unmeasured |
| priority judgment | c < 0.30 was true for -1.00, so it escalated |
negative sentinel is numerically below the floor |
| averages | included -1.0, dragging means down |
no filter before sum/len |
thin_decisions |
-1.0 < threshold, but ordering/composition was accidental |
sentinel not modeled as "unmeasured" |
Concrete pre-fix evidence:
dump: b conf=0.10 | z conf=-1.00 | c conf=0.90
answer echo: Z (conf -1.00)
feedback confident ids: ['z', 'c'] # z was never scored
priority z: escalate # -1.00 < 0.30
Lesson: a store-level sentinel needs a render-level guard at every reader site, not ad-hoc filters. Centralize the guard, then make each semantic site (aggregate, decision, feedback) explicitly drop or classify the sentinel.
All user-facing rendering goes through one helper; all numeric logic filters through _is_measured. The sentinel remains in the store.
"""confidence.py"""
UNMEASURED = -1.0
THIN_THRESHOLD = 0.30
def _is_measured(c):
"""A confidence is measured only if it is a real, non-negative score."""
return c is not None and c >= 0.0
def _conf_render(c):
"""Single render guard for every user-facing surface.
Sentinel (or missing) -> 'n/a'; a real score -> fixed 2-dp string.
This is the ONE place that knows the sentinel is not a number to show.
"""
if not _is_measured(c):
return "n/a"
return f"{c:.2f}"
def _conf_values(records):
"""Every average must be built from this: measured scores only."""
return [r["confidence"] for r in records if _is_measured(r["confidence"])]
def average_confidence(records):
values = _conf_values(records)
if not values:
return None
return sum(values) / len(values)
def thin_decisions(records, threshold=THIN_THRESHOLD):
"""Thin = measured below threshold, then unmeasured.
Order: measured ascending by (confidence, id), then unmeasured ascending by id.
Unmeasured is thin by definition — it cannot be assumed confident.
"""
measured = [r for r in records if _is_measured(r["confidence"]) and r["confidence"] < threshold]
unmeasured = [r for r in records if not _is_measured(r["confidence"])]
measured.sort(key=lambda r: (r["confidence"], r["id"]))
unmeasured.sort(key=lambda r: r["id"])
return measured + unmeasured
def feedback(records, threshold=THIN_THRESHOLD):
"""Confident = measured AND >= threshold. Unmeasured is surfaced separately."""
confident = [r for r in records if _is_measured(r["confidence"]) and r["confidence"] >= threshold]
unmeasured = [r for r in records if not _is_measured(r["confidence"])]
return {"confident": confident, "unmeasured": unmeasured}
def priority(record):
"""Numeric escalation only applies to measured rows below the floor."""
c = record["confidence"]
if _is_measured(c) and c < THIN_THRESHOLD:
return "escalate"
return "normal"
def dump(records):
return "\n".join(f'{r["id"]} conf={_conf_render(r["confidence"])}' for r in records)
def answer_echo(record):
return f'{record["answer"]} (conf {_conf_render(record["confidence"])})'
Key properties:
_conf_render is the only formatter; sentinel → n/a._is_measured is reused everywhere so None and negative sentinels are handled identically.STORE[2]["confidence"] == -1.0 still holds; only the render/decision layer changes.thin_decisions appends unmeasured rows after the measured-below-threshold rows, ordered by id.priority returns normal.needs-score bucket, never the confident bucket.Self-contained test file (run with pytest):
"""test_confidence.py"""
import pytest
from confidence import (
UNMEASURED, answer_echo, average_confidence,
dump, feedback, priority, thin_decisions,
)
STORE = [
{"id": "b", "confidence": 0.10, "answer": "B"},
{"id": "a", "confidence": 0.10, "answer": "A"},
{"id": "z", "confidence": UNMEASURED, "answer": "Z"}, # asserted directly
{"id": "c", "confidence": 0.90, "answer": "C"},
{"id": "y", "confidence": None, "answer": "Y"}, # also unmeasured
{"id": "d", "confidence": 0.30, "answer": "D"},
]
def test_store_sentinel_is_untouched():
assert STORE[2]["confidence"] == -1.0
assert STORE[4]["confidence"] is None
def test_dump_renders_sentinel_as_na():
out = dump(STORE)
assert "-1.00" not in out
assert "z conf=n/a" in out
def test_answer_echo_renders_sentinel_as_na():
assert answer_echo(STORE[2]) == "Z (conf n/a)"
assert answer_echo(STORE[4]) == "Y (conf n/a)"
assert answer_echo(STORE[0]) == "B (conf 0.10)"
def test_conf_values_drops_sentinel_from_average():
# measured 0.10,0.10,0.90,0.30 -> 0.35; sentinel/None excluded
assert average_confidence(STORE) == pytest.approx(0.35)
assert average_confidence([{"confidence": UNMEASURED}]) is None
def test_feedback_does_not_call_unmeasured_confident():
result = feedback(STORE, threshold=0.30)
confident_ids = {r["id"] for r in result["confident"]}
unmeasured_ids = {r["id"] for r in result["unmeasured"]}
assert confident_ids == {"c", "d"}
assert unmeasured_ids == {"z", "y"}
assert confident_ids.isdisjoint(unmeasured_ids)
def test_priority_skips_unmeasured_escalation():
assert priority({"id": "z", "confidence": UNMEASURED}) == "normal"
assert priority({"id": "y", "confidence": None}) == "normal"
assert priority({"id": "b", "confidence": 0.10}) == "escalate"
assert priority({"id": "d", "confidence": 0.30}) == "normal"
def test_thin_decisions_ordering():
ordered = [r["id"] for r in thin_decisions(STORE, threshold=0.30)]
assert ordered == ["a", "b", "y", "z"]
Observed result:
$ python -m pytest -q
....... [100%]
7 passed in 0.01s
Post-fix behavior:
dump: b conf=0.10 | z conf=n/a | c conf=0.90
answer echo: Z (conf n/a)
feedback confident: ['c'] # z is unmeasured, not confident
feedback unmeasured: ['z']
priority z: normal # no numeric escalation
_conf_render and route every user-facing formatter through it (dump, answer echo, any report/UI/CLI output)._is_measured and use it in every numeric path: averages (_conf_values), thresholds, sorting, escalation.thin_decisions) and feedback; never let a negative sentinel fall through a numeric comparison.confidence formatting and comparisons to prevent new leaks:bash
grep -rn "confidence" --include='*.py' . | grep -E "\.2f|<|>=|<=|sum\("
# Evidence - Problem class: confidence-sentinel-leaks-into-user-facing-render - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-25T04:20:44.526Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A stored sentinel confidence (-1 = asserted directly, not scored) leaked into user-facing renders: dump printed a bare conf -1.0, the answer echo repeated it, feedback 0<=c<T filter treated unmeasured as CONFIDENT (reporting every decision at/above threshold while one was never scored), and priority judgment escalated it as -1.00 below the 0.30 floor. Fix shape: one _conf_render helper (sentinel->n/a) at every user-facing render site; _conf_values drops the sentinel from every average; thin_decisions treats unmeasured as thin (ordered measured-by-confidence-then-id, then unmeasured by id); numeric priority escalation skipped for unmeasured rows; sentinel stays in the STORE unchanged. Lesson: a store-level sentinel needs a render-level guard at EVERY reader site, not ad-hoc filters.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "confidence-sentinel-leaks-into-user-facing-render", "provider": "openrouter", "solved_at": "2026-09-25T04:20:44.526Z", "version": ""}