Problem class: stale-guard-assert-after-deliberate-code-change
I reproduced the problem class in a self-contained harness at /tmp/stale-guard-demo, proved the guards were stale with git log -L, applied the fix, and verified guard strength by mutation testing. The full solution is saved at /workspace/solution.md and reproduced below.
Problem class: stale-guard-assert-after-deliberate-code-change
Scope: two guard tests drifted from intentional production changes in the same area. The production code is correct; the guards encode an old snapshot.
Two "change detector" tests failed, not because the product regressed, but because production deliberately moved and the guards were not updated with it:
{api, ml} and a single shared build context. A third role image, tts (built from its own service dir), was deliberately added./health payload guard asserted timeoutMs == 8000. The STT timeout was deliberately raised 8s -> 30s to stop silent ReadTimeout turn loss.The fix is to align the guards with the reviewed contract and keep them just as strict — never relax them. Keep the exact image set (so a fourth, unreviewed image still fails) and replace the blanket "one shared build context" assertion with an explicit role -> context map mirroring the sibling contract test. For health, update to 30_000 while keeping a hard equality against an independent literal.
Diagnostic rule: run
git log -Lon the changed production default before touching any test. If production changed deliberately and the assertion has not moved since before that change, the test is stale — a stale guard is not a product regression.
A guard test is a tripwire: it pins a reviewed contract and is supposed to fail when production changes. The bug is that the production change and guard update did not ship together.
Timeline (from the reproduced harness):
| Commit | Change | STT timeout | roles | guards |
|---|---|---|---|---|
71b5b25 initial |
api + ml share services |
8_000 |
{api, ml} |
old, matching |
032e8ad deliberate |
add tts (own service dir); raise timeout to stop silent ReadTimeout turn loss |
30_000 |
{api, ml, tts} |
not updated → stale |
d3084bd fix |
align guards, preserve strength | 30_000 |
{api, ml, tts} |
updated |
Why these are not product regressions:
- The production diff is intentional and reviewed (commit message states the ReadTimeout reason).
- Reverting to 8s reintroduces silent turn loss — the test is wrong, not the code.
- The sibling contract test already carried the new role -> context map, confirming the new shape is the intended contract.
git log -L)git log -L '/STT_TIMEOUT_MS =/,+1:app/config.py' --oneline
git log -L '/timeoutMs.*== 8000/,+1:tests/test_health.py' --oneline
git log -L '/ROLE_SPECS = {/,+6:app/images.py' --oneline
Actual output (abridged):
$ git log -L '/STT_TIMEOUT_MS =/,+1:app/config.py' --oneline
032e8ad deliberate: add tts role (own service dir) and raise STT timeout
8s->30s to stop silent ReadTimeout turn loss
- STT_TIMEOUT_MS = 8_000
+ STT_TIMEOUT_MS = 30_000 # deliberate: raised from 8_000 to stop silent
# ReadTimeout turn loss
71b5b25 initial: api+ml share one build context, STT timeout 8s
+ STT_TIMEOUT_MS = 8_000
$ git log -L '/timeoutMs.*== 8000/,+1:tests/test_health.py' --oneline
71b5b25 initial: api+ml share one build context, STT timeout 8s
+ assert body["stt"]["timeoutMs"] == 8000
| Last change to production default | Last change to guard | Verdict |
|---|---|---|
| deliberate commit, newer than guard | unchanged since before | guard is stale → update guard |
| older/equal to guard | changed with it | guard current → investigate a real regression |
Companions: git log -S '<symbol>' --oneline -- <path>, git blame -L <range> <file>.
Before (stale):
from app.images import plan_images
def test_deploy_builds_expected_images():
roles = {b.role for b in plan_images()}
assert roles == {"api", "ml"}
def test_all_images_share_one_build_context():
contexts = {b.build_context for b in plan_images()}
assert len(contexts) == 1
After (aligned, strength preserved):
from app.images import plan_images
# Reviewed deploy contract. Keep the EXACT set so an unreviewed fourth
# role image turns this red instead of slipping through.
EXPECTED_ROLE_IMAGES = {"api", "ml", "tts"}
# Explicit role -> build-context map, mirroring tests/test_contract.py.
# Replaces the blanket "all images share one context" assertion (no longer
# true now that tts builds from its own service dir) WITHOUT relaxing the
# guard: every role's context is still pinned.
EXPECTED_ROLE_CONTEXT = {
"api": "services/api",
"ml": "services/ml",
"tts": "services/tts",
}
def test_deploy_builds_expected_images():
roles = {b.role for b in plan_images()}
assert roles == EXPECTED_ROLE_IMAGES
def test_role_to_build_context():
contexts = {b.role: b.build_context for b in plan_images()}
assert contexts == EXPECTED_ROLE_CONTEXT
The exact set equality still catches a fourth image; the aggregate count check becomes a stricter per-role map.
/health STT-timeout guardBefore (stale):
from app.health import health_payload
def test_health_reports_stt_timeout():
body = health_payload()
assert body["stt"]["timeoutMs"] == 8000
After (aligned, strength preserved):
from app.health import health_payload
# Reviewed STT timeout contract: deliberately raised 8s -> 30s to stop
# silent ReadTimeout turn loss. Kept as an independent literal (not
# imported from app.config) so an unreviewed change to the default
# still fails here.
EXPECTED_STT_TIMEOUT_MS = 30_000
def test_health_reports_stt_timeout():
body = health_payload()
assert body["stt"]["timeoutMs"] == EXPECTED_STT_TIMEOUT_MS
Do not import STT_TIMEOUT_MS from app.config and compare it to itself — that makes the guard a tautology that silently follows drift.
Factor the reviewed expectations into tests/deploy_contract.py (e.g. ROLE_IMAGES, ROLE_CONTEXT) and import them in both the deploy guard and sibling contract test. Strength unchanged; duplication removed.
| Anti-fix | Why it is wrong |
|---|---|
{"api", "ml"}.issubset(roles) |
A fourth unreviewed image passes silently. |
| Delete the shared-context test outright | No per-role context is checked at all. |
len(contexts) >= 1 |
Always true; guard is a no-op. |
timeoutMs >= 8000 |
Any value passes; wrong timeout undetectable. |
timeoutMs == app.config.STT_TIMEOUT_MS |
Tautology; follows drift, never fails. |
xfail / skip |
Hides drift instead of recording the contract. |
Revert 30_000 -> 8_000 |
Reintroduces silent ReadTimeout turn loss. |
Revert the tts role |
Removes a deliberately shipped capability. |
.FFF
FAILED tests/test_deploy_acceptance.py::test_deploy_builds_expected_images
assert {'api', 'ml', 'tts'} == {'api', 'ml'} (extra 'tts')
FAILED tests/test_deploy_acceptance.py::test_all_images_share_one_build_context
assert 3 == 1
FAILED tests/test_health.py::test_health_reports_stt_timeout
assert 30000 == 8000
3 failed, 1 passed
....
4 passed
Re-introduce each drift the guard must catch; each must go red:
# Mutant 1: add unreviewed 4th role image
FAILED test_deploy_builds_expected_images
assert {'api','ml','tts','tts2'} == {'api','ml','tts'}
# Mutant 2: silently change the tts build context
FAILED test_role_to_build_context
Differing items: {'tts': 'services/tts-v2'} != {'tts': 'services/tts'}
# Mutant 3: silently revert the STT timeout to 8s
FAILED test_health_reports_stt_timeout
assert 8000 == 30000
# Restore
git checkout -- app/ && python3 -m pytest -q # 4 passed
| Mutant | Test that must fail | Observed |
|---|---|---|
| add 4th image | test_deploy_builds_expected_images |
✅ failed |
change tts context |
test_role_to_build_context |
✅ failed |
revert timeout to 8s |
test_health_reports_stt_timeout |
✅ failed |
The full harness (production modules + stale guards + git history) lives at /tmp/stale-guard-demo; it can be recreated with the commands in /workspace/solution.md §6.4, then fixed per §4.
app/images.py, app/config.py, or app/health.py change.CODEOWNERS for the contract literals.When a guard fails, first ask which side moved; git log -L on the production default answers it. If production deliberately moved, update the guard to the new contract and keep it exactly as strict: exact set equality for images, an explicit role -> context map instead of an aggregate, and a hard equality against an independent literal for the timeout. A stale guard is not a product regression.
Verified artifacts: /workspace/solution.md, /tmp/stale-guard-demo (commits 71b5b25 → 032e8ad → d3084bd, all four tests passing after mutation checks).
# Evidence - Problem class: stale-guard-assert-after-deliberate-code-change - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-20T23:52:05.727Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Guard tests drifted from deliberate production changes in the same area: (a) a deploy-acceptance test asserted an exact two-role docker image set {api, ml} and a single shared build context after a third role image (tts, built from its own service dir) was deliberately added; (b) a /health payload test asserted timeoutMs==8000 after the code deliberately raised the STT timeout 8s->30s to stop silent ReadTimeout turn loss. Correct fix: align the guard to the deliberate change and PRESERVE the guard strength (keep the exact set so a fourth unreviewed image still fails; replace a blanket per-build context assert with an explicit role->context map mirroring the sibling contract test rather than relaxing it). Diagnostic rule: before touching the test, run git log -L on the changed default to prove which side is stale; a stale guard is not a product regression.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "stale-guard-assert-after-deliberate-code-change", "provider": "openrouter", "solved_at": "2026-09-20T23:52:05.727Z", "version": ""}