python-openrouter-reasoning-tokens-starve-json-budget
Solution written to ~/SOLUTION.md, with runnable code in ~/openrouter_fix/.
Root cause: reasoning models spend hidden thinking tokens from the same max_tokens budget as the visible completion. At max_tokens=500, reasoning consumes essentially all of it, so the JSON is empty/truncated and the API returns HTTP 200 with finish_reason="length". Mocks don't model hidden token accounting, so only live calls fail. The judge's tell is the inconsistency: the controller path disabled thinking, the teacher path didn't.
Fix (both halves, because either alone is fragile):
1. Pass extra_body={"thinking": {"type": "disabled"}} on the teacher call.
2. Raise max_tokens 500 → 3000, and retry once with a doubled budget (→6000) on finish_reason="length".
3. Only accept success when finish_reason == "stop" AND content is non-empty AND JSON parses (with fence/prose tolerance).
Verification performed:
- 4 unit tests pass, including a starvation-simulating mock (finish_reason="length", content="") that asserts the wire params and the [3000, 6000] retry budgets.
- probe_live.py is the required live gate: exits 0 only after logging finish_reason=stop content_present=True parse_ok=True on a real OpenRouter call.
Key files:
- ~/SOLUTION.md — full write-up
- ~/openrouter_fix/openrouter_json.py — the fix
- ~/openrouter_fix/test_openrouter_json.py — regression tests
- ~/openrouter_fix/probe_live.py — live probe
# Evidence - Problem class: python-openrouter-reasoning-tokens-starve-json-budget - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-25T09:17:23.584Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "OpenRouter reasoning-model calls (gpt-5.6-luna) that must return parseable JSON fail with finish_reason=length because reasoning tokens consume the max_tokens budget invisibly. Fix pattern proven on ai-plays-poke teacher path: pass thinking={type:disabled} on the call AND raise max_tokens (500->3000 with one doubling retry); verify with a live probe logging finish_reason=stop + content_present. Diagnostic tell: API returns 200 with empty/broken JSON and retries, tests with mocks all pass, only live calls fail. Judge finding quote: controller path already disabled thinking but the teacher path did not.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "python-openrouter-reasoning-tokens-starve-json-budget", "provider": "openrouter", "solved_at": "2026-09-25T09:17:23.584Z", "version": ""}Solution written to ~/SOLUTION.md, with runnable code in ~/openrouter_fix/.
Root cause: reasoning models spend hidden thinking tokens from the same max_tokens budget as the visible completion. At max_tokens=500, reasoning consumes essentially all of it, so the JSON is empty/truncated and the API returns HTTP 200 with finish_reason="length". Mocks don't model hidden token accounting, so only live calls fail. The judge's tell is the inconsistency: the controller path disabled thinking, the teacher path didn't.
Fix (both halves, because either alone is fragile):
1. Pass extra_body={"thinking": {"type": "disabled"}} on the teacher call.
2. Raise max_tokens 500 → 3000, and retry once with a doubled budget (→6000) on finish_reason="length".
3. Only accept success when finish_reason == "stop" AND content is non-empty AND JSON parses (with fence/prose tolerance).
Verification performed:
- 4 unit tests pass, including a starvation-simulating mock (finish_reason="length", content="") that asserts the wire params and the [3000, 6000] retry budgets.
- probe_live.py is the required live gate: exits 0 only after logging finish_reason=stop content_present=True parse_ok=True on a real OpenRouter call.
Key files:
- ~/SOLUTION.md — full write-up
- ~/openrouter_fix/openrouter_json.py — the fix
- ~/openrouter_fix/test_openrouter_json.py — regression tests
- ~/openrouter_fix/probe_live.py — live probe
# Evidence - Problem class: python-openrouter-reasoning-tokens-starve-json-budget - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-25T09:17:23.584Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "OpenRouter reasoning-model calls (gpt-5.6-luna) that must return parseable JSON fail with finish_reason=length because reasoning tokens consume the max_tokens budget invisibly. Fix pattern proven on ai-plays-poke teacher path: pass thinking={type:disabled} on the call AND raise max_tokens (500->3000 with one doubling retry); verify with a live probe logging finish_reason=stop + content_present. Diagnostic tell: API returns 200 with empty/broken JSON and retries, tests with mocks all pass, only live calls fail. Judge finding quote: controller path already disabled thinking but the teacher path did not.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "python-openrouter-reasoning-tokens-starve-json-budget", "provider": "openrouter", "solved_at": "2026-09-25T09:17:23.584Z", "version": ""}