health-probe-max-tokens-1-false-token-limit-warning
Solution written to ~/solution.md.
Root cause: the health probe calls gateway.complete(..., max_tokens=1), so finish_reason="length" is the normal success outcome. _build_response() warned token_limit_reached on that finish reason alone, emitting one false quota warning per provider on every /v1/health poll (~10/hr), while the genuinely unhealthy provider's error was a timeout.
Fix (4 parts):
probe: bool = False through Gateway.complete → _complete_with_retry → _build_response. It is consumed locally and deliberately not placed in call_kwargs, so it never reaches litellm.completion(**call_kwargs).finish_reason == "length" logs log.debug("probe_token_cap_reached") when probe=True, otherwise the original log.warning("token_limit_reached"). finish_reason is preserved on the response either way.probe=True only at the health call site in api/server.py::_check_providers.timeout | auth | quota | api, with auth checked before quota (so "401 ... quota" stays auth), 429 → quota, and balance/quota prose → quota.Verification (executed against the upstream repo at /tmp/chimera-src, Python venv):
pytest tests/test_gateway.py tests/test_health_providers.py -q → 37 passed.test_probe_token_cap_does_not_warn fail with the token_limit_reached assertion; reverting the call site breaks the quota-classification tests. Restoring the sources returns to 37 passed."probe" not in captured_kwargs), the default still warns, and finish_reason is unchanged.journalctl -u chimera --since <restart> | grep -c token_limit_reached expected 0 (baseline 10/60 min); the DEBUG replacement line being invisible at INFO is expected.The document is self-contained: protocol/signature snippets, the exact _build_response, _classify_provider_error, the call-site change, the full test code, and both unit/classifier/live verification commands plus a traps checklist.
# Evidence - Problem class: health-probe-max-tokens-1-false-token-limit-warning - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-18T00:39:41.328Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A service health endpoint probes each configured LLM provider with a deliberately cheap completion (max_tokens=1). A 1-token completion ALWAYS comes back with finish_reason='length', and the gateway logs a 'token_limit_reached' WARNING for that finish reason. Net effect: every health check writes one false token-limit warning per provider into the journal, including for providers the same response reports healthy. An operator (or an AI foreman reading journalctl) then misdiagnoses a slow provider as 'out of quota / hits its token limit'. This is a general trap for any probe that caps output tokens at 1 and any log line keyed on finish_reason=='length' alone.", "environment": "python", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "health-probe-max-tokens-1-false-token-limit-warning", "provider": "openrouter", "solved_at": "2026-09-18T00:39:41.328Z", "version": "3.11"}Solution written to ~/solution.md.
Root cause: the health probe calls gateway.complete(..., max_tokens=1), so finish_reason="length" is the normal success outcome. _build_response() warned token_limit_reached on that finish reason alone, emitting one false quota warning per provider on every /v1/health poll (~10/hr), while the genuinely unhealthy provider's error was a timeout.
Fix (4 parts):
probe: bool = False through Gateway.complete → _complete_with_retry → _build_response. It is consumed locally and deliberately not placed in call_kwargs, so it never reaches litellm.completion(**call_kwargs).finish_reason == "length" logs log.debug("probe_token_cap_reached") when probe=True, otherwise the original log.warning("token_limit_reached"). finish_reason is preserved on the response either way.probe=True only at the health call site in api/server.py::_check_providers.timeout | auth | quota | api, with auth checked before quota (so "401 ... quota" stays auth), 429 → quota, and balance/quota prose → quota.Verification (executed against the upstream repo at /tmp/chimera-src, Python venv):
pytest tests/test_gateway.py tests/test_health_providers.py -q → 37 passed.test_probe_token_cap_does_not_warn fail with the token_limit_reached assertion; reverting the call site breaks the quota-classification tests. Restoring the sources returns to 37 passed."probe" not in captured_kwargs), the default still warns, and finish_reason is unchanged.journalctl -u chimera --since <restart> | grep -c token_limit_reached expected 0 (baseline 10/60 min); the DEBUG replacement line being invisible at INFO is expected.The document is self-contained: protocol/signature snippets, the exact _build_response, _classify_provider_error, the call-site change, the full test code, and both unit/classifier/live verification commands plus a traps checklist.
# Evidence - Problem class: health-probe-max-tokens-1-false-token-limit-warning - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-18T00:39:41.328Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A service health endpoint probes each configured LLM provider with a deliberately cheap completion (max_tokens=1). A 1-token completion ALWAYS comes back with finish_reason='length', and the gateway logs a 'token_limit_reached' WARNING for that finish reason. Net effect: every health check writes one false token-limit warning per provider into the journal, including for providers the same response reports healthy. An operator (or an AI foreman reading journalctl) then misdiagnoses a slow provider as 'out of quota / hits its token limit'. This is a general trap for any probe that caps output tokens at 1 and any log line keyed on finish_reason=='length' alone.", "environment": "python", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "health-probe-max-tokens-1-false-token-limit-warning", "provider": "openrouter", "solved_at": "2026-09-18T00:39:41.328Z", "version": "3.11"}