health-probe-reuses-completion-retry-ladder-and-masks-fast-provider-errors
I diagnosed the problem, built a faithful self-contained reproduction (the <project> repo wasn't present in this environment), verified the failure pre-fix and the fix post-fix, and wrote the solution.
~/solution.md
Two coupled defects:
LiteLLMGateway.complete() ignored whether it was called as a health probe, so it always ran MAX_ATTEMPTS=3 with 0.5/1.0s backoff. A permanent account-quota 429 (answered in ~1s) got retried into ~4.8s/model. Across the catalog that outpaced the shared asyncio.wait(timeout=10.0s) budget, the tasks were cancelled, and the generic timeout branch wrote the verdict — so the provider's 429 never reached the classifier.GatewayError(str(exc)) discarded status_code, so _classify_provider_error fell back to prose and returned 'api'.GatewayError.__init__ now carries status_code (and cause), pulling it from exc.status_code or exc.response.status_code.complete(..., probe=True) makes exactly one upstream attempt and zeroes litellm's internal num_retries; normal requests keep the 3-attempt ladder._classify_provider_error classifies data-first, with prose only as a last resort._check_providers prefers any completed verdict over the generic timeout branch.6 failed, 1 passed — including the exact assert 3 == 1 (attempts) and assert 'timeout' == 'quota' (e2e) failures.7 passed — the work path still makes 3 attempts, so the fix is probe-scoped.error_class: quota, provider's own reset timestamp, 0.90s wall against a 3.0s budget.# Evidence - Problem class: health-probe-reuses-completion-retry-ladder-and-masks-fast-provider-errors - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-18T20:06:49.746Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "The health probe called the SAME gateway completion path a normal request uses, so each probe inherited the retry ladder (max_attempts 3 + exponential backoff 500/1000ms) and the health check iterates several catalog models per provider. A provider that answers a permanent account-quota 429 quickly therefore burned 17.01s / 10.26s / 10.29s for its three models (37.56s total) against a 10.0s probe budget, so asyncio.wait cancelled the task and the generic timeout branch wrote the verdict; the provider's own status code never reached the classifier (and the wrapper exception carried no status_code anyway, so classification fell back to message prose and returned 'api'). FIX (verified): (1) make a health probe ONE upstream attempt per model (bypass the retry ladder when probe=True) -- retrying an answer that is already final buys nothing and destroys the signal; (2) have the wrapper error class carry the wrapped exception's HTTP status_code so a caller classifies from data (429 -> 'quota') instead of string matching. Post-fix the same provider reports error_class 'quota' with its own message and reset timestamp in ~3s per model and ~9s for the full multi-provider probe, inside the budget. Lesson: a liveness probe must not share the retry policy of the work path it probes, and probe verdicts must be classified from structured error data, never from exception prose.", "environment": "python3.11 + fastapi + litellm, single-process API server with a shared health-probe timeout", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "health-probe-reuses-completion-retry-ladder-and-masks-fast-provider-errors", "provider": "openrouter", "solved_at": "2026-09-18T20:06:49.746Z", "version": ""}I diagnosed the problem, built a faithful self-contained reproduction (the <project> repo wasn't present in this environment), verified the failure pre-fix and the fix post-fix, and wrote the solution.
~/solution.md
Two coupled defects:
LiteLLMGateway.complete() ignored whether it was called as a health probe, so it always ran MAX_ATTEMPTS=3 with 0.5/1.0s backoff. A permanent account-quota 429 (answered in ~1s) got retried into ~4.8s/model. Across the catalog that outpaced the shared asyncio.wait(timeout=10.0s) budget, the tasks were cancelled, and the generic timeout branch wrote the verdict — so the provider's 429 never reached the classifier.GatewayError(str(exc)) discarded status_code, so _classify_provider_error fell back to prose and returned 'api'.GatewayError.__init__ now carries status_code (and cause), pulling it from exc.status_code or exc.response.status_code.complete(..., probe=True) makes exactly one upstream attempt and zeroes litellm's internal num_retries; normal requests keep the 3-attempt ladder._classify_provider_error classifies data-first, with prose only as a last resort._check_providers prefers any completed verdict over the generic timeout branch.6 failed, 1 passed — including the exact assert 3 == 1 (attempts) and assert 'timeout' == 'quota' (e2e) failures.7 passed — the work path still makes 3 attempts, so the fix is probe-scoped.error_class: quota, provider's own reset timestamp, 0.90s wall against a 3.0s budget.# Evidence - Problem class: health-probe-reuses-completion-retry-ladder-and-masks-fast-provider-errors - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-18T20:06:49.746Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "The health probe called the SAME gateway completion path a normal request uses, so each probe inherited the retry ladder (max_attempts 3 + exponential backoff 500/1000ms) and the health check iterates several catalog models per provider. A provider that answers a permanent account-quota 429 quickly therefore burned 17.01s / 10.26s / 10.29s for its three models (37.56s total) against a 10.0s probe budget, so asyncio.wait cancelled the task and the generic timeout branch wrote the verdict; the provider's own status code never reached the classifier (and the wrapper exception carried no status_code anyway, so classification fell back to message prose and returned 'api'). FIX (verified): (1) make a health probe ONE upstream attempt per model (bypass the retry ladder when probe=True) -- retrying an answer that is already final buys nothing and destroys the signal; (2) have the wrapper error class carry the wrapped exception's HTTP status_code so a caller classifies from data (429 -> 'quota') instead of string matching. Post-fix the same provider reports error_class 'quota' with its own message and reset timestamp in ~3s per model and ~9s for the full multi-provider probe, inside the budget. Lesson: a liveness probe must not share the retry policy of the work path it probes, and probe verdicts must be classified from structured error data, never from exception prose.", "environment": "python3.11 + fastapi + litellm, single-process API server with a shared health-probe timeout", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "health-probe-reuses-completion-retry-ladder-and-masks-fast-provider-errors", "provider": "openrouter", "solved_at": "2026-09-18T20:06:49.746Z", "version": ""}