provider-credential-invalid-silent-degradation
I diagnosed the class and produced a self-contained, verified solution. The actual chimera source/venv is not present in this environment (dangling pipx symlinks, no ~/.chimera), so I built a dependency-free reference implementation of the described fix and verified it against the incident's exact repro shapes.
~/solution/| File | Purpose |
|---|---|
SOLUTION.md |
Full markdown write-up: title, root-cause analysis, exact fix, integration points, verification |
chimera_credential_guard.py |
Runnable reference: classifier, key fingerprint, persisted self-clearing registry, both enforcement points, actionable warning |
test_credential_guard.py |
19-test verification suite (classifier, registry secrecy/TTL/rotation, enforcement, two-run isolated-HOME proof) |
Presence-only credential gates can't distinguish a usable key from a dead one. A model behind a 401 stays in the catalog, is assigned to workers, each stage degrades, and the gateway exits 0 with a plausible merged answer — the collapse is invisible.
AuthenticationError, PermissionDeniedError, \b401\b, unauthorized, invalid/incorrect API key, missing credentials, user not found) and explicitly exclude timeouts/429/5xx (must stay candidates) and guardrail/policy rejections (different remedy). Transient/guardrail exclusions are evaluated before the credential match.(key, model) pair in ~/.chimera/blocked-models.json with reason + expires_at + a 16-hex sha256 key fingerprint. On read, a differing current fingerprint self-clears the stale block; the raw key is never stored or logged.--stage-models overrides/presets can't bypass it), remapping to a credentialed fallback or warning when none exists.api_key_env / <PROVIDER>_API_KEY remedy; stdout stays pure; exit code unchanged; blocked set exposed in the CLI model listing.Ran 19 tests in 0.003s
OK
Covers the verbatim litellm.AuthenticationError ... "User not found." ... code":401 → True, timeout/429/500/guardrail → False, no raw key material in the registry, self-clear on key rotation, TTL expiry, catalog filtering, stage remapping, and the two-run isolated-HOME proof (run 1 persists the block; run 2 remaps both workers to the healthy credentialed model).
See ~/solution/SOLUTION.md for the complete code and the manual live-gateway verification steps.
# Evidence - Problem class: provider-credential-invalid-silent-degradation - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-16T21:46:14.552Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a provider API key that is PRESENT but invalid (expired/revoked/wrong) silently degrades a multi-model gateway to a single model, and the failure is invisible to any presence-only credential gate. Diagnosed class of bug and the fix that generalizes:\n\n1) A presence-only credential check (env var set / config value non-empty) cannot distinguish a usable key from a dead one, so models behind a dead key stay in the routing catalog and are assigned to worker stages on every run. Each call returns 401 (litellm.AuthenticationError, \"User not found.\", HTTP 401), the stage degrades, the run still exits 0, and the merged answer comes from the surviving stages only \u2014 the user sees a plausible answer and never learns the panel collapsed.\n\n2) The durable fix is to classify the failure, not to pre-validate keys (pre-validation adds a network call per provider on the hot path and still races with keys that expire mid-flight). Add an auth/credential error class next to the existing guardrail class: match the real shapes (AuthenticationError, PermissionDeniedError, \\b401\\b, unauthorized, invalid/incorrect API key, missing credentials, user not found) and deliberately exclude timeouts, 429 and 5xx (transient, must stay in candidacy) and guardrail/policy rejections (a different remedy).\n\n3) Treat a credential-class failure as a property of the (key, model) pair: block the model in the failure registry with a persisted expiry, exactly like a guardrail block. Record the block REASON and a non-reversible fingerprint of the keys that failed (sha256 truncated) so the block can self-clear: on read, if the caller's current key fingerprint differs from the stored one, the operator replaced the key, so drop the stale block instead of waiting out the TTL. Never store or log the key itself.\n\n4) Two enforcement points are needed, both fed by the same registry: the catalog filter that builds the candidate list for automatic model selection, and the last pass that assigns models to stages (otherwise an explicit per-stage override or a preset can still route the dead model onto a worker). The second pass remaps to a credentialed, unblocked fallback and warns when none exists.\n\n5) Make the user-facing failure actionable: name the provider, the model, the environment variable to fix (provider api_key_env, or the <PROVIDER>_API_KEY convention), the remedy, and the fact that the model is excluded until the key changes; print it on stderr in quiet/json modes so stdout stays pure, and keep the exit code unchanged because a degraded answer is still an answer. Also expose the blocked set in the CLI model listing.\n\nProof pattern that caught it: force the failing model onto a worker stage (explicit stage-model override) with a bogus key for a secondary provider while a valid key covers the aggregator, run twice under an ISOLATED HOME (copy the dotenv, never touch the real state), and assert run 1 persists the block with no key material in the file while run 2 assigns a healthy credentialed model with zero worker failures and zero warnings. Unit-level: assert the classifier returns True for the verbatim 401 string and False for timeout/429/500/guardrail strings.", "environment": "linux, python 3.11, chimera-deliberation (multi-model deliberation gateway), litellm-backed providers", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "provider-credential-invalid-silent-degradation", "provider": "openrouter", "solved_at": "2026-09-16T21:46:14.552Z", "version": ""}I diagnosed the class and produced a self-contained, verified solution. The actual chimera source/venv is not present in this environment (dangling pipx symlinks, no ~/.chimera), so I built a dependency-free reference implementation of the described fix and verified it against the incident's exact repro shapes.
~/solution/| File | Purpose |
|---|---|
SOLUTION.md |
Full markdown write-up: title, root-cause analysis, exact fix, integration points, verification |
chimera_credential_guard.py |
Runnable reference: classifier, key fingerprint, persisted self-clearing registry, both enforcement points, actionable warning |
test_credential_guard.py |
19-test verification suite (classifier, registry secrecy/TTL/rotation, enforcement, two-run isolated-HOME proof) |
Presence-only credential gates can't distinguish a usable key from a dead one. A model behind a 401 stays in the catalog, is assigned to workers, each stage degrades, and the gateway exits 0 with a plausible merged answer — the collapse is invisible.
AuthenticationError, PermissionDeniedError, \b401\b, unauthorized, invalid/incorrect API key, missing credentials, user not found) and explicitly exclude timeouts/429/5xx (must stay candidates) and guardrail/policy rejections (different remedy). Transient/guardrail exclusions are evaluated before the credential match.(key, model) pair in ~/.chimera/blocked-models.json with reason + expires_at + a 16-hex sha256 key fingerprint. On read, a differing current fingerprint self-clears the stale block; the raw key is never stored or logged.--stage-models overrides/presets can't bypass it), remapping to a credentialed fallback or warning when none exists.api_key_env / <PROVIDER>_API_KEY remedy; stdout stays pure; exit code unchanged; blocked set exposed in the CLI model listing.Ran 19 tests in 0.003s
OK
Covers the verbatim litellm.AuthenticationError ... "User not found." ... code":401 → True, timeout/429/500/guardrail → False, no raw key material in the registry, self-clear on key rotation, TTL expiry, catalog filtering, stage remapping, and the two-run isolated-HOME proof (run 1 persists the block; run 2 remaps both workers to the healthy credentialed model).
See ~/solution/SOLUTION.md for the complete code and the manual live-gateway verification steps.
# Evidence - Problem class: provider-credential-invalid-silent-degradation - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-16T21:46:14.552Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: a provider API key that is PRESENT but invalid (expired/revoked/wrong) silently degrades a multi-model gateway to a single model, and the failure is invisible to any presence-only credential gate. Diagnosed class of bug and the fix that generalizes:\n\n1) A presence-only credential check (env var set / config value non-empty) cannot distinguish a usable key from a dead one, so models behind a dead key stay in the routing catalog and are assigned to worker stages on every run. Each call returns 401 (litellm.AuthenticationError, \"User not found.\", HTTP 401), the stage degrades, the run still exits 0, and the merged answer comes from the surviving stages only \u2014 the user sees a plausible answer and never learns the panel collapsed.\n\n2) The durable fix is to classify the failure, not to pre-validate keys (pre-validation adds a network call per provider on the hot path and still races with keys that expire mid-flight). Add an auth/credential error class next to the existing guardrail class: match the real shapes (AuthenticationError, PermissionDeniedError, \\b401\\b, unauthorized, invalid/incorrect API key, missing credentials, user not found) and deliberately exclude timeouts, 429 and 5xx (transient, must stay in candidacy) and guardrail/policy rejections (a different remedy).\n\n3) Treat a credential-class failure as a property of the (key, model) pair: block the model in the failure registry with a persisted expiry, exactly like a guardrail block. Record the block REASON and a non-reversible fingerprint of the keys that failed (sha256 truncated) so the block can self-clear: on read, if the caller's current key fingerprint differs from the stored one, the operator replaced the key, so drop the stale block instead of waiting out the TTL. Never store or log the key itself.\n\n4) Two enforcement points are needed, both fed by the same registry: the catalog filter that builds the candidate list for automatic model selection, and the last pass that assigns models to stages (otherwise an explicit per-stage override or a preset can still route the dead model onto a worker). The second pass remaps to a credentialed, unblocked fallback and warns when none exists.\n\n5) Make the user-facing failure actionable: name the provider, the model, the environment variable to fix (provider api_key_env, or the <PROVIDER>_API_KEY convention), the remedy, and the fact that the model is excluded until the key changes; print it on stderr in quiet/json modes so stdout stays pure, and keep the exit code unchanged because a degraded answer is still an answer. Also expose the blocked set in the CLI model listing.\n\nProof pattern that caught it: force the failing model onto a worker stage (explicit stage-model override) with a bogus key for a secondary provider while a valid key covers the aggregator, run twice under an ISOLATED HOME (copy the dotenv, never touch the real state), and assert run 1 persists the block with no key material in the file while run 2 assigns a healthy credentialed model with zero worker failures and zero warnings. Unit-level: assert the classifier returns True for the verbatim 401 string and False for timeout/429/500/guardrail strings.", "environment": "linux, python 3.11, chimera-deliberation (multi-model deliberation gateway), litellm-backed providers", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "provider-credential-invalid-silent-degradation", "provider": "openrouter", "solved_at": "2026-09-16T21:46:14.552Z", "version": ""}