provider-swallows-error-mock-fallback-never-reached
Done. The environment had no repository (only /workspace/problem.json), so I built a self-contained, runnable reproduction of the described Amadeus provider paths and verified the fix on Node 22 before writing the solution.
/workspace/solution/SOLUTION.md — the requested self-contained writeup: title, root-cause analysis (RC1 swallowing, RC2 unbounded auth fetch, RC3 no total budget), exact TypeScript fix for all five files, and a verification section./workspace/solution/harness/ — executable proof, matching the real filenames:amadeus-http-client.mjs — total budget + per-attempt clamp + AMAD_BUDGET_EXCEEDED/AMAD_ABORTEDamadeus-auth.mjs — per-attempt AbortController + authTimeoutMsamadeus-search.mjs — swallowing search vs non-swallowing searchStrictamadeus-providers.mjs — reachable mock fallback + {reason, code, attempts, elapsedMs} diagnosticsverify.mjs, real-blackhole.mjs, mutation.mjsPASS 1503ms budget caps total time and fallback result is returned
PASS 3656ms MUTATION: budget disabled overruns the cap
PASS 3003ms search swallows but searchStrict reaches fallback
PASS 302ms auth fetch is bounded by authTimeoutMs
PASS 50ms caller signal yields AMAD_ABORTED quickly
Real blackhole <ip-address>:9999 with fallbackBudgetMs: 1500 → elapsedMs: 1502, provider: mock, code: AMAD_BUDGET_EXCEEDED.
Mutation proof: budget enabled 302ms vs disabled 3356ms — so an elapsed <= 800 assertion fails with expected 3356 to be less than or equal to 800 when the cap is removed, demonstrating the timing assertion (not just result shape) guards the budget.
The key correctness point is captured throughout: the fallback test asserts both the mock result shape and the measured elapsed time vs budget, avoiding the "mock data returned passes even with no budget" pitfall.
# Evidence - Problem class: provider-swallows-error-mock-fallback-never-reached - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-18T20:31:10.615Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "CLASS: a provider integration whose 'graceful fallback' is dead code because an inner wrapper swallows the transport error, and whose auth/token fetch has no timeout at all.\n\nSYMPTOM (measured, blackholed endpoint <ip-address>:9999 where SYN is dropped): fetchAvailability() returned status=empty / provider=amadeus after 20991ms and NEVER reached its mock fallback; the raw HTTP client took 17514ms (4 attempts x 4000ms abort + backoff) and the OAuth token fetch 20974ms (2 attempts x the OS connect timeout, no AbortController, no signal). Under a cheaper failure mode (bogus proxy / connect refused) the same path burned 9.5s inside a unit test.\n\nROOT CAUSE 1 (masking): the provider's search() wrapped the HTTP call in `catch { return [] }`, so the error never reached the caller's catch/fallback branch - the orchestrator's `catch (err) { breaker.onFailure(); return mock }` was unreachable for that provider, and a total provider outage looked like an honest empty result. Any 'our fallback is tested' claim is false when the inner layer swallows: the fallback test must assert the mock/fallback RESULT, not just 'no crash'.\nROOT CAUSE 2 (no bound): the token/auth fetch used globalThis.fetch with no timeout and no signal, so a blackholed peer hangs for the OS connect timeout (2 x ~10s) before anything can even classify the failure.\nROOT CAUSE 3 (no total budget): retry loops that bound per-attempt time but not TOTAL elapsed time multiply the timeout by the attempt count (4 x 4s + backoff). A caller-side budget that is not passed down cannot stop an in-flight attempt.\n\nFIX THAT WORKED: (a) a total wall-clock budget on the HTTP client covering attempts + backoff, with the per-attempt timeout clamped to the remaining budget and a distinct structured code (AMAD_BUDGET_EXCEEDED) when exhausted; (b) a per-attempt AbortController on the auth fetch plus an optional caller signal; (c) a non-swallowing strict variant of the provider call (searchStrict) so a transport failure reaches the caller's catch; (d) the caller's AbortSignal threaded end-to-end (AMAD_ABORTED); (e) config knobs (request timeout / max retries / fallback budget / auth timeout) so the degraded-mode budget is tunable; (f) additive fallback diagnostics ({reason, code, attempts, elapsedMs}) so the fallback is visible in results.\n\nVERIFICATION RECIPE (cheap, no fixtures): point the provider base URL at an unroutable/blackholed address (TEST-NET-1 <ip-address>:9999 - SYN dropped, so you measure timeouts rather than instants refusals), time the public entry point, and assert both the latency bound and the fallback RESULT SHAPE. Then run the suite with HTTPS_PROXY/HTTP_PROXY/ALL_PROXY=<ip-address>:9 to prove no test depends on outbound network: any test that really hits the vendor host is the flaky, slow one - replace it with an injected fetcher (hanging / failing / healthy) and add an injection seam instead of mocking the module. Finally, mutation-check the cap: disable the budget and confirm the new test FAILS (assert on elapsed vs budget, e.g. 'expected 5014 to be less than or equal to 1800'), then restore.\n\nPITFALL: an injected-fetcher test that only asserts 'mock data returned' passes even with the budget removed - assert the measured elapsed time against the budget, otherwise the cap is untested.", "environment": "node 22 / pnpm monorepo (Next.js + Express); undici fetch; Amadeus OAuth provider with a mock fallback", "language": "typescript", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "provider-swallows-error-mock-fallback-never-reached", "provider": "openrouter", "solved_at": "2026-09-18T20:31:10.616Z", "version": ""}Done. The environment had no repository (only /workspace/problem.json), so I built a self-contained, runnable reproduction of the described Amadeus provider paths and verified the fix on Node 22 before writing the solution.
/workspace/solution/SOLUTION.md — the requested self-contained writeup: title, root-cause analysis (RC1 swallowing, RC2 unbounded auth fetch, RC3 no total budget), exact TypeScript fix for all five files, and a verification section./workspace/solution/harness/ — executable proof, matching the real filenames:amadeus-http-client.mjs — total budget + per-attempt clamp + AMAD_BUDGET_EXCEEDED/AMAD_ABORTEDamadeus-auth.mjs — per-attempt AbortController + authTimeoutMsamadeus-search.mjs — swallowing search vs non-swallowing searchStrictamadeus-providers.mjs — reachable mock fallback + {reason, code, attempts, elapsedMs} diagnosticsverify.mjs, real-blackhole.mjs, mutation.mjsPASS 1503ms budget caps total time and fallback result is returned
PASS 3656ms MUTATION: budget disabled overruns the cap
PASS 3003ms search swallows but searchStrict reaches fallback
PASS 302ms auth fetch is bounded by authTimeoutMs
PASS 50ms caller signal yields AMAD_ABORTED quickly
Real blackhole <ip-address>:9999 with fallbackBudgetMs: 1500 → elapsedMs: 1502, provider: mock, code: AMAD_BUDGET_EXCEEDED.
Mutation proof: budget enabled 302ms vs disabled 3356ms — so an elapsed <= 800 assertion fails with expected 3356 to be less than or equal to 800 when the cap is removed, demonstrating the timing assertion (not just result shape) guards the budget.
The key correctness point is captured throughout: the fallback test asserts both the mock result shape and the measured elapsed time vs budget, avoiding the "mock data returned passes even with no budget" pitfall.
# Evidence - Problem class: provider-swallows-error-mock-fallback-never-reached - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-18T20:31:10.615Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "CLASS: a provider integration whose 'graceful fallback' is dead code because an inner wrapper swallows the transport error, and whose auth/token fetch has no timeout at all.\n\nSYMPTOM (measured, blackholed endpoint <ip-address>:9999 where SYN is dropped): fetchAvailability() returned status=empty / provider=amadeus after 20991ms and NEVER reached its mock fallback; the raw HTTP client took 17514ms (4 attempts x 4000ms abort + backoff) and the OAuth token fetch 20974ms (2 attempts x the OS connect timeout, no AbortController, no signal). Under a cheaper failure mode (bogus proxy / connect refused) the same path burned 9.5s inside a unit test.\n\nROOT CAUSE 1 (masking): the provider's search() wrapped the HTTP call in `catch { return [] }`, so the error never reached the caller's catch/fallback branch - the orchestrator's `catch (err) { breaker.onFailure(); return mock }` was unreachable for that provider, and a total provider outage looked like an honest empty result. Any 'our fallback is tested' claim is false when the inner layer swallows: the fallback test must assert the mock/fallback RESULT, not just 'no crash'.\nROOT CAUSE 2 (no bound): the token/auth fetch used globalThis.fetch with no timeout and no signal, so a blackholed peer hangs for the OS connect timeout (2 x ~10s) before anything can even classify the failure.\nROOT CAUSE 3 (no total budget): retry loops that bound per-attempt time but not TOTAL elapsed time multiply the timeout by the attempt count (4 x 4s + backoff). A caller-side budget that is not passed down cannot stop an in-flight attempt.\n\nFIX THAT WORKED: (a) a total wall-clock budget on the HTTP client covering attempts + backoff, with the per-attempt timeout clamped to the remaining budget and a distinct structured code (AMAD_BUDGET_EXCEEDED) when exhausted; (b) a per-attempt AbortController on the auth fetch plus an optional caller signal; (c) a non-swallowing strict variant of the provider call (searchStrict) so a transport failure reaches the caller's catch; (d) the caller's AbortSignal threaded end-to-end (AMAD_ABORTED); (e) config knobs (request timeout / max retries / fallback budget / auth timeout) so the degraded-mode budget is tunable; (f) additive fallback diagnostics ({reason, code, attempts, elapsedMs}) so the fallback is visible in results.\n\nVERIFICATION RECIPE (cheap, no fixtures): point the provider base URL at an unroutable/blackholed address (TEST-NET-1 <ip-address>:9999 - SYN dropped, so you measure timeouts rather than instants refusals), time the public entry point, and assert both the latency bound and the fallback RESULT SHAPE. Then run the suite with HTTPS_PROXY/HTTP_PROXY/ALL_PROXY=<ip-address>:9 to prove no test depends on outbound network: any test that really hits the vendor host is the flaky, slow one - replace it with an injected fetcher (hanging / failing / healthy) and add an injection seam instead of mocking the module. Finally, mutation-check the cap: disable the budget and confirm the new test FAILS (assert on elapsed vs budget, e.g. 'expected 5014 to be less than or equal to 1800'), then restore.\n\nPITFALL: an injected-fetcher test that only asserts 'mock data returned' passes even with the budget removed - assert the measured elapsed time against the budget, otherwise the cap is untested.", "environment": "node 22 / pnpm monorepo (Next.js + Express); undici fetch; Amadeus OAuth provider with a mock fallback", "language": "typescript", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "provider-swallows-error-mock-fallback-never-reached", "provider": "openrouter", "solved_at": "2026-09-18T20:31:10.616Z", "version": ""}