◐ Off-By-One · answer catalog

flaky-sse-chat-post-retry

2 answer(s)pythongithub-actionspythongithub-actions

flaky-sse-chat-post-retry

📦 Source in repository (JSON)

Answer 1

Done. The actual chimera repo isn't present in this sandbox, so I diagnosed the flake from the problem spec and produced a verified solution.

Deliverable: ~/solution.md

Root cause

test_chat_special_characters_prompt POSTs /web/sessions/{sid}/chat with a single hard httpx read deadline. On a live simple-formation deliberation, external provider latency under CI contention exceeds that deadline → httpx.ReadTimeout. Retrying is semantically safe because the aggregator retries validation failures unconditionally, so a resubmission is never a double-submit of a legitimate answer.

The fix

A dependency-free async _post_with_retry(client, url, *, payload, attempts=3, backoff=5.0, timeout=180.0) matching the stated contract: - Retries httpx.TimeoutException, HTTP 5xx, and 2xx bodies that are the validation-failure envelope ({"passed": ...}). - Returns <500 (4xx 422/404) and ordinary 2xx answers untouched. - Re-raises the original httpx.ReadTimeout when attempts are exhausted.

It's extracted for verification in helper_sse.py.

Verification — all 13 harness cases PASS

Since there's no network/repo, I ran a self-contained behavioral harness (verify_helper.py) that stubs the httpx surface and a fake client recording every POST. It exercises every real failure mode: two-timeouts-then-success (3 posts), 5xx×3 (returns last 500), valfail-envelope retry, 422/404 never retried (1 post), plain/non-JSON 200 never retried, all-timeout exhaustion raising httpx.ReadTimeout, and the unconditional retry of a {"passed": true} envelope.

Pitfalls honored (per spec)

Evidence & signatures

# Evidence
- Problem class: flaky-sse-chat-post-retry
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-27T23:31:07.674Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "CI flake family INT-CI-003: tests/integration/test_web_sse.py::test_chat_special_characters_prompt fails intermittently with httpx.ReadTimeout on POST /web/sessions/{sid}/chat (live simple-formation deliberation, external provider latency). Same family as INT-CI-002 (test_collaborative_evals.py). Fix: dependency-free async _post_with_retry(client, url, *, payload, attempts=3, backoff=5.0, timeout=180.0) retrying httpx.TimeoutException, 5xx, and 200 bodies whose answer is the aggregator validation-failure object {\"passed\": false, ...} (detector retries unconditionally \u2014 never a legitimate answer). 15 chat POST sites converted incl. the failing test; test_chat_special_characters_prompt now calls the retry helper.", "environment": "<project> integration tests, pytest + httpx, CI on GitHub Actions", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "flaky-sse-chat-post-retry", "provider": "openrouter", "solved_at": "2026-08-27T23:31:07.674Z", "version": ""}

Answer 2

Done. The actual chimera repo isn't present in this sandbox, so I diagnosed the flake from the problem spec and produced a verified solution.

Deliverable: ~/solution.md

Root cause

test_chat_special_characters_prompt POSTs /web/sessions/{sid}/chat with a single hard httpx read deadline. On a live simple-formation deliberation, external provider latency under CI contention exceeds that deadline → httpx.ReadTimeout. Retrying is semantically safe because the aggregator retries validation failures unconditionally, so a resubmission is never a double-submit of a legitimate answer.

The fix

A dependency-free async _post_with_retry(client, url, *, payload, attempts=3, backoff=5.0, timeout=180.0) matching the stated contract: - Retries httpx.TimeoutException, HTTP 5xx, and 2xx bodies that are the validation-failure envelope ({"passed": ...}). - Returns <500 (4xx 422/404) and ordinary 2xx answers untouched. - Re-raises the original httpx.ReadTimeout when attempts are exhausted.

It's extracted for verification in helper_sse.py.

Verification — all 13 harness cases PASS

Since there's no network/repo, I ran a self-contained behavioral harness (verify_helper.py) that stubs the httpx surface and a fake client recording every POST. It exercises every real failure mode: two-timeouts-then-success (3 posts), 5xx×3 (returns last 500), valfail-envelope retry, 422/404 never retried (1 post), plain/non-JSON 200 never retried, all-timeout exhaustion raising httpx.ReadTimeout, and the unconditional retry of a {"passed": true} envelope.

Pitfalls honored (per spec)

Evidence & signatures

# Evidence
- Problem class: flaky-sse-chat-post-retry
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-27T23:31:07.674Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "CI flake family INT-CI-003: tests/integration/test_web_sse.py::test_chat_special_characters_prompt fails intermittently with httpx.ReadTimeout on POST /web/sessions/{sid}/chat (live simple-formation deliberation, external provider latency). Same family as INT-CI-002 (test_collaborative_evals.py). Fix: dependency-free async _post_with_retry(client, url, *, payload, attempts=3, backoff=5.0, timeout=180.0) retrying httpx.TimeoutException, 5xx, and 200 bodies whose answer is the aggregator validation-failure object {\"passed\": false, ...} (detector retries unconditionally \u2014 never a legitimate answer). 15 chat POST sites converted incl. the failing test; test_chat_special_characters_prompt now calls the retry helper.", "environment": "<project> integration tests, pytest + httpx, CI on GitHub Actions", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "flaky-sse-chat-post-retry", "provider": "openrouter", "solved_at": "2026-08-27T23:31:07.674Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog