Status: Open tracking row (do not file per-occurrence fix rows)
Diagnosed and solved. Since no target repo was mounted (only the unrelated pi tree), I produced the deliverable as a self-contained CI runbook and validated its artifacts: all Python snippets compile, all YAML parses, and the pytest marker config was proven to deselect live tests by default.
Written to: ~/solution/INT-CI-006-live-provider-flake.md
Verification performed:
- pytest config addopts = '-m "not live"' → confirmed 1 passed, 1 deselected, and -m live re-selects it.
- All 3 Python code blocks pass py_compile.
- All 3 YAML workflow blocks parse with yaml.safe_load.
Status: Open tracking row (do not file per-occurrence fix rows)
Owner: CI / integrations
Closure criteria: three consecutive push runs green or the leg is made deterministic (recorded fixtures / opt-in marker / network-free double)
Related: provider-side httpx.ReadTimeout in the push-only live-provider integration job
The default branch shows consecutive red runs, but the red is confined to a single job that calls a real third-party LLM endpoint. Lint, the full unit-test matrix, and packaging-smoke are all green in the same run. The scheduled run of the same head SHA is green because scheduled runs skip the integration job entirely.
This is provider flakiness (httpx.ReadTimeout, rate limits, 5xx), not a regression.
It must be tracked once as INT-CI-006, and the durable fix is to stop the push
event from depending on a live network call.
gh run list reports a run conclusion, which is the AND of all jobs. One flaky job
turns the whole run red, which reads as "the repo is broken" to every contributor.
The failing job is the push-only live-provider integration leg. Its test performs an
outbound httpx call to the upstream LLM API. Under load or provider degradation the
call raises httpx.ReadTimeout and the test does not catch it, so the job exits
non-zero.
push and green on scheduleThe workflow gates the live-integration job to push (and/or provides it the live
secret only on that event):
integration-live:
if: github.event_name == 'push'
schedule runs therefore never execute the job, so the same commit passes on the
scheduled trigger. This event-flavored split is the signature of a live-dependency
flake, not a deterministic code failure (a real regression would also fail the
hermetic lint/test matrix in the same run, and would fail on every event that runs the
job).
A retry loop reduces the probability of a red run but leaves the push gate coupled to
a third-party availability SLO. The run stays nondeterministic, and every incident still
produces a duplicate board row. The correct fix is to make the push path
deterministic and move the live smoke to a non-blocking trigger.
A real regression would show at least one of:
None of these hold:
* job-level check shows the integration job red and everything else green,
* schedule of the same SHA is green,
* the traceback bottoms out in httpx during a provider call.
Never judge CI by run conclusion. Use job-level + same-SHA/different-event grouping.
REPO=<org>/<repo>
# 1. Job-level verdict for the suspect run.
gh run view "<push-run-id>" --repo "$REPO" \
--json jobs --jq '.jobs[] | "\(.name): \(.conclusion)"'
# 2. Group recent runs by head SHA and event.
gh run list --repo "$REPO" --limit 30 \
--json databaseId,event,conclusion,headSha \
--jq 'group_by(.headSha)
| map({
sha: .[0].headSha[0:12],
events: (map({event, conclusion}) | unique)
})
| .[] | "\(.sha) \(.events)"'
Expected output for this class:
push: failure # repo/name = ci, integration-live = failure
schedule: success # no integration-live job
a1b2c3d4e5f6 [{"event":"push","conclusion":"failure"},{"event":"schedule","conclusion":"success"}]
# 3. Find the provider-side error in the failing job log.
gh run view "<push-run-id>" --repo "$REPO" --log-failed \
| grep -Ei 'ReadTimeout|ConnectTimeout|RateLimit|429|50[0-9]|httpx' \
| head -40
If steps 1–3 show provider errors only in the push live leg while lint / matrix / packaging-smoke are green in the same run → this row. Do not open a fix row per occurrence.
Apply A + B + C. A makes the push path deterministic; B keeps a real live smoke available on demand; C removes the push dependency so a provider outage can never turn the default branch red.
Use respx to intercept the httpx calls the provider client makes, so the integration
test exercises the real request/response handling without the network.
pyproject.toml:
[project.optional-dependencies]
test = [
"pytest",
"pytest-asyncio",
"respx>=0.21",
"tenacity>=8",
]
[tool.pytest.ini_options]
addopts = '-m "not live"'
markers = [
"live: hits a real third-party endpoint; deselected by default",
]
tests/integration/conftest.py:
import httpx
import pytest
import respx
LLM_BASE_URL = "https://api.llm-provider.example"
# Deterministic canned completion; keeps the test hermetic.
CANNED_COMPLETION = {
"id": "chatcmpl-ci",
"object": "chat.completion",
"model": "provider-model-1",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "pong"},
"finish_reason": "stop",
}
],
"usage": {"prompt_tokens": 3, "completion_tokens": 1, "total_tokens": 4},
}
@pytest.fixture
def fake_llm():
"""Intercept every provider call for the duration of the test."""
with respx.mock(base_url=LLM_BASE_URL, assert_all_called=False) as router:
router.post("/v1/chat/completions").mock(
return_value=httpx.Response(200, json=CANNED_COMPLETION)
)
yield router
tests/integration/test_llm_roundtrip.py:
import pytest
@pytest.mark.integration
def test_llm_roundtrip(fake_llm):
from mypkg.llm import complete # first-party client, uses httpx under the hood
out = complete("ping")
assert out == "pong"
# Prove the network double was actually used.
assert fake_llm["POST"].called
@pytest.mark.live
def test_llm_roundtrip_live():
"""Opt-in only. Requires MYPPKG_LIVE=1 and real credentials."""
from mypkg.llm import complete
assert complete("ping")
Run locally:
pytest tests/integration -m "not live" -q # hermetic, no network
Record cassettes so even the "live" suite has a deterministic replay mode:
pip install pytest-recording
pytest tests/integration/test_llm_roundtrip.py::test_llm_roundtrip_live \
--record-mode=once # writes tests/integration/cassettes/*.yaml
After recording, default runs replay the cassette and never touch the network. Real network runs require an explicit opt-in:
MYPPKG_LIVE=1 pytest -m live --record-mode=none
Guard the live test with pytest.mark.skipif(not os.getenv("MYPPKG_LIVE"), ...) if
pytest-recording is not adopted.
Optional hardening for the genuinely live run — bounded retries for transient errors only (never for 4xx):
from tenacity import (
retry, retry_if_exception_type, stop_after_attempt, wait_exponential,
)
import httpx
@retry(
retry=retry_if_exception_type((httpx.ReadTimeout, httpx.ConnectTimeout, httpx.HTTPStatusError)),
wait=wait_exponential(multiplier=1, min=1, max=15),
stop=stop_after_attempt(4),
reraise=True,
)
def _post_with_retry(client, *args, **kwargs):
resp = client.post(*args, **kwargs)
if resp.status_code >= 500 or resp.status_code == 429:
raise httpx.HTTPStatusError("transient", request=resp.request, response=resp)
return resp
push path.github/workflows/ci.yml:
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install -e '.[test]'
- run: ruff check .
- run: mypy mypkg
test:
strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.11", "3.12"]
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "${{ matrix.python-version }}" }
- run: pip install -e '.[test]'
# Hermetic: live tests are deselected by addopts + -m below.
- run: pytest -m "not live" -q
packaging-smoke:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install build && python -m build
# Live provider smoke runs on schedule/manual only. It is NOT part of the
# push gate, so provider availability cannot redden the default branch.
integration-live:
if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
continue-on-error: true # belt and braces; do not block anything
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install -e '.[test]'
- run: pytest tests/integration -m live -q --record-mode=none
env:
MYPPKG_LIVE: "1"
LLM_API_KEY: ${{ secrets.LLM_API_KEY }}
Also flip the on: block if the integration job was previously reached from push
only:
on:
push:
pull_request:
schedule:
- cron: "17 6 * * *" # daily live smoke
workflow_dispatch:
If the org requires the live leg on
pushfor compliance, make the fixture-based test the gate and run the live check as a non-required job (continue-on-error: true) plus a required-status list that excludes it.
Create/keep exactly one standing row:
| Field | Value |
|---|---|
| ID | INT-CI-006 |
| Type | CI flake / third-party dependency |
| Trigger | push-only live-provider integration leg |
| Symptom | httpx.ReadTimeout from upstream LLM API |
| Evidence | job-level red on live leg; lint/matrix/packaging green; same SHA green on schedule |
| Fix | hermetic respx double + live opt-in marker + live leg moved off push |
| Close when | 3 consecutive push runs green or leg made deterministic |
| Policy | do not file a duplicate fix row per occurrence; append a dated comment instead |
Append each occurrence as a one-line comment (keeps the run history without row churn):
2026-09-17 push run 123456789 sha a1b2c3d4 httpx.ReadTimeout (job: integration-live)
cd <repo>
pip install -e '.[test]'
# Must pass with the network severed.
pytest -m "not live" -q
# Prove no outbound socket is attempted by the integration test.
# On Linux:
unshare -rn pytest tests/integration -m "not live" -q
# Prove the live test is deselected by default.
pytest tests/integration -q --collect-only | grep -c "test_llm_roundtrip_live" || true
# expected: 0
REPO=<org>/<repo>
# Push run: integration-live must be absent or non-blocking; everything else green.
gh run view "<push-run-id>" --repo "$REPO" \
--json jobs --jq '.jobs[] | "\(.name): \(.conclusion)"'
# Same SHA on schedule: live leg runs (or is skipped intentionally) and is green.
gh run list --repo "$REPO" --limit 30 \
--json databaseId,event,conclusion,headSha \
--jq 'group_by(.headSha)
| map(select(.[0].headSha | startswith("a1b2c3d4")))
| .[] | "\(.[0].headSha[0:12]): \(map({event, conclusion}))"'
Expected:
lint: success
test (3.10): success
test (3.11): success
test (3.12): success
packaging-smoke: success
integration-live: skipped
# Count consecutive green push runs (stop at first failure).
gh run list --repo "$REPO" --branch main --event push --limit 10 \
--json databaseId,conclusion \
--jq '.[] | "\(.databaseId) \(.conclusion)"'
Close INT-CI-006 when the top three consecutive push entries are all
success, or immediately once the hermetic fixture test is the required gate and the
live leg is off the push path (§4). Record the closing comment and date.
The change is additive and reversible:
git revert <commit> # restores the previous CI gating
The respx fixture only affects tests marked/using fake_llm; it does not alter
production code. If the provider API changes, refresh the canned response shape in
tests/integration/conftest.py — no network access required.
A red run on push whose only red job is a live-provider integration leg, while
the same SHA is green on schedule, is a third-party flake: track it as INT-CI-006,
make the push path hermetic with a network double, and move the real live call to an
opt-in, non-blocking trigger.
# Evidence - Problem class: ci-push-only-live-provider-flake - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-17T11:48:09.637Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: `gh run list` shows consecutive red CI runs on the default branch even though every hermetic job is green and the full suite passes locally, so it looks like a code break.\n\nDISCRIMINATION (do this before filing anything):\n1. Never judge CI by RUN conclusion \u2014 judge it by JOB. `gh run view <id> --json jobs --jq '.jobs[] | \"\\(.name): \\(.conclusion)\"'`.\n2. Compare the failing run against a run at the SAME head sha with a DIFFERENT event: `gh run list --json databaseId,event,conclusion,headSha` then group by headSha. If the same sha is red on `push` and green on `schedule`, the failure is event-flavored, not a code break.\n3. Read the failing job's log for a provider-side error (httpx.ReadTimeout / rate limit / 5xx from the upstream LLM API) in a live-provider integration test.\n4. If the red is confined to the push-only live-provider integration leg while lint / unit tests (all matrix versions) / packaging-smoke are green in the SAME run, treat it as flake and TRACK it with a standing row (e.g. INT-CI-<n>) instead of filing a duplicate fix row per occurrence. Close the tracking row only when three consecutive push runs pass, or the leg is made deterministic (recorded fixtures / opt-in marker / network-free double).\n\nWHY IT MATTERS: a run-level red on the default branch reads as 'the repo is broken' to every contributor, and each occurrence otherwise generates a duplicate board row that nobody can fix because the dependency is a third-party endpoint. The discriminator (job-level + same-sha/different-event) is what separates provider flakiness from a real regression.", "environment": "github-actions", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-push-only-live-provider-flake", "provider": "openrouter", "solved_at": "2026-09-17T11:48:09.637Z", "version": "2026-09-17"}Diagnosed and solved. Since no target repo was mounted (only the unrelated pi tree), I produced the deliverable as a self-contained CI runbook and validated its artifacts: all Python snippets compile, all YAML parses, and the pytest marker config was proven to deselect live tests by default.
Written to: ~/solution/INT-CI-006-live-provider-flake.md
Verification performed:
- pytest config addopts = '-m "not live"' → confirmed 1 passed, 1 deselected, and -m live re-selects it.
- All 3 Python code blocks pass py_compile.
- All 3 YAML workflow blocks parse with yaml.safe_load.
Status: Open tracking row (do not file per-occurrence fix rows)
Owner: CI / integrations
Closure criteria: three consecutive push runs green or the leg is made deterministic (recorded fixtures / opt-in marker / network-free double)
Related: provider-side httpx.ReadTimeout in the push-only live-provider integration job
The default branch shows consecutive red runs, but the red is confined to a single job that calls a real third-party LLM endpoint. Lint, the full unit-test matrix, and packaging-smoke are all green in the same run. The scheduled run of the same head SHA is green because scheduled runs skip the integration job entirely.
This is provider flakiness (httpx.ReadTimeout, rate limits, 5xx), not a regression.
It must be tracked once as INT-CI-006, and the durable fix is to stop the push
event from depending on a live network call.
gh run list reports a run conclusion, which is the AND of all jobs. One flaky job
turns the whole run red, which reads as "the repo is broken" to every contributor.
The failing job is the push-only live-provider integration leg. Its test performs an
outbound httpx call to the upstream LLM API. Under load or provider degradation the
call raises httpx.ReadTimeout and the test does not catch it, so the job exits
non-zero.
push and green on scheduleThe workflow gates the live-integration job to push (and/or provides it the live
secret only on that event):
integration-live:
if: github.event_name == 'push'
schedule runs therefore never execute the job, so the same commit passes on the
scheduled trigger. This event-flavored split is the signature of a live-dependency
flake, not a deterministic code failure (a real regression would also fail the
hermetic lint/test matrix in the same run, and would fail on every event that runs the
job).
A retry loop reduces the probability of a red run but leaves the push gate coupled to
a third-party availability SLO. The run stays nondeterministic, and every incident still
produces a duplicate board row. The correct fix is to make the push path
deterministic and move the live smoke to a non-blocking trigger.
A real regression would show at least one of:
None of these hold:
* job-level check shows the integration job red and everything else green,
* schedule of the same SHA is green,
* the traceback bottoms out in httpx during a provider call.
Never judge CI by run conclusion. Use job-level + same-SHA/different-event grouping.
REPO=<org>/<repo>
# 1. Job-level verdict for the suspect run.
gh run view "<push-run-id>" --repo "$REPO" \
--json jobs --jq '.jobs[] | "\(.name): \(.conclusion)"'
# 2. Group recent runs by head SHA and event.
gh run list --repo "$REPO" --limit 30 \
--json databaseId,event,conclusion,headSha \
--jq 'group_by(.headSha)
| map({
sha: .[0].headSha[0:12],
events: (map({event, conclusion}) | unique)
})
| .[] | "\(.sha) \(.events)"'
Expected output for this class:
push: failure # repo/name = ci, integration-live = failure
schedule: success # no integration-live job
a1b2c3d4e5f6 [{"event":"push","conclusion":"failure"},{"event":"schedule","conclusion":"success"}]
# 3. Find the provider-side error in the failing job log.
gh run view "<push-run-id>" --repo "$REPO" --log-failed \
| grep -Ei 'ReadTimeout|ConnectTimeout|RateLimit|429|50[0-9]|httpx' \
| head -40
If steps 1–3 show provider errors only in the push live leg while lint / matrix / packaging-smoke are green in the same run → this row. Do not open a fix row per occurrence.
Apply A + B + C. A makes the push path deterministic; B keeps a real live smoke available on demand; C removes the push dependency so a provider outage can never turn the default branch red.
Use respx to intercept the httpx calls the provider client makes, so the integration
test exercises the real request/response handling without the network.
pyproject.toml:
[project.optional-dependencies]
test = [
"pytest",
"pytest-asyncio",
"respx>=0.21",
"tenacity>=8",
]
[tool.pytest.ini_options]
addopts = '-m "not live"'
markers = [
"live: hits a real third-party endpoint; deselected by default",
]
tests/integration/conftest.py:
import httpx
import pytest
import respx
LLM_BASE_URL = "https://api.llm-provider.example"
# Deterministic canned completion; keeps the test hermetic.
CANNED_COMPLETION = {
"id": "chatcmpl-ci",
"object": "chat.completion",
"model": "provider-model-1",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "pong"},
"finish_reason": "stop",
}
],
"usage": {"prompt_tokens": 3, "completion_tokens": 1, "total_tokens": 4},
}
@pytest.fixture
def fake_llm():
"""Intercept every provider call for the duration of the test."""
with respx.mock(base_url=LLM_BASE_URL, assert_all_called=False) as router:
router.post("/v1/chat/completions").mock(
return_value=httpx.Response(200, json=CANNED_COMPLETION)
)
yield router
tests/integration/test_llm_roundtrip.py:
import pytest
@pytest.mark.integration
def test_llm_roundtrip(fake_llm):
from mypkg.llm import complete # first-party client, uses httpx under the hood
out = complete("ping")
assert out == "pong"
# Prove the network double was actually used.
assert fake_llm["POST"].called
@pytest.mark.live
def test_llm_roundtrip_live():
"""Opt-in only. Requires MYPPKG_LIVE=1 and real credentials."""
from mypkg.llm import complete
assert complete("ping")
Run locally:
pytest tests/integration -m "not live" -q # hermetic, no network
Record cassettes so even the "live" suite has a deterministic replay mode:
pip install pytest-recording
pytest tests/integration/test_llm_roundtrip.py::test_llm_roundtrip_live \
--record-mode=once # writes tests/integration/cassettes/*.yaml
After recording, default runs replay the cassette and never touch the network. Real network runs require an explicit opt-in:
MYPPKG_LIVE=1 pytest -m live --record-mode=none
Guard the live test with pytest.mark.skipif(not os.getenv("MYPPKG_LIVE"), ...) if
pytest-recording is not adopted.
Optional hardening for the genuinely live run — bounded retries for transient errors only (never for 4xx):
from tenacity import (
retry, retry_if_exception_type, stop_after_attempt, wait_exponential,
)
import httpx
@retry(
retry=retry_if_exception_type((httpx.ReadTimeout, httpx.ConnectTimeout, httpx.HTTPStatusError)),
wait=wait_exponential(multiplier=1, min=1, max=15),
stop=stop_after_attempt(4),
reraise=True,
)
def _post_with_retry(client, *args, **kwargs):
resp = client.post(*args, **kwargs)
if resp.status_code >= 500 or resp.status_code == 429:
raise httpx.HTTPStatusError("transient", request=resp.request, response=resp)
return resp
push path.github/workflows/ci.yml:
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install -e '.[test]'
- run: ruff check .
- run: mypy mypkg
test:
strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.11", "3.12"]
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "${{ matrix.python-version }}" }
- run: pip install -e '.[test]'
# Hermetic: live tests are deselected by addopts + -m below.
- run: pytest -m "not live" -q
packaging-smoke:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install build && python -m build
# Live provider smoke runs on schedule/manual only. It is NOT part of the
# push gate, so provider availability cannot redden the default branch.
integration-live:
if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
continue-on-error: true # belt and braces; do not block anything
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install -e '.[test]'
- run: pytest tests/integration -m live -q --record-mode=none
env:
MYPPKG_LIVE: "1"
LLM_API_KEY: ${{ secrets.LLM_API_KEY }}
Also flip the on: block if the integration job was previously reached from push
only:
on:
push:
pull_request:
schedule:
- cron: "17 6 * * *" # daily live smoke
workflow_dispatch:
If the org requires the live leg on
pushfor compliance, make the fixture-based test the gate and run the live check as a non-required job (continue-on-error: true) plus a required-status list that excludes it.
Create/keep exactly one standing row:
| Field | Value |
|---|---|
| ID | INT-CI-006 |
| Type | CI flake / third-party dependency |
| Trigger | push-only live-provider integration leg |
| Symptom | httpx.ReadTimeout from upstream LLM API |
| Evidence | job-level red on live leg; lint/matrix/packaging green; same SHA green on schedule |
| Fix | hermetic respx double + live opt-in marker + live leg moved off push |
| Close when | 3 consecutive push runs green or leg made deterministic |
| Policy | do not file a duplicate fix row per occurrence; append a dated comment instead |
Append each occurrence as a one-line comment (keeps the run history without row churn):
2026-09-17 push run 123456789 sha a1b2c3d4 httpx.ReadTimeout (job: integration-live)
cd <repo>
pip install -e '.[test]'
# Must pass with the network severed.
pytest -m "not live" -q
# Prove no outbound socket is attempted by the integration test.
# On Linux:
unshare -rn pytest tests/integration -m "not live" -q
# Prove the live test is deselected by default.
pytest tests/integration -q --collect-only | grep -c "test_llm_roundtrip_live" || true
# expected: 0
REPO=<org>/<repo>
# Push run: integration-live must be absent or non-blocking; everything else green.
gh run view "<push-run-id>" --repo "$REPO" \
--json jobs --jq '.jobs[] | "\(.name): \(.conclusion)"'
# Same SHA on schedule: live leg runs (or is skipped intentionally) and is green.
gh run list --repo "$REPO" --limit 30 \
--json databaseId,event,conclusion,headSha \
--jq 'group_by(.headSha)
| map(select(.[0].headSha | startswith("a1b2c3d4")))
| .[] | "\(.[0].headSha[0:12]): \(map({event, conclusion}))"'
Expected:
lint: success
test (3.10): success
test (3.11): success
test (3.12): success
packaging-smoke: success
integration-live: skipped
# Count consecutive green push runs (stop at first failure).
gh run list --repo "$REPO" --branch main --event push --limit 10 \
--json databaseId,conclusion \
--jq '.[] | "\(.databaseId) \(.conclusion)"'
Close INT-CI-006 when the top three consecutive push entries are all
success, or immediately once the hermetic fixture test is the required gate and the
live leg is off the push path (§4). Record the closing comment and date.
The change is additive and reversible:
git revert <commit> # restores the previous CI gating
The respx fixture only affects tests marked/using fake_llm; it does not alter
production code. If the provider API changes, refresh the canned response shape in
tests/integration/conftest.py — no network access required.
A red run on push whose only red job is a live-provider integration leg, while
the same SHA is green on schedule, is a third-party flake: track it as INT-CI-006,
make the push path hermetic with a network double, and move the real live call to an
opt-in, non-blocking trigger.
# Evidence - Problem class: ci-push-only-live-provider-flake - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-17T11:48:09.637Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: `gh run list` shows consecutive red CI runs on the default branch even though every hermetic job is green and the full suite passes locally, so it looks like a code break.\n\nDISCRIMINATION (do this before filing anything):\n1. Never judge CI by RUN conclusion \u2014 judge it by JOB. `gh run view <id> --json jobs --jq '.jobs[] | \"\\(.name): \\(.conclusion)\"'`.\n2. Compare the failing run against a run at the SAME head sha with a DIFFERENT event: `gh run list --json databaseId,event,conclusion,headSha` then group by headSha. If the same sha is red on `push` and green on `schedule`, the failure is event-flavored, not a code break.\n3. Read the failing job's log for a provider-side error (httpx.ReadTimeout / rate limit / 5xx from the upstream LLM API) in a live-provider integration test.\n4. If the red is confined to the push-only live-provider integration leg while lint / unit tests (all matrix versions) / packaging-smoke are green in the SAME run, treat it as flake and TRACK it with a standing row (e.g. INT-CI-<n>) instead of filing a duplicate fix row per occurrence. Close the tracking row only when three consecutive push runs pass, or the leg is made deterministic (recorded fixtures / opt-in marker / network-free double).\n\nWHY IT MATTERS: a run-level red on the default branch reads as 'the repo is broken' to every contributor, and each occurrence otherwise generates a duplicate board row that nobody can fix because the dependency is a third-party endpoint. The discriminator (job-level + same-sha/different-event) is what separates provider flakiness from a real regression.", "environment": "github-actions", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-push-only-live-provider-flake", "provider": "openrouter", "solved_at": "2026-09-17T11:48:09.637Z", "version": "2026-09-17"}