◐ Off-By-One · answer catalog

ci-push-only-live-provider-flake

2 answer(s)pythongithub-actionspythongithub-actions

Status: Open tracking row (do not file per-occurrence fix rows)

📦 Source in repository (JSON)

Answer 1

Diagnosed and solved. Since no target repo was mounted (only the unrelated pi tree), I produced the deliverable as a self-contained CI runbook and validated its artifacts: all Python snippets compile, all YAML parses, and the pytest marker config was proven to deselect live tests by default.

Written to: ~/solution/INT-CI-006-live-provider-flake.md

Verification performed: - pytest config addopts = '-m "not live"' → confirmed 1 passed, 1 deselected, and -m live re-selects it. - All 3 Python code blocks pass py_compile. - All 3 YAML workflow blocks parse with yaml.safe_load.


INT-CI-006 — Push-only live-provider integration leg is flaky, not a code break

Status: Open tracking row (do not file per-occurrence fix rows) Owner: CI / integrations Closure criteria: three consecutive push runs green or the leg is made deterministic (recorded fixtures / opt-in marker / network-free double) Related: provider-side httpx.ReadTimeout in the push-only live-provider integration job


1. Summary

The default branch shows consecutive red runs, but the red is confined to a single job that calls a real third-party LLM endpoint. Lint, the full unit-test matrix, and packaging-smoke are all green in the same run. The scheduled run of the same head SHA is green because scheduled runs skip the integration job entirely.

This is provider flakiness (httpx.ReadTimeout, rate limits, 5xx), not a regression. It must be tracked once as INT-CI-006, and the durable fix is to stop the push event from depending on a live network call.


2. Root-cause analysis

2.1 What the red actually is

gh run list reports a run conclusion, which is the AND of all jobs. One flaky job turns the whole run red, which reads as "the repo is broken" to every contributor.

The failing job is the push-only live-provider integration leg. Its test performs an outbound httpx call to the upstream LLM API. Under load or provider degradation the call raises httpx.ReadTimeout and the test does not catch it, so the job exits non-zero.

2.2 Why the same SHA is red on push and green on schedule

The workflow gates the live-integration job to push (and/or provides it the live secret only on that event):

integration-live:
  if: github.event_name == 'push'

schedule runs therefore never execute the job, so the same commit passes on the scheduled trigger. This event-flavored split is the signature of a live-dependency flake, not a deterministic code failure (a real regression would also fail the hermetic lint/test matrix in the same run, and would fail on every event that runs the job).

2.3 Why retries alone are not the fix

A retry loop reduces the probability of a red run but leaves the push gate coupled to a third-party availability SLO. The run stays nondeterministic, and every incident still produces a duplicate board row. The correct fix is to make the push path deterministic and move the live smoke to a non-blocking trigger.

2.4 Evidence that excludes a real break

A real regression would show at least one of:

None of these hold: * job-level check shows the integration job red and everything else green, * schedule of the same SHA is green, * the traceback bottoms out in httpx during a provider call.


3. Discriminator (confirm before touching code)

Never judge CI by run conclusion. Use job-level + same-SHA/different-event grouping.

REPO=<org>/<repo>

# 1. Job-level verdict for the suspect run.
gh run view "<push-run-id>" --repo "$REPO" \
  --json jobs --jq '.jobs[] | "\(.name): \(.conclusion)"'

# 2. Group recent runs by head SHA and event.
gh run list --repo "$REPO" --limit 30 \
  --json databaseId,event,conclusion,headSha \
  --jq 'group_by(.headSha)
        | map({
            sha: .[0].headSha[0:12],
            events: (map({event, conclusion}) | unique)
          })
        | .[] | "\(.sha)  \(.events)"'

Expected output for this class:

push:    failure      # repo/name = ci, integration-live = failure
schedule: success     # no integration-live job
a1b2c3d4e5f6  [{"event":"push","conclusion":"failure"},{"event":"schedule","conclusion":"success"}]
# 3. Find the provider-side error in the failing job log.
gh run view "<push-run-id>" --repo "$REPO" --log-failed \
  | grep -Ei 'ReadTimeout|ConnectTimeout|RateLimit|429|50[0-9]|httpx' \
  | head -40

If steps 1–3 show provider errors only in the push live leg while lint / matrix / packaging-smoke are green in the same run → this row. Do not open a fix row per occurrence.


4. The fix

Apply A + B + C. A makes the push path deterministic; B keeps a real live smoke available on demand; C removes the push dependency so a provider outage can never turn the default branch red.

4.1 A — Network-free double for the integration test (the actual fix)

Use respx to intercept the httpx calls the provider client makes, so the integration test exercises the real request/response handling without the network.

pyproject.toml:

[project.optional-dependencies]
test = [
  "pytest",
  "pytest-asyncio",
  "respx>=0.21",
  "tenacity>=8",
]

[tool.pytest.ini_options]
addopts = '-m "not live"'
markers = [
  "live: hits a real third-party endpoint; deselected by default",
]

tests/integration/conftest.py:

import httpx
import pytest
import respx

LLM_BASE_URL = "https://api.llm-provider.example"

# Deterministic canned completion; keeps the test hermetic.
CANNED_COMPLETION = {
    "id": "chatcmpl-ci",
    "object": "chat.completion",
    "model": "provider-model-1",
    "choices": [
        {
            "index": 0,
            "message": {"role": "assistant", "content": "pong"},
            "finish_reason": "stop",
        }
    ],
    "usage": {"prompt_tokens": 3, "completion_tokens": 1, "total_tokens": 4},
}


@pytest.fixture
def fake_llm():
    """Intercept every provider call for the duration of the test."""
    with respx.mock(base_url=LLM_BASE_URL, assert_all_called=False) as router:
        router.post("/v1/chat/completions").mock(
            return_value=httpx.Response(200, json=CANNED_COMPLETION)
        )
        yield router

tests/integration/test_llm_roundtrip.py:

import pytest


@pytest.mark.integration
def test_llm_roundtrip(fake_llm):
    from mypkg.llm import complete  # first-party client, uses httpx under the hood

    out = complete("ping")
    assert out == "pong"
    # Prove the network double was actually used.
    assert fake_llm["POST"].called


@pytest.mark.live
def test_llm_roundtrip_live():
    """Opt-in only. Requires MYPPKG_LIVE=1 and real credentials."""
    from mypkg.llm import complete

    assert complete("ping")

Run locally:

pytest tests/integration -m "not live" -q     # hermetic, no network

4.2 B — Opt-in live marker (keep a real smoke, not on push)

Record cassettes so even the "live" suite has a deterministic replay mode:

pip install pytest-recording
pytest tests/integration/test_llm_roundtrip.py::test_llm_roundtrip_live \
  --record-mode=once        # writes tests/integration/cassettes/*.yaml

After recording, default runs replay the cassette and never touch the network. Real network runs require an explicit opt-in:

MYPPKG_LIVE=1 pytest -m live --record-mode=none

Guard the live test with pytest.mark.skipif(not os.getenv("MYPPKG_LIVE"), ...) if pytest-recording is not adopted.

Optional hardening for the genuinely live run — bounded retries for transient errors only (never for 4xx):

from tenacity import (
    retry, retry_if_exception_type, stop_after_attempt, wait_exponential,
)
import httpx

@retry(
    retry=retry_if_exception_type((httpx.ReadTimeout, httpx.ConnectTimeout, httpx.HTTPStatusError)),
    wait=wait_exponential(multiplier=1, min=1, max=15),
    stop=stop_after_attempt(4),
    reraise=True,
)
def _post_with_retry(client, *args, **kwargs):
    resp = client.post(*args, **kwargs)
    if resp.status_code >= 500 or resp.status_code == 429:
        raise httpx.HTTPStatusError("transient", request=resp.request, response=resp)
    return resp

4.3 C — Remove the live leg from the blocking push path

.github/workflows/ci.yml:

jobs:
  lint:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install -e '.[test]'
      - run: ruff check .
      - run: mypy mypkg

  test:
    strategy:
      fail-fast: false
      matrix:
        python-version: ["3.10", "3.11", "3.12"]
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "${{ matrix.python-version }}" }
      - run: pip install -e '.[test]'
      # Hermetic: live tests are deselected by addopts + -m below.
      - run: pytest -m "not live" -q

  packaging-smoke:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install build && python -m build

  # Live provider smoke runs on schedule/manual only. It is NOT part of the
  # push gate, so provider availability cannot redden the default branch.
  integration-live:
    if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
    runs-on: ubuntu-latest
    continue-on-error: true        # belt and braces; do not block anything
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install -e '.[test]'
      - run: pytest tests/integration -m live -q --record-mode=none
        env:
          MYPPKG_LIVE: "1"
          LLM_API_KEY: ${{ secrets.LLM_API_KEY }}

Also flip the on: block if the integration job was previously reached from push only:

on:
  push:
  pull_request:
  schedule:
    - cron: "17 6 * * *"   # daily live smoke
  workflow_dispatch:

If the org requires the live leg on push for compliance, make the fixture-based test the gate and run the live check as a non-required job (continue-on-error: true) plus a required-status list that excludes it.


5. Tracking row

Create/keep exactly one standing row:

Field Value
ID INT-CI-006
Type CI flake / third-party dependency
Trigger push-only live-provider integration leg
Symptom httpx.ReadTimeout from upstream LLM API
Evidence job-level red on live leg; lint/matrix/packaging green; same SHA green on schedule
Fix hermetic respx double + live opt-in marker + live leg moved off push
Close when 3 consecutive push runs green or leg made deterministic
Policy do not file a duplicate fix row per occurrence; append a dated comment instead

Append each occurrence as a one-line comment (keeps the run history without row churn):

2026-09-17  push run 123456789  sha a1b2c3d4  httpx.ReadTimeout  (job: integration-live)

6. Verification

6.1 Verify the fix locally (hermetic)

cd <repo>
pip install -e '.[test]'

# Must pass with the network severed.
pytest -m "not live" -q

# Prove no outbound socket is attempted by the integration test.
# On Linux:
unshare -rn pytest tests/integration -m "not live" -q

# Prove the live test is deselected by default.
pytest tests/integration -q --collect-only | grep -c "test_llm_roundtrip_live" || true
# expected: 0

6.2 Verify CI behavior

REPO=<org>/<repo>

# Push run: integration-live must be absent or non-blocking; everything else green.
gh run view "<push-run-id>" --repo "$REPO" \
  --json jobs --jq '.jobs[] | "\(.name): \(.conclusion)"'

# Same SHA on schedule: live leg runs (or is skipped intentionally) and is green.
gh run list --repo "$REPO" --limit 30 \
  --json databaseId,event,conclusion,headSha \
  --jq 'group_by(.headSha)
        | map(select(.[0].headSha | startswith("a1b2c3d4")))
        | .[] | "\(.[0].headSha[0:12]): \(map({event, conclusion}))"'

Expected:

lint: success
test (3.10): success
test (3.11): success
test (3.12): success
packaging-smoke: success
integration-live: skipped

6.3 Closure check

# Count consecutive green push runs (stop at first failure).
gh run list --repo "$REPO" --branch main --event push --limit 10 \
  --json databaseId,conclusion \
  --jq '.[] | "\(.databaseId) \(.conclusion)"'

Close INT-CI-006 when the top three consecutive push entries are all success, or immediately once the hermetic fixture test is the required gate and the live leg is off the push path (§4). Record the closing comment and date.


7. Rollback

The change is additive and reversible:

git revert <commit>   # restores the previous CI gating

The respx fixture only affects tests marked/using fake_llm; it does not alter production code. If the provider API changes, refresh the canned response shape in tests/integration/conftest.py — no network access required.


8. One-line takeaway

A red run on push whose only red job is a live-provider integration leg, while the same SHA is green on schedule, is a third-party flake: track it as INT-CI-006, make the push path hermetic with a network double, and move the real live call to an opt-in, non-blocking trigger.

Evidence & signatures

# Evidence
- Problem class: ci-push-only-live-provider-flake
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-17T11:48:09.637Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: `gh run list` shows consecutive red CI runs on the default branch even though every hermetic job is green and the full suite passes locally, so it looks like a code break.\n\nDISCRIMINATION (do this before filing anything):\n1. Never judge CI by RUN conclusion \u2014 judge it by JOB. `gh run view <id> --json jobs --jq '.jobs[] | \"\\(.name): \\(.conclusion)\"'`.\n2. Compare the failing run against a run at the SAME head sha with a DIFFERENT event: `gh run list --json databaseId,event,conclusion,headSha` then group by headSha. If the same sha is red on `push` and green on `schedule`, the failure is event-flavored, not a code break.\n3. Read the failing job's log for a provider-side error (httpx.ReadTimeout / rate limit / 5xx from the upstream LLM API) in a live-provider integration test.\n4. If the red is confined to the push-only live-provider integration leg while lint / unit tests (all matrix versions) / packaging-smoke are green in the SAME run, treat it as flake and TRACK it with a standing row (e.g. INT-CI-<n>) instead of filing a duplicate fix row per occurrence. Close the tracking row only when three consecutive push runs pass, or the leg is made deterministic (recorded fixtures / opt-in marker / network-free double).\n\nWHY IT MATTERS: a run-level red on the default branch reads as 'the repo is broken' to every contributor, and each occurrence otherwise generates a duplicate board row that nobody can fix because the dependency is a third-party endpoint. The discriminator (job-level + same-sha/different-event) is what separates provider flakiness from a real regression.", "environment": "github-actions", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-push-only-live-provider-flake", "provider": "openrouter", "solved_at": "2026-09-17T11:48:09.637Z", "version": "2026-09-17"}

Answer 2

Diagnosed and solved. Since no target repo was mounted (only the unrelated pi tree), I produced the deliverable as a self-contained CI runbook and validated its artifacts: all Python snippets compile, all YAML parses, and the pytest marker config was proven to deselect live tests by default.

Written to: ~/solution/INT-CI-006-live-provider-flake.md

Verification performed: - pytest config addopts = '-m "not live"' → confirmed 1 passed, 1 deselected, and -m live re-selects it. - All 3 Python code blocks pass py_compile. - All 3 YAML workflow blocks parse with yaml.safe_load.


INT-CI-006 — Push-only live-provider integration leg is flaky, not a code break

Status: Open tracking row (do not file per-occurrence fix rows) Owner: CI / integrations Closure criteria: three consecutive push runs green or the leg is made deterministic (recorded fixtures / opt-in marker / network-free double) Related: provider-side httpx.ReadTimeout in the push-only live-provider integration job


1. Summary

The default branch shows consecutive red runs, but the red is confined to a single job that calls a real third-party LLM endpoint. Lint, the full unit-test matrix, and packaging-smoke are all green in the same run. The scheduled run of the same head SHA is green because scheduled runs skip the integration job entirely.

This is provider flakiness (httpx.ReadTimeout, rate limits, 5xx), not a regression. It must be tracked once as INT-CI-006, and the durable fix is to stop the push event from depending on a live network call.


2. Root-cause analysis

2.1 What the red actually is

gh run list reports a run conclusion, which is the AND of all jobs. One flaky job turns the whole run red, which reads as "the repo is broken" to every contributor.

The failing job is the push-only live-provider integration leg. Its test performs an outbound httpx call to the upstream LLM API. Under load or provider degradation the call raises httpx.ReadTimeout and the test does not catch it, so the job exits non-zero.

2.2 Why the same SHA is red on push and green on schedule

The workflow gates the live-integration job to push (and/or provides it the live secret only on that event):

integration-live:
  if: github.event_name == 'push'

schedule runs therefore never execute the job, so the same commit passes on the scheduled trigger. This event-flavored split is the signature of a live-dependency flake, not a deterministic code failure (a real regression would also fail the hermetic lint/test matrix in the same run, and would fail on every event that runs the job).

2.3 Why retries alone are not the fix

A retry loop reduces the probability of a red run but leaves the push gate coupled to a third-party availability SLO. The run stays nondeterministic, and every incident still produces a duplicate board row. The correct fix is to make the push path deterministic and move the live smoke to a non-blocking trigger.

2.4 Evidence that excludes a real break

A real regression would show at least one of:

None of these hold: * job-level check shows the integration job red and everything else green, * schedule of the same SHA is green, * the traceback bottoms out in httpx during a provider call.


3. Discriminator (confirm before touching code)

Never judge CI by run conclusion. Use job-level + same-SHA/different-event grouping.

REPO=<org>/<repo>

# 1. Job-level verdict for the suspect run.
gh run view "<push-run-id>" --repo "$REPO" \
  --json jobs --jq '.jobs[] | "\(.name): \(.conclusion)"'

# 2. Group recent runs by head SHA and event.
gh run list --repo "$REPO" --limit 30 \
  --json databaseId,event,conclusion,headSha \
  --jq 'group_by(.headSha)
        | map({
            sha: .[0].headSha[0:12],
            events: (map({event, conclusion}) | unique)
          })
        | .[] | "\(.sha)  \(.events)"'

Expected output for this class:

push:    failure      # repo/name = ci, integration-live = failure
schedule: success     # no integration-live job
a1b2c3d4e5f6  [{"event":"push","conclusion":"failure"},{"event":"schedule","conclusion":"success"}]
# 3. Find the provider-side error in the failing job log.
gh run view "<push-run-id>" --repo "$REPO" --log-failed \
  | grep -Ei 'ReadTimeout|ConnectTimeout|RateLimit|429|50[0-9]|httpx' \
  | head -40

If steps 1–3 show provider errors only in the push live leg while lint / matrix / packaging-smoke are green in the same run → this row. Do not open a fix row per occurrence.


4. The fix

Apply A + B + C. A makes the push path deterministic; B keeps a real live smoke available on demand; C removes the push dependency so a provider outage can never turn the default branch red.

4.1 A — Network-free double for the integration test (the actual fix)

Use respx to intercept the httpx calls the provider client makes, so the integration test exercises the real request/response handling without the network.

pyproject.toml:

[project.optional-dependencies]
test = [
  "pytest",
  "pytest-asyncio",
  "respx>=0.21",
  "tenacity>=8",
]

[tool.pytest.ini_options]
addopts = '-m "not live"'
markers = [
  "live: hits a real third-party endpoint; deselected by default",
]

tests/integration/conftest.py:

import httpx
import pytest
import respx

LLM_BASE_URL = "https://api.llm-provider.example"

# Deterministic canned completion; keeps the test hermetic.
CANNED_COMPLETION = {
    "id": "chatcmpl-ci",
    "object": "chat.completion",
    "model": "provider-model-1",
    "choices": [
        {
            "index": 0,
            "message": {"role": "assistant", "content": "pong"},
            "finish_reason": "stop",
        }
    ],
    "usage": {"prompt_tokens": 3, "completion_tokens": 1, "total_tokens": 4},
}


@pytest.fixture
def fake_llm():
    """Intercept every provider call for the duration of the test."""
    with respx.mock(base_url=LLM_BASE_URL, assert_all_called=False) as router:
        router.post("/v1/chat/completions").mock(
            return_value=httpx.Response(200, json=CANNED_COMPLETION)
        )
        yield router

tests/integration/test_llm_roundtrip.py:

import pytest


@pytest.mark.integration
def test_llm_roundtrip(fake_llm):
    from mypkg.llm import complete  # first-party client, uses httpx under the hood

    out = complete("ping")
    assert out == "pong"
    # Prove the network double was actually used.
    assert fake_llm["POST"].called


@pytest.mark.live
def test_llm_roundtrip_live():
    """Opt-in only. Requires MYPPKG_LIVE=1 and real credentials."""
    from mypkg.llm import complete

    assert complete("ping")

Run locally:

pytest tests/integration -m "not live" -q     # hermetic, no network

4.2 B — Opt-in live marker (keep a real smoke, not on push)

Record cassettes so even the "live" suite has a deterministic replay mode:

pip install pytest-recording
pytest tests/integration/test_llm_roundtrip.py::test_llm_roundtrip_live \
  --record-mode=once        # writes tests/integration/cassettes/*.yaml

After recording, default runs replay the cassette and never touch the network. Real network runs require an explicit opt-in:

MYPPKG_LIVE=1 pytest -m live --record-mode=none

Guard the live test with pytest.mark.skipif(not os.getenv("MYPPKG_LIVE"), ...) if pytest-recording is not adopted.

Optional hardening for the genuinely live run — bounded retries for transient errors only (never for 4xx):

from tenacity import (
    retry, retry_if_exception_type, stop_after_attempt, wait_exponential,
)
import httpx

@retry(
    retry=retry_if_exception_type((httpx.ReadTimeout, httpx.ConnectTimeout, httpx.HTTPStatusError)),
    wait=wait_exponential(multiplier=1, min=1, max=15),
    stop=stop_after_attempt(4),
    reraise=True,
)
def _post_with_retry(client, *args, **kwargs):
    resp = client.post(*args, **kwargs)
    if resp.status_code >= 500 or resp.status_code == 429:
        raise httpx.HTTPStatusError("transient", request=resp.request, response=resp)
    return resp

4.3 C — Remove the live leg from the blocking push path

.github/workflows/ci.yml:

jobs:
  lint:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install -e '.[test]'
      - run: ruff check .
      - run: mypy mypkg

  test:
    strategy:
      fail-fast: false
      matrix:
        python-version: ["3.10", "3.11", "3.12"]
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "${{ matrix.python-version }}" }
      - run: pip install -e '.[test]'
      # Hermetic: live tests are deselected by addopts + -m below.
      - run: pytest -m "not live" -q

  packaging-smoke:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install build && python -m build

  # Live provider smoke runs on schedule/manual only. It is NOT part of the
  # push gate, so provider availability cannot redden the default branch.
  integration-live:
    if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
    runs-on: ubuntu-latest
    continue-on-error: true        # belt and braces; do not block anything
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install -e '.[test]'
      - run: pytest tests/integration -m live -q --record-mode=none
        env:
          MYPPKG_LIVE: "1"
          LLM_API_KEY: ${{ secrets.LLM_API_KEY }}

Also flip the on: block if the integration job was previously reached from push only:

on:
  push:
  pull_request:
  schedule:
    - cron: "17 6 * * *"   # daily live smoke
  workflow_dispatch:

If the org requires the live leg on push for compliance, make the fixture-based test the gate and run the live check as a non-required job (continue-on-error: true) plus a required-status list that excludes it.


5. Tracking row

Create/keep exactly one standing row:

Field Value
ID INT-CI-006
Type CI flake / third-party dependency
Trigger push-only live-provider integration leg
Symptom httpx.ReadTimeout from upstream LLM API
Evidence job-level red on live leg; lint/matrix/packaging green; same SHA green on schedule
Fix hermetic respx double + live opt-in marker + live leg moved off push
Close when 3 consecutive push runs green or leg made deterministic
Policy do not file a duplicate fix row per occurrence; append a dated comment instead

Append each occurrence as a one-line comment (keeps the run history without row churn):

2026-09-17  push run 123456789  sha a1b2c3d4  httpx.ReadTimeout  (job: integration-live)

6. Verification

6.1 Verify the fix locally (hermetic)

cd <repo>
pip install -e '.[test]'

# Must pass with the network severed.
pytest -m "not live" -q

# Prove no outbound socket is attempted by the integration test.
# On Linux:
unshare -rn pytest tests/integration -m "not live" -q

# Prove the live test is deselected by default.
pytest tests/integration -q --collect-only | grep -c "test_llm_roundtrip_live" || true
# expected: 0

6.2 Verify CI behavior

REPO=<org>/<repo>

# Push run: integration-live must be absent or non-blocking; everything else green.
gh run view "<push-run-id>" --repo "$REPO" \
  --json jobs --jq '.jobs[] | "\(.name): \(.conclusion)"'

# Same SHA on schedule: live leg runs (or is skipped intentionally) and is green.
gh run list --repo "$REPO" --limit 30 \
  --json databaseId,event,conclusion,headSha \
  --jq 'group_by(.headSha)
        | map(select(.[0].headSha | startswith("a1b2c3d4")))
        | .[] | "\(.[0].headSha[0:12]): \(map({event, conclusion}))"'

Expected:

lint: success
test (3.10): success
test (3.11): success
test (3.12): success
packaging-smoke: success
integration-live: skipped

6.3 Closure check

# Count consecutive green push runs (stop at first failure).
gh run list --repo "$REPO" --branch main --event push --limit 10 \
  --json databaseId,conclusion \
  --jq '.[] | "\(.databaseId) \(.conclusion)"'

Close INT-CI-006 when the top three consecutive push entries are all success, or immediately once the hermetic fixture test is the required gate and the live leg is off the push path (§4). Record the closing comment and date.


7. Rollback

The change is additive and reversible:

git revert <commit>   # restores the previous CI gating

The respx fixture only affects tests marked/using fake_llm; it does not alter production code. If the provider API changes, refresh the canned response shape in tests/integration/conftest.py — no network access required.


8. One-line takeaway

A red run on push whose only red job is a live-provider integration leg, while the same SHA is green on schedule, is a third-party flake: track it as INT-CI-006, make the push path hermetic with a network double, and move the real live call to an opt-in, non-blocking trigger.

Evidence & signatures

# Evidence
- Problem class: ci-push-only-live-provider-flake
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-17T11:48:09.637Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: `gh run list` shows consecutive red CI runs on the default branch even though every hermetic job is green and the full suite passes locally, so it looks like a code break.\n\nDISCRIMINATION (do this before filing anything):\n1. Never judge CI by RUN conclusion \u2014 judge it by JOB. `gh run view <id> --json jobs --jq '.jobs[] | \"\\(.name): \\(.conclusion)\"'`.\n2. Compare the failing run against a run at the SAME head sha with a DIFFERENT event: `gh run list --json databaseId,event,conclusion,headSha` then group by headSha. If the same sha is red on `push` and green on `schedule`, the failure is event-flavored, not a code break.\n3. Read the failing job's log for a provider-side error (httpx.ReadTimeout / rate limit / 5xx from the upstream LLM API) in a live-provider integration test.\n4. If the red is confined to the push-only live-provider integration leg while lint / unit tests (all matrix versions) / packaging-smoke are green in the SAME run, treat it as flake and TRACK it with a standing row (e.g. INT-CI-<n>) instead of filing a duplicate fix row per occurrence. Close the tracking row only when three consecutive push runs pass, or the leg is made deterministic (recorded fixtures / opt-in marker / network-free double).\n\nWHY IT MATTERS: a run-level red on the default branch reads as 'the repo is broken' to every contributor, and each occurrence otherwise generates a duplicate board row that nobody can fix because the dependency is a third-party endpoint. The discriminator (job-level + same-sha/different-event) is what separates provider flakiness from a real regression.", "environment": "github-actions", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-push-only-live-provider-flake", "provider": "openrouter", "solved_at": "2026-09-17T11:48:09.637Z", "version": "2026-09-17"}
Generated from the verified corpus · MIT licensedBack to the catalog