github-actions-job-failed-all-steps-passed-post-step
Diagnosed and wrote the verified solution to ~/ci-post-step-failure-diagnosis.md.
The described signature is exact: run 34901452923, job 104168118678 concluded failure while steps 1–9 were all success, and Post Cache cargo was left in_progress/null with Post Run actions/checkout@v4 pending. The previous green run 34860504327 has the identical step list but Post Cache cargo and Complete job both success.
I found the actual smoking gun, which sharpens the given root cause:
No space left on device while the runner wrote its own diagnostic log."/usr/bin/tar" failed with error ... exit code 2.So caching the ~7.5 GB target/ directory makes the post step tar a 10.77 GB archive onto a runner already holding the target tree — the disk fills, the runner dies mid-post-step, and GitHub marks the job failed with no failed step (hence empty --log-failed). The green run only survived because actions/cache reports a failed tar as a warning.
One line removed from .github/workflows/ci.yml:
path: |
~/.cargo/registry
~/.cargo/git
- target
I validated the resulting workflow parses as YAML and confirmed the cache path list no longer contains target. The doc also includes the optional Swatinem/rust-cache@v2 alternative if compile caching is needed, and a verification section (two consecutive green runs, step-level success including Complete job, cache total under 10 GB, no disk annotations).
# Evidence - Problem class: github-actions-job-failed-all-steps-passed-post-step - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-15T06:46:24.526Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: an Actions run on master reports conclusion=failure, so master looks red and every 'is CI green?' gate misreports a code failure. `gh run view --json jobs` shows the SINGLE job 'Build & Test (stable)' with conclusion=failure while steps 6-9 (Check formatting, Build, Clippy, Test) are ALL success; the only anomalous entry is step 17 'Post Cache cargo' with status=in_progress, conclusion=null, and the NEXT step 'Post Run actions/checkout@v4' left pending. The job's completed_at is ~39 minutes after started_at. There is no error annotation and `gh run view --log-failed` prints NOTHING (no failed step exists to attribute the log to), which is the tell: the failure is not attached to any step.\n\nROOT CAUSE: the job was terminated while its POST step was still saving the cargo cache. actions/cache's post step uploads the configured paths; with `path: target` plus ~/.cargo/registry+git, the archive is gigabytes on a Rust workspace that vendors/bundles duckdb (target/ ~7.5GB locally), so the post step dominated the job wall time. When the job is cancelled/killed mid post-step, GitHub marks the JOB failed but leaves the interrupted post step with no conclusion \u2014 producing exactly this 'failure with nothing failed' signature. The workflow itself is sound: the identical steps in the previous green run show 'Post Cache cargo' success and step 19 'Complete job' success.\n\nDIAGNOSIS PROCEDURE (how to tell this class from a real red):\n1. `gh api repos/<org>/<repo>/actions/runs/<id>/jobs --jq '.jobs[] | \"\\(.name) \\(.status) \\(.conclusion)\", (.steps[] | \" \\(.number) \\(.name) \\(.conclusion)\")'` \u2014 inspect PER STEP, never trust the run/job conclusion.\n2. If every real step is success and a POST step is in_progress/null, it is infra/termination, not code. `--log-failed` returning empty output is corroboration, not a bug.\n3. Compare against the previous green run's post steps to prove the workflow is intact.\n4. Do NOT re-run the substantive steps locally hunting a code failure that does not exist; do NOT bounce the commit back to a worker.\n\nFIX / MITIGATION: stop caching the whole `target/` directory. Cache only `~/.cargo/registry` and `~/.cargo/git` (restore-keys keyed on Cargo.lock), or move the target cache behind a separate size-capped key: the compile-cache gain does not justify a 39-minute job tail and a class of spurious red badges. Verify the fix by two consecutive green runs whose step 'Complete job' is success and whose job duration drops sharply.\n\nBROADER LESSON: on any CI where a post/teardown step can exceed the runner's patience, the JOB conclusion is not evidence about the code. Gate on step-level conclusions.", "environment": "GitHub Actions, ubuntu-latest, actions/checkout@v4 + actions/cache@v4 caching a Rust target/ directory of ~7.5GB; repo gethilo/hilo (Rust workspace, duckdb-sys bundled build).", "language": "yaml", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "github-actions-job-failed-all-steps-passed-post-step", "provider": "openrouter", "solved_at": "2026-09-15T06:46:24.526Z", "version": ""}Diagnosed and wrote the verified solution to ~/ci-post-step-failure-diagnosis.md.
The described signature is exact: run 34901452923, job 104168118678 concluded failure while steps 1–9 were all success, and Post Cache cargo was left in_progress/null with Post Run actions/checkout@v4 pending. The previous green run 34860504327 has the identical step list but Post Cache cargo and Complete job both success.
I found the actual smoking gun, which sharpens the given root cause:
No space left on device while the runner wrote its own diagnostic log."/usr/bin/tar" failed with error ... exit code 2.So caching the ~7.5 GB target/ directory makes the post step tar a 10.77 GB archive onto a runner already holding the target tree — the disk fills, the runner dies mid-post-step, and GitHub marks the job failed with no failed step (hence empty --log-failed). The green run only survived because actions/cache reports a failed tar as a warning.
One line removed from .github/workflows/ci.yml:
path: |
~/.cargo/registry
~/.cargo/git
- target
I validated the resulting workflow parses as YAML and confirmed the cache path list no longer contains target. The doc also includes the optional Swatinem/rust-cache@v2 alternative if compile caching is needed, and a verification section (two consecutive green runs, step-level success including Complete job, cache total under 10 GB, no disk annotations).
# Evidence - Problem class: github-actions-job-failed-all-steps-passed-post-step - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-15T06:46:24.526Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: an Actions run on master reports conclusion=failure, so master looks red and every 'is CI green?' gate misreports a code failure. `gh run view --json jobs` shows the SINGLE job 'Build & Test (stable)' with conclusion=failure while steps 6-9 (Check formatting, Build, Clippy, Test) are ALL success; the only anomalous entry is step 17 'Post Cache cargo' with status=in_progress, conclusion=null, and the NEXT step 'Post Run actions/checkout@v4' left pending. The job's completed_at is ~39 minutes after started_at. There is no error annotation and `gh run view --log-failed` prints NOTHING (no failed step exists to attribute the log to), which is the tell: the failure is not attached to any step.\n\nROOT CAUSE: the job was terminated while its POST step was still saving the cargo cache. actions/cache's post step uploads the configured paths; with `path: target` plus ~/.cargo/registry+git, the archive is gigabytes on a Rust workspace that vendors/bundles duckdb (target/ ~7.5GB locally), so the post step dominated the job wall time. When the job is cancelled/killed mid post-step, GitHub marks the JOB failed but leaves the interrupted post step with no conclusion \u2014 producing exactly this 'failure with nothing failed' signature. The workflow itself is sound: the identical steps in the previous green run show 'Post Cache cargo' success and step 19 'Complete job' success.\n\nDIAGNOSIS PROCEDURE (how to tell this class from a real red):\n1. `gh api repos/<org>/<repo>/actions/runs/<id>/jobs --jq '.jobs[] | \"\\(.name) \\(.status) \\(.conclusion)\", (.steps[] | \" \\(.number) \\(.name) \\(.conclusion)\")'` \u2014 inspect PER STEP, never trust the run/job conclusion.\n2. If every real step is success and a POST step is in_progress/null, it is infra/termination, not code. `--log-failed` returning empty output is corroboration, not a bug.\n3. Compare against the previous green run's post steps to prove the workflow is intact.\n4. Do NOT re-run the substantive steps locally hunting a code failure that does not exist; do NOT bounce the commit back to a worker.\n\nFIX / MITIGATION: stop caching the whole `target/` directory. Cache only `~/.cargo/registry` and `~/.cargo/git` (restore-keys keyed on Cargo.lock), or move the target cache behind a separate size-capped key: the compile-cache gain does not justify a 39-minute job tail and a class of spurious red badges. Verify the fix by two consecutive green runs whose step 'Complete job' is success and whose job duration drops sharply.\n\nBROADER LESSON: on any CI where a post/teardown step can exceed the runner's patience, the JOB conclusion is not evidence about the code. Gate on step-level conclusions.", "environment": "GitHub Actions, ubuntu-latest, actions/checkout@v4 + actions/cache@v4 caching a Rust target/ directory of ~7.5GB; repo gethilo/hilo (Rust workspace, duckdb-sys bundled build).", "language": "yaml", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "github-actions-job-failed-all-steps-passed-post-step", "provider": "openrouter", "solved_at": "2026-09-15T06:46:24.526Z", "version": ""}