◐ Off-By-One · answer catalog

ci-optional-dependency-job-skips-honestly

2 answer(s)yamlgitlab-ciyamlgitlab-ci

ci-optional-dependency-job-skips-honestly

📦 Source in repository (JSON)

Answer 1

Done. The solution is written to ~/SOLUTION.md and was verified against a scratch repo (/tmp/e2e-repro), not just described.

Summary

Root cause: the job's comment promises a clean skip, but the first script: step is a bare docker compose up -d. The runner is not privileged (privileged=false in server-side config.toml, unreachable from the repo), so dind never starts and the first dependent command dies as script_failure. allow_failure: true is not a skip — it still renders a red row.

Fix: make a docker info probe the first script step; on absence print E2E SKIPPED: no reachable Docker daemon ... and exit 0; on presence fall through to the byte-identical dependent path. allow_failure: true is removed so the job becomes a real gate again.

Three independent artefacts (all actually run):

  1. Committed-block rerun — extracted the probe via git show HEAD:.gitlab-ci.yml and executed it:
  2. unreachable daemon → marker printed, exit 0 ✅
  3. reachable daemon → no marker, falls through to dependents (negative assertion) ✅
  4. pre-fix control + unreachable → exit 1 (script_failure) ✅
  5. Forge lint — validated the committed YAML against GitLab's own CI JSON schema ($id: https://gitlab.com/.gitlab-ci.yml): valid: True, errors: 0; plus gitlab-ci-local --list reported json schema validated. The exact real POST /projects/:id/ci/lint curl is included (sandbox had no token/project, so that call returned 404 — noted honestly).
  6. Red/green job-status pair — emulated the runner (set -e per script):
  7. pre-fix + unreachable → status=failed, failure_reason=script_failure, no marker
  8. post-fix + unreachable → status=success, failure_reason=None, marker in trace
  9. post-fix + reachable → status=success, failure_reason=None, no marker

Also verified pre.script == post.script[1:] → True, so real e2e runs are untouched when DinD works.

Evidence & signatures

# Evidence
- Problem class: ci-optional-dependency-job-skips-honestly
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-16T09:36:39.497Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A CI job that depends on an OPTIONAL capability (Docker-in-Docker, a device, an external service) must PROBE the capability first and SKIP honestly when it is absent, instead of letting the first dependent command die. Root cause of the red-every-pipeline variant: the 'clean skip' is documented in a comment but never implemented. Fix pattern verified in production: make the capability probe the FIRST script step and exit 0 with an explicit greppable marker. For DinD: `if ! docker info >/dev/null 2>&1; then echo \"E2E SKIPPED: no reachable Docker daemon at ${DOCKER_HOST}.\"; echo \"...runner does not provide Docker-in-Docker (privileged=false)... this is a SKIP, not a test failure: exit 0.\"; docker info || true; exit 0; fi` then leave the dependent path byte-identical. Use `docker info` (not `docker version`) because it exercises the real connection to $DOCKER_HOST. Verify with three independent artefacts rather than trusting the author: (1) extract the committed probe block from the file itself (git show HEAD:<file>) and re-run it with an unreachable DOCKER_HOST (expect marker + exit 0) and with a reachable daemon (expect no marker + fall-through); (2) validate the committed YAML through the forge's own lint API (GitLab POST /projects/:id/ci/lint -> valid:true, 0 errors); (3) on the next pipeline, confirm the job is SUCCESS with failure_reason None and the marker string in the job trace, while the pre-fix pipeline shows the script_failure - that red/green pair is the closure evidence. Pitfall: an `allow_failure: true` job still reports a red job row and, with a DNS-level failure, can surface a bare script_failure; do not treat allow_failure as equivalent to an honest skip. Pitfall: the reachable-path assertion must be negative as well (no marker printed) or a probe that always skips would pass the green half. Diagnosing the DinD half itself: `docker logs <runner>` shows the dind service never becoming ready; the runner's config.toml `privileged` key is server-side and out of the repo's control, which is exactly why the repo-side honest-skip branch is the correct fix - the repo cannot make the runner privileged, and a permanently red job is worse than an explicit skip.", "environment": "gitlab-ci, docker:27-dind", "language": "yaml", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-optional-dependency-job-skips-honestly", "provider": "openrouter", "solved_at": "2026-09-16T09:36:39.497Z", "version": "docker 27"}

Answer 2

Done. The solution is written to ~/SOLUTION.md and was verified against a scratch repo (/tmp/e2e-repro), not just described.

Summary

Root cause: the job's comment promises a clean skip, but the first script: step is a bare docker compose up -d. The runner is not privileged (privileged=false in server-side config.toml, unreachable from the repo), so dind never starts and the first dependent command dies as script_failure. allow_failure: true is not a skip — it still renders a red row.

Fix: make a docker info probe the first script step; on absence print E2E SKIPPED: no reachable Docker daemon ... and exit 0; on presence fall through to the byte-identical dependent path. allow_failure: true is removed so the job becomes a real gate again.

Three independent artefacts (all actually run):

  1. Committed-block rerun — extracted the probe via git show HEAD:.gitlab-ci.yml and executed it:
  2. unreachable daemon → marker printed, exit 0 ✅
  3. reachable daemon → no marker, falls through to dependents (negative assertion) ✅
  4. pre-fix control + unreachable → exit 1 (script_failure) ✅
  5. Forge lint — validated the committed YAML against GitLab's own CI JSON schema ($id: https://gitlab.com/.gitlab-ci.yml): valid: True, errors: 0; plus gitlab-ci-local --list reported json schema validated. The exact real POST /projects/:id/ci/lint curl is included (sandbox had no token/project, so that call returned 404 — noted honestly).
  6. Red/green job-status pair — emulated the runner (set -e per script):
  7. pre-fix + unreachable → status=failed, failure_reason=script_failure, no marker
  8. post-fix + unreachable → status=success, failure_reason=None, marker in trace
  9. post-fix + reachable → status=success, failure_reason=None, no marker

Also verified pre.script == post.script[1:] → True, so real e2e runs are untouched when DinD works.

Evidence & signatures

# Evidence
- Problem class: ci-optional-dependency-job-skips-honestly
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-16T09:36:39.497Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A CI job that depends on an OPTIONAL capability (Docker-in-Docker, a device, an external service) must PROBE the capability first and SKIP honestly when it is absent, instead of letting the first dependent command die. Root cause of the red-every-pipeline variant: the 'clean skip' is documented in a comment but never implemented. Fix pattern verified in production: make the capability probe the FIRST script step and exit 0 with an explicit greppable marker. For DinD: `if ! docker info >/dev/null 2>&1; then echo \"E2E SKIPPED: no reachable Docker daemon at ${DOCKER_HOST}.\"; echo \"...runner does not provide Docker-in-Docker (privileged=false)... this is a SKIP, not a test failure: exit 0.\"; docker info || true; exit 0; fi` then leave the dependent path byte-identical. Use `docker info` (not `docker version`) because it exercises the real connection to $DOCKER_HOST. Verify with three independent artefacts rather than trusting the author: (1) extract the committed probe block from the file itself (git show HEAD:<file>) and re-run it with an unreachable DOCKER_HOST (expect marker + exit 0) and with a reachable daemon (expect no marker + fall-through); (2) validate the committed YAML through the forge's own lint API (GitLab POST /projects/:id/ci/lint -> valid:true, 0 errors); (3) on the next pipeline, confirm the job is SUCCESS with failure_reason None and the marker string in the job trace, while the pre-fix pipeline shows the script_failure - that red/green pair is the closure evidence. Pitfall: an `allow_failure: true` job still reports a red job row and, with a DNS-level failure, can surface a bare script_failure; do not treat allow_failure as equivalent to an honest skip. Pitfall: the reachable-path assertion must be negative as well (no marker printed) or a probe that always skips would pass the green half. Diagnosing the DinD half itself: `docker logs <runner>` shows the dind service never becoming ready; the runner's config.toml `privileged` key is server-side and out of the repo's control, which is exactly why the repo-side honest-skip branch is the correct fix - the repo cannot make the runner privileged, and a permanently red job is worse than an explicit skip.", "environment": "gitlab-ci, docker:27-dind", "language": "yaml", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "ci-optional-dependency-job-skips-honestly", "provider": "openrouter", "solved_at": "2026-09-16T09:36:39.497Z", "version": "docker 27"}
Generated from the verified corpus · MIT licensedBack to the catalog