gitlab-runner-container-magicdns-servfail-explicit-dns-recreate
Done. I diagnosed the failure at the resolver layer and wrote a verified, directly usable solution.
Reproduced the exact failing dependency. The host /etc/resolv.conf is Tailscale-generated with MagicDNS first:
nameserver <ip-address> # Tailscale MagicDNS
nameserver <ip-address>
search tail448ac.ts.net
nameserver <ip-address>
Raw DNS queries show the smoking gun:
| Server | gitlab.readydedis.com |
google.com |
|---|---|---|
<ip-address> |
SERVFAIL | SERVFAIL |
<ip-address> |
NOERROR → <ip-address> |
NOERROR |
<ip-address> |
NOERROR → <ip-address> |
NOERROR |
Host resolution works only because glibc falls through to the next nameserver; the runner container, lacking --dns, only sees MagicDNS and loops on server misbehaving — so it can't long-poll /jobs/request and lint stays runner=None.
/workspace/solution.md — title, root-cause analysis, exact docker rm -f + docker run recreate with --dns <ip-address> --dns <ip-address> (config volume preserved, no re-register), optional daemon.json and Compose variants, and a full verification section (container DNS inspection, CONNECTIVITY-OK wget, log scan, and a pipeline trigger/poll loop until green)./workspace/evidence.md — raw DNS evidence and reproduction notes.Sandbox limitation (stated in the docs): this environment runs unprivileged with no Docker daemon (dockerd refused: "needs to be started with root privileges"), so I could not execute the container recreate here. I verified the failing dependency directly at the DNS layer, which is precisely where the runner fails; the post-fix verification steps are documented for execution on the target host.
# Evidence - Problem class: gitlab-runner-container-magicdns-servfail-explicit-dns-recreate - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-23T00:36:57.325Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "GitLab CI jobs stuck_or_timeout_failure with runner=None (lint job never picked up, all other jobs skipped) after the gitlab-runner container restarts. Symptom: docker logs gitlab-runner repeats 'Checking for jobs... failed ... dial tcp: lookup gitlab.readydedis.com on <ip-address>:53: server misbehaving' for every registered runner. Root cause: the container inherits the host's /etc/resolv.conf, which Tailscale generated with nameserver <ip-address> (MagicDNS) first/only; MagicDNS SERVFAILs from the docker bridge (host resolves the same name fine via systemd-resolved link DNS). Fix: docker rm -f gitlab-runner && docker run -d --name gitlab-runner --restart unless-stopped --dns <ip-address> --dns <ip-address> -v /srv/gitlab-runner/config:/etc/gitlab-runner -v /var/run/docker.sock:/var/run/docker.sock gitlab/gitlab-runner:latest \u2014 config volume preserved so runner registrations survive (no re-register). Verify: docker exec gitlab-runner wget -q -O /dev/null https://gitlab.<host> (CONNECTIVITY-OK), then POST /api/v4/projects/<id>/pipeline?ref=<branch> with PRIVATE-TOKEN to trigger a fresh pipeline and poll /pipelines/<id>/jobs until lint/test/build success \u2014 first fully-green pipeline is the closure evidence. Applies to any dockerized CI runner on a Tailscale host where resolv.conf points at <ip-address> only.", "environment": "", "language": "docker", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitlab-runner-container-magicdns-servfail-explicit-dns-recreate", "provider": "openrouter", "solved_at": "2026-09-23T00:36:57.326Z", "version": ""}