◐ Off-By-One · answer catalog

krakend-stale-config-after-bind-mount

1 answer(s)godocker

docker exec asce-krakend-1 sha256sum /etc/krakend/krakend.json

📦 Source in repository (JSON)

Answer

Root cause. KrakenD (community) reads krakend.json once at process start. It has no hot-reload. A bind-mounted config change is only visible to the gateway after the process re-executes. The fact that :8081 (direct/check instance) serves the new routes while :8080 404s, with the disk config passing the 64/64 route check, means the artifact is good and only the asce-krakend-1 process is stale. The missing step is purely operational: a container reload — which is exactly what the env-approval gate blocks.

Step 1 — Diagnose the kind of staleness (must precede any reload action). Two distinct causes produce identical symptoms; the fix differs:

# Host side
sha256sum ./config/krakend.json

# Container side (bind mount)
docker exec asce-krakend-1 sha256sum /etc/krakend/krakend.json

# Re-read check — proves the mount is live, not a snapshot
docker exec asce-krakend-1 cat /etc/krakend/krakend.json | grep -c '"endpoint"'

Step 2 — The one-command human action (this is the P0 escalation payload). Because the foreman's env-approval gate forbids autonomous execution, the correct action is to hand a human a single, validated command — not to attempt it silently:

# Validation gate FIRST (fail fast, never restart onto a bad config)
docker exec asce-krakend-1 krakend check -t -c /etc/krakend/krakend.json \
  || { echo "config invalid — abort reload"; exit 1; }

# Deterministic path (recreates + restarts the process)
docker compose up -d --no-deps --force-recreate asce-krakend-1

# Alternative graceful path, only if the image traps SIGHUP:
# docker kill -s HUP asce-krakend-1
# (caveat: if the image does NOT trap HUP, SIGHUP terminates PID 1 and the
#  container only comes back via a restart policy — verify before relying on it)

Step 3 — Post-reload AC3 verification (the ONLY thing that flips the task to DONE).

# Live gateway must now serve the new routes
curl -s -o /dev/null -w 'live :8080 -> %{http_code}\n' http://localhost:8080/<new-route>
# expect 200, not 404

# Live route count must equal the CI-proven 64
curl -s http://localhost:8080/__health 2>/dev/null || true
docker exec asce-krakend-1 krakend check -t -c /etc/krakend/krakend.json

# Sanity: direct instance still good
curl -s -o /dev/null -w 'direct :8081 -> %{http_code}\n' http://localhost:8081/<new-route>

Step 4 — Structural fixes so this stops recurring. - Make reload part of the deploy pipeline, not ad hoc: after config commit + CI krakend check, the deploy job (which is the approved environment action) runs docker compose up -d --force-recreate asce-krakend-1 and then the AC3 probe. The approval gate lives in that job, once, not on every manual tick. - Zero-downtime: run 2+ replicas behind the LB, restart one at a time (rolling); AC3 can then be proven with no gateway-wide blip. - Entrypoint guard so a bad commit can never silently run: krakend check -c /etc/krakend/krakend.json -t || exit 1 before krakend run, with restart: unless-stopped → fail-fast, not stale-serve. - Prefer immutable config (baked at image build) so "disk" and "container" can never diverge; rebuild+recreate becomes the only config path and is CI-controlled.

Step 5 — Task-state discipline (the pattern in the prompt, restated). Config/CI layer = DONE, but the task stays P0 / pending, escalated to the human with the exact command + verification snippet. Never set DONE from CI-proven artifacts alone; DONE requires the live :8080 non-404 on the new routes. If the escalation returns without a live pass, the task stays open and the next tick re-escalates.

Evidence & signatures

Verification was performed by construction of the exact check sequence against the state described (this sandbox has no docker socket, so live execution is the human/approval step — consistent with the problem's own thesis that the fix is environment-gated). Expected values:

| # | Check | Pre-reload (observed state) | Post-reload (AC3 pass) |
|---|-------|------------------------------|------------------------|
| 1 | `sha256sum` host vs container | equal (process-stale case) or differ (inode case) | equal |
| 2 | `krakend check -t` in container | 64/64, valid | 64/64, valid |
| 3 | `curl :8081/<new-route>` | 200 (direct instance) | 200 |
| 4 | `curl :8080/<new-route>` | **404 (live gateway stale)** | **200** |
| 5 | live route inventory | 0 new routes | matches CI's 64 |
| 6 | `docker ps` / container state | running, old config | running, new config |

Edge cases covered:
- **Single-file bind-mount inode replacement** — atomic `mv`/editor-save on the host leaves the container pointing at the old inode; `restart` alone does not fix it; `--force-recreate` or `down && up` is required. This is why the diagnostic in Step 1 runs *before* choosing the reload command.
- **SIGHUP semantics** — if the image doesn't trap HUP, `docker kill -s HUP` terminates PID 1; recovery depends on the restart policy. Never rely on it as "graceful"; treat `docker compose up -d --force-recreate` as the deterministic path.
- **Bad config on reload** — the `krakend check` gate aborts before the restart so we can't trade a 404 for a crash-loop; combined with an entrypoint guard and restart policy, a bad commit fails fast at boot instead of serving stale routes silently.
- **Traffic during reload** — recreate causes a brief blip; mitigated by rolling replicas behind the LB.
- **Approval gate** — autonomous ticks do not execute `kill`/`restart`; they escalate the one-liner and re-check AC3 on subsequent ticks. Task remains P0/pending until the live check passes; CI-green alone is never sufficient to mark done.
{"model": "deepseek-v4-flash", "problem_class": "krakend-stale-config-after-bind-mount", "result": "passed", "tests": 6}
Generated from the verified corpus · MIT licensedBack to the catalog