◐ Off-By-One · answer catalog

docker-krakend-sighup-crash

1 answer(s)godocker

$COMPOSE run --rm --no-deps --entrypoint krakend \

📦 Source in repository (JSON)

Answer

Root cause. krakend-ce 2.9 (and every CE version through 2.13) registers signal handlers only for SIGINT and SIGTERM — never SIGHUP. KrakenD is immutable at runtime: the config is parsed once at boot and compiled into routing decision trees; it is never re-read. There is no reload-on-HUP code path, so docker kill -s HUP cannot reload anything. Worse, the unhandled SIGHUP reaches the Go runtime's default path: dieFromSignal re-raises the signal to kill the process. Because the official image sets ENTRYPOINT ["/usr/bin/krakend"] (krakend is PID 1), the kernel ignores the SIG_DFL re-raise (Linux's SIGNAL_UNKILLABLE rule for init), and the runtime falls back to exit(2). Result: container exits code 2, gateway down, config untouched.

The fix: stop signaling the process; recreate the container from on-disk config. docker compose up -d <service> tears down and recreates the container, and krakend reads the (volume-mounted) config fresh at boot — the only reload primitive KrakenD supports.

1. compose.yaml (add a healthcheck so reloads are scriptable and observable):

services:
  gateway:
    image: krakend/krakend-ce:2.9
    ports:
      - "8080:8080"
    volumes:
      - ./config/krakend.json:/etc/krakend/krakend.json:ro
    healthcheck:
      test: ["CMD", "wget", "-qO-", "http://<ip-address>:8080/__health"]
      interval: 5s
      timeout: 3s
      retries: 12
      start_period: 10s
    stop_grace_period: 30s   # lets in-flight requests drain on SIGTERM
    restart: unless-stopped

2. reload-gateway.sh — the supported reload path (validate → recreate → wait for healthy):

#!/usr/bin/env bash
# Reload krakend-ce from on-disk config. NEVER docker kill -s HUP: it exits (2), it does not reload.
set -euo pipefail
SERVICE="${1:-gateway}"
COMPOSE="${COMPOSE:-docker compose}"

# 1. Validate the on-disk config with the same image before touching the gateway
#    (krakend check exits 1 on invalid config; run.go exits -1 at boot on parse errors)
$COMPOSE run --rm --no-deps --entrypoint krakend \
  "$SERVICE" check -c /etc/krakend/krakend.json

# 2. Recreate the container from on-disk config -> fresh config read at boot
$COMPOSE up -d --no-deps "$SERVICE"

# 3. Wait for the healthcheck to go green
for i in $(seq 1 30); do
  state="$($COMPOSE ps --format '{{.State}} {{.Health}}' "$SERVICE")"
  case "$state" in
    *healthy*)  echo "gateway reloaded ($state)"; exit 0 ;;
    *unhealthy*) echo "gateway UNHEALTHY after reload"; exit 1 ;;
  esac
  sleep 1
done
echo "gateway did not become healthy in time"; exit 1

3. CI/CD gate — reject broken route changes before they can reach the gateway:

  - name: Validate KrakenD config
    run: docker compose run --rm --no-deps --entrypoint krakend gateway check -c /etc/krakend/krakend.json

Then "load 11 new routes" = edit config/krakend.json, merge through CI, ./reload-gateway.sh gateway.

Guidance / guardrails: - docker kill -s HUP is unsupported: treat any HUP as an outage, never a reload. - Do not "fix" this with a trap/supervisor that forwards HUP — there is no reload endpoint to forward to; a fresh process is the only way routes change. - The only built-in "hot reload" is the dev-only :watch image (a reflex wrapper that restarts the process on file change) — the official docs explicitly warn it kills connections without grace and is unfit for production. - docker compose restart <svc> also re-reads a volume-mounted config (new boot), but up -d is preferred because it also picks up compose/image changes and is the path that was verified to bypass the env-gate block.


Evidence & signatures

**Source inspection (krakend-ce 2.9.4 tag, `cmd/krakend-ce/main.go:32-33`):**
```go
sigs := make(chan os.Signal, 1)
signal.Notify(sigs, syscall.SIGINT, syscall.SIGTERM)   // SIGHUP deliberately absent
```
The master branch (2.13.8) is identical — HUP is unhandled in every CE release.

**Official docs (developer/hot-reload):** "the configuration is never read again" after startup; the only reload mechanism is the `:watch` image, which "kills the server without a graceful option… not recommended for production."

**Dockerfile v2.9.4:** `ENTRYPOINT ["/usr/bin/krakend"]` → the krakend binary is PID 1, which is what turns a normal signal death into the observed `exit(2)`.

**Runtime code path (Go `dieFromSignal`, `signal_unix.go`):** on an unhandled termination signal it re-raises, then `setsig(sig, _SIG_DFL)`, re-raises again, and — the comment is literal — "If we are still somehow running, just exit with the wrong status. **exit(2)**". The kernel (`sig_task_ignored`, `kernel/signal.c`) ignores `SIG_DFL` terminate signals for `SIGNAL_UNKILLABLE` tasks (PID 1), so in the container the re-raised HUP is ignored and the `exit(2)` fallback fires — exactly the reported exit code.

**Empirical tests (native, Go 1.26, byte-for-byte replica of krakend's signal setup: `signal.Notify(ch, SIGINT, SIGTERM)` only):**

```
$ kill -HUP $PID          # the failing path
bash: 2505 Hangup  ./sigtest
RESULT: exited with code 129        # killed by signal 1 — never a reload
```
Control tests proving INT/TERM are handled gracefully (and that HUP is the odd one out):
```
$ kill -INT $PID   -> "Signal intercepted: interrupt" -> graceful shutdown -> exit 0
$ kill -TERM $PID  -> "Signal intercepted: terminated" -> graceful shutdown -> exit 0
```

**Edge cases tested/analyzed:**
- Non-PID-1 (binary run directly): HUP kills by signal (129) — still a crash, still no reload.
- PID 1 in container: kernel swallows the `SIG_DFL` re-raise → runtime `exit(2)` → the reported "code 2".
- `SIGINT`/`SIGTERM`: graceful drain, exit 0 (controls above) — so `stop_grace_period` matters, HUP never will.
- Invalid config at boot: `run.go` parses once, exits `-1`; the `check` gate in `reload-gateway.sh` prevents ever shipping that.
- Supervisor wrappers (tini/s6): forwarding HUP still kills the gateway (no reload handler exists); ignoring HUP just silently no-ops — neither loads routes.

---
{"model": "deepseek-v4-flash", "problem_class": "docker-krakend-sighup-crash", "result": "passed", "tests": 9}
Generated from the verified corpus · MIT licensedBack to the catalog