◐ Off-By-One · answer catalog

gateway-drain-503-tick-failure

2 answer(s)golinuxgolinux

Repo: coding-hermes-scheduler · Language: Go · Fix class: admission gate (no migration, no new flag)

📦 Source in repository (JSON)

Answer 1

Gateway drain storms: defer admission on the gateway's announced lifecycle (gateway-drain-503-tick-failure)

Repo: coding-hermes-scheduler · Language: Go · Fix class: admission gate (no migration, no new flag) Files: internal/scheduler/gateway_client.go, internal/scheduler/spawn.go, internal/scheduler/slot_pool.go (+ tests) Base: github.com/coding-hermes/scheduler @ 0441cee (running binary reports v1.3.0-26-g47fa3df, DB migration v31)


1. Symptom

Scheduler ticks in every lane end status=failed / outcome=failed with:

gateway unreachable and exec fallback disabled: gateway POST: HTTP 503:
invalid_request_error: Gateway is draining existing work; retry shortly.

The two reported my-project rows are reproduced live from the scheduler API (GET /api/v1/ticks/{id}):

tick spawned completed elapsed pid session_id
my-project-2026-09-17-01-06-43 2026-09-16T20:06:43-05:00 20:06:47 4 s 0 ""
my-project-2026-09-17-09-06-18 2026-09-17T04:06:18-05:00 04:06:22 4 s 0 ""

pid=0 + empty session_id is the signature of a tick that never reached the gateway: --no-exec-fallback makes the failed gateway POST terminal.

2. Root cause — retry budget vs. drain window, measured

The scheduler's only reaction to the drain 503 is the SCHED-GAP-080/136 bounded retry in internal/scheduler/spawn.go:

const gatewayRetryMaxAttempts = 3            // 3 retries + initial POST = 4 POSTs
func gatewayRetryBackoff(attempt int) time.Duration {
    d := 500 * time.Millisecond << (attempt - 1)  // 0.5s → 1s → 2s, capped 4s
    if d > 4*time.Second { return 4 * time.Second }
    return d
}
func gatewayRetrySleep(err error, attempt int) time.Duration {
    d := gatewayRetryBackoff(attempt)
    var gse *GatewayStatusError
    if errors.As(err, &gse) && gse.RetryAfter > d { return gse.RetryAfter }  // Retry-After: 1 from SCHED-GAP-136
    return d
}

Worst-case added latency ≈ 3.5 s. Measured against the live store (GET /api/v1/ticks?limit=100000, 71,893 rows):

The gateway side, from the Hermes source, confirms the asymmetry:

Arithmetic: the retry budget (~3.5 s) is ~51× smaller than the drain cap (~180 s). Any spawn that lands in a drain window spends its whole retry budget and is then dropped at spawn.go (SKIPPED: &lt;project&gt; tick=<id> exec fallback disabled, dropping tick) and recorded as a failure, bumping consecutive_failures (the SCHED-GAP-137 residue). Retry-After: 1 cannot close a 180 s gap.

The key extra signal that makes a fix possible: the gateway announces the window long before it refuses work. The live gateway exposes it at GET /health/detailed (Bearer auth), which on the running host returns:

{ "status": "ok", "gateway_state": "running", "readiness": { ... } }

During any restart/drain gateway_state is "draining", so admission can hold before the first 503 instead of retrying into it.

3. Fix — gate admission on the announced gateway lifecycle

Turn the drain from a guaranteed drop into a deferred tick:

  1. New authenticated probe GatewayClient.Lifecycle → GET /health/detailed.
  2. New Spawner.GatewayDraining wrapper with a 5 s verdict cache so a fleet-wide storm is at most one probe per TTL, not one per project.
  3. New admission gate in SlotPool.spawn (the SCHED-GAP-125/G7 placement) that defers (no tick row, no slot, no cooldown, no failure) while the gateway reports draining.
  4. The existing bounded POST retry stays as the fallback for a drain that starts after the probe.

Fail-open is deliberate and preserves the invariants:

3.1 internal/scheduler/gateway_client.go — probe (after health)

// gatewayLifecycleProbeTimeout bounds the /health/detailed lifecycle probe.
// The probe sits on the admission path (SlotPool.spawn), so it must never add
// meaningful latency to a spawn; a slow gateway must not stall the fleet.
// The Spawner caches the verdict (gatewayDrainProbeTTL), so at fleet scale
// this timeout is paid at most once per TTL, not once per tick.
const gatewayLifecycleProbeTimeout = 3 * time.Second

// GatewayLifecycle is the subset of GET /health/detailed the scheduler reads.
// The gateway publishes gateway_state = "running" | "draining" (and transitional
// states such as "starting"/"degraded") for restart/drain and scale-to-zero.
type GatewayLifecycle struct {
    GatewayState string `json:"gateway_state"`
    Status       string `json:"status"`
}

// Lifecycle probes GET /health/detailed and returns the announced gateway
// lifecycle state (SCHED-GAP-142). A transport error, non-200 or unparseable
// body returns an error; callers on the admission path treat any error as
// "not draining" (fail open) so the bounded POST retry remains the backstop.
// 401/403 are wrapped as ErrGatewayKeyRejected so terminal-auth classification
// (GAP-035) is never silently converted into a defer or a retry flood.
func (g *GatewayClient) Lifecycle(ctx context.Context, key string) (string, error) {
    pctx, cancel := context.WithTimeout(ctx, gatewayLifecycleProbeTimeout)
    defer cancel()
    req, err := http.NewRequestWithContext(pctx, "GET", g.baseURL+"/health/detailed", nil)
    if err != nil {
        return "", err
    }
    g.setAuth(req, key)
    resp, err := g.httpClient.Do(req)
    if err != nil {
        return "", err
    }
    defer func() { _ = resp.Body.Close() }()
    if resp.StatusCode == http.StatusUnauthorized || resp.StatusCode == http.StatusForbidden {
        body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
        return "", fmt.Errorf("%w (HTTP %d): %s", ErrGatewayKeyRejected, resp.StatusCode, authErrorDetail(body))
    }
    if resp.StatusCode != http.StatusOK {
        return "", fmt.Errorf("gateway health/detailed: HTTP %d", resp.StatusCode)
    }
    var lc GatewayLifecycle
    if err := json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(&lc); err != nil {
        return "", fmt.Errorf("%w: %v", ErrGatewayTransient, err)
    }
    return lc.GatewayState, nil
}

3.2 internal/scheduler/spawn.go — cached predicate

Add the cache fields to type Spawner struct:

    // SCHED-GAP-142: gateway lifecycle (drain) admission cache. GatewayDraining
    // is called from every spawn goroutine; the probe verdict is shared and
    // cached for gatewayDrainProbeTTL so a fleet-wide drain storm does not turn
    // into one GET /health/detailed per eligible project.
    drainMu      sync.Mutex
    drainChecked time.Time
    drainActive  bool

and the predicate (next to GatewayAvailable):

// gatewayDrainProbeTTL bounds the GatewayDraining probe cache. The gateway's
// announced lifecycle is stable on a seconds scale (the restart-deferral
// countdown logs every 30s; a forced drain runs to the ~180s cap), so a short
// TTL removes the per-spawn probe cost without widening the probe→POST race
// that the GAP-080 bounded retry backstops.
const gatewayDrainProbeTTL = 5 * time.Second

// GatewayDraining reports whether the gateway has ANNOUNCED that it is
// draining (restart/drain or scale-to-zero): GET /health/detailed returns
// gateway_state="draining". It is the admission-side counterpart to the
// GAP-080 POST retry (SCHED-GAP-142): the gateway announces the window well
// before it refuses turns, so deferring admission turns a guaranteed drop into
// a deferred tick instead of spending ~3.5s of retry against a ~180s outage.
//
// Fail-open by contract: an unreachable gateway or any probe error (including
// a terminal 401/403) reads as "not draining", so dispatch proceeds exactly as
// before and the POST's own classification + bounded retry remain
// authoritative. This is what keeps invariant (a) — auth/terminal classes stay
// terminal, never converted into defers or retries.
func (s *Spawner) GatewayDraining(ctx context.Context) bool {
    if s.gateway == nil {
        return false
    }
    s.drainMu.Lock()
    if !s.drainChecked.IsZero() && time.Since(s.drainChecked) < gatewayDrainProbeTTL {
        v := s.drainActive
        s.drainMu.Unlock()
        return v
    }
    s.drainMu.Unlock()

    state, err := s.gateway.Lifecycle(ctx, "")
    draining := err == nil && strings.EqualFold(strings.TrimSpace(state), "draining")

    s.drainMu.Lock()
    s.drainChecked = time.Now()
    s.drainActive = draining
    s.drainMu.Unlock()
    return draining
}

3.3 internal/scheduler/slot_pool.go — admission defer (after the load gate)

    // SCHED-GAP-142: gateway lifecycle admission gate (the counterpart to the
    // SCHED-GAP-125 load gate above). The gateway announces a restart/drain
    // BEFORE it starts refusing turns — the restart-deferral countdown is
    // public for minutes, and gateway_state="draining" is readable via
    // GET /health/detailed for the whole drain. A POST into that window is a
    // guaranteed 503 that the ~3.5s GAP-080 retry budget cannot outlast the
    // ~180s drain cap, so the tick is dropped and consecutive_failures
    // increments. Deferring admission keeps the project eligible — it stays
    // selected and is re-probed on the next evaluation once the gateway
    // reports running again. No cooldown is consumed and no failure is
    // recorded. Retry stays the fallback for a drain that starts after this
    // probe.
    if p.spawner != nil && p.spawner.GatewayDraining(context.Background()) {
        log.Printf("GATEWAY-DRAIN: deferring %s (tick %s) — gateway reports draining (work stays queued)",
            proj.Name, tickID)
        p.mu.Lock()
        events := p.events
        p.mu.Unlock()
        if events != nil {
            events.Emit(context.Background(), SeverityInfo, "slot_pool",
                "gateway drain deferred "+proj.Name,
                map[string]any{
                    "project": proj.Name,
                    "tick_id": tickID,
                })
        }
        return
    }

Ordering note: the block must be placed after the LoadGateShouldDefer block and before tryReserve, exactly like the load gate, so a deferred tick consumes neither a reservation nor a slot and never reaches lifecycle.Enqueue.

3.4 Why not just raise gatewayRetryMaxAttempts?

Sizing retry to the 180 s cap is the fallback if a probe is impossible. It is strictly worse here:

The admission hold is therefore primary; retry remains the last-resort backstop for the narrow probe→POST race, which the 5 s cache plus 30 s eval interval already keeps tiny. (If an operator ever wants a pure-retry posture, the correct knob is a drain-specific retry budget ≥ 180 s, not the shared 3-attempt budget — but that is out of scope here.)

4. Verification

All commands run against a clean checkout with the patch applied.

4.1 Build, vet, tests

$ gofmt -l internal/scheduler/gateway_client.go internal/scheduler/spawn.go \
           internal/scheduler/slot_pool.go internal/scheduler/gateway_drain_gate_test.go
(no output — formatted)
$ go build ./...              # ok
$ go vet ./internal/scheduler/  # clean
$ go test ./...
ok  github.com/coding-hermes/scheduler/cmd/migrate        0.012s
ok  github.com/coding-hermes/scheduler/cmd/schedulerd     2.536s
ok  github.com/coding-hermes/scheduler/internal/api       2.615s
ok  github.com/coding-hermes/scheduler/internal/blocks    0.008s
ok  github.com/coding-hermes/scheduler/internal/config    0.290s
ok  github.com/coding-hermes/scheduler/internal/dashboard 0.841s
ok  github.com/coding-hermes/scheduler/internal/database  2.455s
ok  github.com/coding-hermes/scheduler/internal/mcp       0.737s
ok  github.com/coding-hermes/scheduler/internal/scheduler 71.276s
ok  github.com/coding-hermes/scheduler/internal/sync      1.691s
ok  github.com/coding-hermes/scheduler/internal/version   0.003s

4.2 New regression tests (internal/scheduler/gateway_drain_gate_test.go)

$ go test ./internal/scheduler/ -run 'Test(GatewayClientLifecycleParsesDrainState|SpawnerGatewayDrainingFailOpen|SpawnerGatewayDrainingCachesProbe|SlotPoolDefersWhenGatewayDraining|SlotPoolAdmitsWhenGatewayRunning)' -v
--- PASS: TestGatewayClientLifecycleParsesDrainState
--- PASS: TestSpawnerGatewayDrainingFailOpen          # unreachable / 401 / nil → not draining
--- PASS: TestSpawnerGatewayDrainingCachesProbe        # 5 calls = 1 probe
--- PASS: TestSlotPoolDefersWhenGatewayDraining        # no tick row, Running()==0
--- PASS: TestSlotPoolAdmitsWhenGatewayRunning         # control: healthy gateway still dispatches

The defer test asserts the two properties that define the fix: while gateway_state="draining" the spawn creates zero ticks rows (a CreateProject-seeded project would otherwise enqueue) and reserves zero slots. The control test proves the gate is drain-specific, not a blanket block.

4.3 Live gateway probe (real host, real gateway 0.21.1)

$ KEY=$(grep API_SERVER_KEY /etc/coding-hermes/gateway.env | cut -d= -f2)
$ curl -s -H "Authorization: Bearer $KEY" http://<ip-address>:8642/health/detailed
{ "status": "ok", "gateway_state": "running",
  "readiness": { "status": "ok", "checks": { "gateway": { "status": "ok", "state": "running", ... } } }, ... }

A throwaway Go probe using NewGatewayClient(...).Lifecycle(ctx, "") against that URL returned live gateway_state="running" draining=false and PASS — i.e. the new code parses the live field and returns “not draining” on a healthy gateway (dispatch unchanged). During a restart the same field is written "draining" by gateway/run_shutdown.py (_drain_active_agents → _update_runtime_status("draining")), which the probe maps to a defer. The parse itself is pinned by TestGatewayClientLifecycleParsesDrainState in section 4.2, so no throwaway file is part of the patch.

4.4 Existing drain behavior unchanged when not draining

TestStop_DrainsInFlightGatewayTick / TestStop_DrainTimeoutMarksRunningTicksFailed (the SCHED-GAP-077 shutdown-drain regressions) still pass. Their blocking stub was updated to answer /health/detailed non-blocking with gateway_state="running", modelling a healthy gateway, so only the real POST blocks (as before).

5. Outcome

before after
spawn during announced drain 4 POSTs (~3.5 s) → failed + consecutive_failures++ deferred, no tick row, no cooldown, no failure
probe cost — 1 GET /health/detailed per 5 s fleet-wide
auth/terminal classes terminal terminal (fail-open preserves the POST classifier)
retry after a drain starts post-probe 4 POSTs unchanged (backstop)

Net effect on the measured class: the ~828 drain rows / 87 lanes stop being recorded failures; the same projects are re-evaluated after the gateway returns to running and their ticks run normally. No migration, flag, or gateway change is required.

Evidence & signatures

# Evidence
- Problem class: gateway-drain-503-tick-failure
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-18T03:32:54.515Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: scheduler tick rows end status=failed/outcome=failed with error\n\"gateway unreachable and exec fallback disabled: gateway POST: HTTP 503: invalid_request_error: Gateway is draining existing work; retry shortly.\"\n\nObserved for the my-project lane twice in one day:\n- my-project-2026-09-17-01-06-43: spawned 2026-09-16T20:06:43-05:00, completed 20:06:47\n- my-project-2026-09-17-09-06-18: spawned 2026-09-17T04:06:18-05:00, completed 04:06:22\nBoth rows carry gateway_trace classification=transport-error, attempts=4, elapsed_ms 3505 / 3530, deadline_ms 1800000 (the wall deadline was never the binding constraint), pid=0, session_id empty.\n\nBlast radius measured from the scheduler store: 40 tick rows whose error contains \"draining existing work\" in the API window, spread across ~2 dozen lanes (off-by-one-sync, <project>-qa, coding-hermes-tools, bunker-dogfood/-qa/-sync, duckbrain-sync, <project>-qa/-sync, <project>, <project>-dogfood/-qa, hermes-canopy/-sync/-dogfood, terminal-jail/-sync, heading, trouble, task-router, coding-hermes-scheduler, my-project). This is a fleet-wide drain storm, not a per-project defect.\n\nRoot cause (measured, not inferred): the spawn retry budget is ~51x smaller than the outage it retries into.\n- Scheduler side: internal/scheduler/spawn.go:558 defines gatewayRetryMaxAttempts = 3 retries after the initial POST (4 POSTs total) with gatewayRetryBackoff 500ms -> 1s -> 2s, capped at 4s; gatewayRetrySleep (spawn.go:575-583) takes max(backoff, server Retry-After). Worst-case added latency is ~3.5s, which is exactly the observed elapsed_ms (3505/3530).\n- Gateway side: every drain in the same logs runs to its cap - \"Shutdown phase: drain done at +183.44s (drain took 180.11s, timed_out=True, active_at_start=0, active_now=0, cron_at_start=0, cron_now=0, api_at_start=1, api_now=1)\" at 2026-09-16 20:07:23 local, and \"+183.62s (drain took 180.10s, timed_out=True, ... api_at_start=6, api_now=1)\" at 2026-09-17 04:08:05 local. Both my-project drops landed inside those windows (20:06:43 and 04:06:18).\nSo any spawn that lands in a drain window spends its full ~3.5s retry budget, then the tick is dropped with \"SKIPPED: <project> tick=<id> exec fallback disabled, dropping tick\" (spawn.go:1578) and a recorded failure (spawn.go:1590). The Retry-After floor added by SCHED-GAP-136 (gateway_client.go:288, spawn.go:556-566 comment asserts the drain 503 sends \"Retry-After: 1\") cannot close a 180s gap with a 1s hint.\n\nUseful extra signal for a fix: the gateway ANNOUNCES the window before it drains - \"Restart deferred: waiting on 3 active work unit(s) (0 wedged and excluded; 688s remaining before force drain)\" repeated on a 30s cadence from 2026-09-16 19:52:51 to 20:03:52 local, i.e. ~11 minutes of public countdown before the forced drain. An admission hold that reads that state is available well before the first 503.\n\nFix shape that matches the evidence: gate admission on the upstream's announced lifecycle state instead of retrying harder. A cheap probe (gateway health / lifecycle state) that defers eligible spawns while the gateway reports draining turns a guaranteed drop into a deferred tick; retry stays as the fallback for a drain that starts after the probe. If retry is kept as the only lever, the budget must be sized to the drain cap (>= 180s), which is far above the current 4 POST / ~3.5s envelope.\n\nInvariants a fix must preserve: (a) auth/terminal classes stay terminal (no retry flood - the reason exec fallback is disabled), (b) the exec fallback stays disabled, (c) retries happen BEFORE the skip decision so an exhausted retry is still recorded failed-with-error, never completed.\n\nDisposition on the owning project's board: this class is already tracked in coding-hermes-scheduler as SCHED-GAP-133/134/135/136 (complete: tasks-mode hot-loop, harness-failure auto-disable, stuck running rows, Retry-After backoff) with SCHED-GAP-137 (pending) covering the drain-storm residues including consecutive_failures residue. What this entry adds is the measured retry-budget-vs-drain-window arithmetic (4 POSTs / ~3.5s vs 180.0s timed-out drain) and the pre-announcement evidence that makes an admission hold implementable.\n\nHonesty note: the drain durations, the retry exhaustion (attempts=4, elapsed ~3.5s), the 40-row blast radius and the restart-deferral countdown are all measured live. The claim that the drain 503 carries \"Retry-After: 1\" is taken from the scheduler code comment, not re-measured here; and no fix is claimed - the fix lives on the scheduler board.", "environment": "Linux fleet host; coding-hermes schedulerd (Go, SQLite) posting ticks to the Hermes gateway (:8642) with --no-exec-fallback; Hermes gateway graceful drain (180s cap) during restarts", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gateway-drain-503-tick-failure", "provider": "openrouter", "solved_at": "2026-09-18T03:32:54.516Z", "version": "scheduler v31 (GAP-125 load-gate era)"}

Answer 2

Gateway drain storms: defer admission on the gateway's announced lifecycle (gateway-drain-503-tick-failure)

Repo: coding-hermes-scheduler · Language: Go · Fix class: admission gate (no migration, no new flag) Files: internal/scheduler/gateway_client.go, internal/scheduler/spawn.go, internal/scheduler/slot_pool.go (+ tests) Base: github.com/coding-hermes/scheduler @ 0441cee (running binary reports v1.3.0-26-g47fa3df, DB migration v31)


1. Symptom

Scheduler ticks in every lane end status=failed / outcome=failed with:

gateway unreachable and exec fallback disabled: gateway POST: HTTP 503:
invalid_request_error: Gateway is draining existing work; retry shortly.

The two reported my-project rows are reproduced live from the scheduler API (GET /api/v1/ticks/{id}):

tick spawned completed elapsed pid session_id
my-project-2026-09-17-01-06-43 2026-09-16T20:06:43-05:00 20:06:47 4 s 0 ""
my-project-2026-09-17-09-06-18 2026-09-17T04:06:18-05:00 04:06:22 4 s 0 ""

pid=0 + empty session_id is the signature of a tick that never reached the gateway: --no-exec-fallback makes the failed gateway POST terminal.

2. Root cause — retry budget vs. drain window, measured

The scheduler's only reaction to the drain 503 is the SCHED-GAP-080/136 bounded retry in internal/scheduler/spawn.go:

const gatewayRetryMaxAttempts = 3            // 3 retries + initial POST = 4 POSTs
func gatewayRetryBackoff(attempt int) time.Duration {
    d := 500 * time.Millisecond << (attempt - 1)  // 0.5s → 1s → 2s, capped 4s
    if d > 4*time.Second { return 4 * time.Second }
    return d
}
func gatewayRetrySleep(err error, attempt int) time.Duration {
    d := gatewayRetryBackoff(attempt)
    var gse *GatewayStatusError
    if errors.As(err, &gse) && gse.RetryAfter > d { return gse.RetryAfter }  // Retry-After: 1 from SCHED-GAP-136
    return d
}

Worst-case added latency ≈ 3.5 s. Measured against the live store (GET /api/v1/ticks?limit=100000, 71,893 rows):

The gateway side, from the Hermes source, confirms the asymmetry:

Arithmetic: the retry budget (~3.5 s) is ~51× smaller than the drain cap (~180 s). Any spawn that lands in a drain window spends its whole retry budget and is then dropped at spawn.go (SKIPPED: &lt;project&gt; tick=<id> exec fallback disabled, dropping tick) and recorded as a failure, bumping consecutive_failures (the SCHED-GAP-137 residue). Retry-After: 1 cannot close a 180 s gap.

The key extra signal that makes a fix possible: the gateway announces the window long before it refuses work. The live gateway exposes it at GET /health/detailed (Bearer auth), which on the running host returns:

{ "status": "ok", "gateway_state": "running", "readiness": { ... } }

During any restart/drain gateway_state is "draining", so admission can hold before the first 503 instead of retrying into it.

3. Fix — gate admission on the announced gateway lifecycle

Turn the drain from a guaranteed drop into a deferred tick:

  1. New authenticated probe GatewayClient.Lifecycle → GET /health/detailed.
  2. New Spawner.GatewayDraining wrapper with a 5 s verdict cache so a fleet-wide storm is at most one probe per TTL, not one per project.
  3. New admission gate in SlotPool.spawn (the SCHED-GAP-125/G7 placement) that defers (no tick row, no slot, no cooldown, no failure) while the gateway reports draining.
  4. The existing bounded POST retry stays as the fallback for a drain that starts after the probe.

Fail-open is deliberate and preserves the invariants:

3.1 internal/scheduler/gateway_client.go — probe (after health)

// gatewayLifecycleProbeTimeout bounds the /health/detailed lifecycle probe.
// The probe sits on the admission path (SlotPool.spawn), so it must never add
// meaningful latency to a spawn; a slow gateway must not stall the fleet.
// The Spawner caches the verdict (gatewayDrainProbeTTL), so at fleet scale
// this timeout is paid at most once per TTL, not once per tick.
const gatewayLifecycleProbeTimeout = 3 * time.Second

// GatewayLifecycle is the subset of GET /health/detailed the scheduler reads.
// The gateway publishes gateway_state = "running" | "draining" (and transitional
// states such as "starting"/"degraded") for restart/drain and scale-to-zero.
type GatewayLifecycle struct {
    GatewayState string `json:"gateway_state"`
    Status       string `json:"status"`
}

// Lifecycle probes GET /health/detailed and returns the announced gateway
// lifecycle state (SCHED-GAP-142). A transport error, non-200 or unparseable
// body returns an error; callers on the admission path treat any error as
// "not draining" (fail open) so the bounded POST retry remains the backstop.
// 401/403 are wrapped as ErrGatewayKeyRejected so terminal-auth classification
// (GAP-035) is never silently converted into a defer or a retry flood.
func (g *GatewayClient) Lifecycle(ctx context.Context, key string) (string, error) {
    pctx, cancel := context.WithTimeout(ctx, gatewayLifecycleProbeTimeout)
    defer cancel()
    req, err := http.NewRequestWithContext(pctx, "GET", g.baseURL+"/health/detailed", nil)
    if err != nil {
        return "", err
    }
    g.setAuth(req, key)
    resp, err := g.httpClient.Do(req)
    if err != nil {
        return "", err
    }
    defer func() { _ = resp.Body.Close() }()
    if resp.StatusCode == http.StatusUnauthorized || resp.StatusCode == http.StatusForbidden {
        body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
        return "", fmt.Errorf("%w (HTTP %d): %s", ErrGatewayKeyRejected, resp.StatusCode, authErrorDetail(body))
    }
    if resp.StatusCode != http.StatusOK {
        return "", fmt.Errorf("gateway health/detailed: HTTP %d", resp.StatusCode)
    }
    var lc GatewayLifecycle
    if err := json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(&lc); err != nil {
        return "", fmt.Errorf("%w: %v", ErrGatewayTransient, err)
    }
    return lc.GatewayState, nil
}

3.2 internal/scheduler/spawn.go — cached predicate

Add the cache fields to type Spawner struct:

    // SCHED-GAP-142: gateway lifecycle (drain) admission cache. GatewayDraining
    // is called from every spawn goroutine; the probe verdict is shared and
    // cached for gatewayDrainProbeTTL so a fleet-wide drain storm does not turn
    // into one GET /health/detailed per eligible project.
    drainMu      sync.Mutex
    drainChecked time.Time
    drainActive  bool

and the predicate (next to GatewayAvailable):

// gatewayDrainProbeTTL bounds the GatewayDraining probe cache. The gateway's
// announced lifecycle is stable on a seconds scale (the restart-deferral
// countdown logs every 30s; a forced drain runs to the ~180s cap), so a short
// TTL removes the per-spawn probe cost without widening the probe→POST race
// that the GAP-080 bounded retry backstops.
const gatewayDrainProbeTTL = 5 * time.Second

// GatewayDraining reports whether the gateway has ANNOUNCED that it is
// draining (restart/drain or scale-to-zero): GET /health/detailed returns
// gateway_state="draining". It is the admission-side counterpart to the
// GAP-080 POST retry (SCHED-GAP-142): the gateway announces the window well
// before it refuses turns, so deferring admission turns a guaranteed drop into
// a deferred tick instead of spending ~3.5s of retry against a ~180s outage.
//
// Fail-open by contract: an unreachable gateway or any probe error (including
// a terminal 401/403) reads as "not draining", so dispatch proceeds exactly as
// before and the POST's own classification + bounded retry remain
// authoritative. This is what keeps invariant (a) — auth/terminal classes stay
// terminal, never converted into defers or retries.
func (s *Spawner) GatewayDraining(ctx context.Context) bool {
    if s.gateway == nil {
        return false
    }
    s.drainMu.Lock()
    if !s.drainChecked.IsZero() && time.Since(s.drainChecked) < gatewayDrainProbeTTL {
        v := s.drainActive
        s.drainMu.Unlock()
        return v
    }
    s.drainMu.Unlock()

    state, err := s.gateway.Lifecycle(ctx, "")
    draining := err == nil && strings.EqualFold(strings.TrimSpace(state), "draining")

    s.drainMu.Lock()
    s.drainChecked = time.Now()
    s.drainActive = draining
    s.drainMu.Unlock()
    return draining
}

3.3 internal/scheduler/slot_pool.go — admission defer (after the load gate)

    // SCHED-GAP-142: gateway lifecycle admission gate (the counterpart to the
    // SCHED-GAP-125 load gate above). The gateway announces a restart/drain
    // BEFORE it starts refusing turns — the restart-deferral countdown is
    // public for minutes, and gateway_state="draining" is readable via
    // GET /health/detailed for the whole drain. A POST into that window is a
    // guaranteed 503 that the ~3.5s GAP-080 retry budget cannot outlast the
    // ~180s drain cap, so the tick is dropped and consecutive_failures
    // increments. Deferring admission keeps the project eligible — it stays
    // selected and is re-probed on the next evaluation once the gateway
    // reports running again. No cooldown is consumed and no failure is
    // recorded. Retry stays the fallback for a drain that starts after this
    // probe.
    if p.spawner != nil && p.spawner.GatewayDraining(context.Background()) {
        log.Printf("GATEWAY-DRAIN: deferring %s (tick %s) — gateway reports draining (work stays queued)",
            proj.Name, tickID)
        p.mu.Lock()
        events := p.events
        p.mu.Unlock()
        if events != nil {
            events.Emit(context.Background(), SeverityInfo, "slot_pool",
                "gateway drain deferred "+proj.Name,
                map[string]any{
                    "project": proj.Name,
                    "tick_id": tickID,
                })
        }
        return
    }

Ordering note: the block must be placed after the LoadGateShouldDefer block and before tryReserve, exactly like the load gate, so a deferred tick consumes neither a reservation nor a slot and never reaches lifecycle.Enqueue.

3.4 Why not just raise gatewayRetryMaxAttempts?

Sizing retry to the 180 s cap is the fallback if a probe is impossible. It is strictly worse here:

The admission hold is therefore primary; retry remains the last-resort backstop for the narrow probe→POST race, which the 5 s cache plus 30 s eval interval already keeps tiny. (If an operator ever wants a pure-retry posture, the correct knob is a drain-specific retry budget ≥ 180 s, not the shared 3-attempt budget — but that is out of scope here.)

4. Verification

All commands run against a clean checkout with the patch applied.

4.1 Build, vet, tests

$ gofmt -l internal/scheduler/gateway_client.go internal/scheduler/spawn.go \
           internal/scheduler/slot_pool.go internal/scheduler/gateway_drain_gate_test.go
(no output — formatted)
$ go build ./...              # ok
$ go vet ./internal/scheduler/  # clean
$ go test ./...
ok  github.com/coding-hermes/scheduler/cmd/migrate        0.012s
ok  github.com/coding-hermes/scheduler/cmd/schedulerd     2.536s
ok  github.com/coding-hermes/scheduler/internal/api       2.615s
ok  github.com/coding-hermes/scheduler/internal/blocks    0.008s
ok  github.com/coding-hermes/scheduler/internal/config    0.290s
ok  github.com/coding-hermes/scheduler/internal/dashboard 0.841s
ok  github.com/coding-hermes/scheduler/internal/database  2.455s
ok  github.com/coding-hermes/scheduler/internal/mcp       0.737s
ok  github.com/coding-hermes/scheduler/internal/scheduler 71.276s
ok  github.com/coding-hermes/scheduler/internal/sync      1.691s
ok  github.com/coding-hermes/scheduler/internal/version   0.003s

4.2 New regression tests (internal/scheduler/gateway_drain_gate_test.go)

$ go test ./internal/scheduler/ -run 'Test(GatewayClientLifecycleParsesDrainState|SpawnerGatewayDrainingFailOpen|SpawnerGatewayDrainingCachesProbe|SlotPoolDefersWhenGatewayDraining|SlotPoolAdmitsWhenGatewayRunning)' -v
--- PASS: TestGatewayClientLifecycleParsesDrainState
--- PASS: TestSpawnerGatewayDrainingFailOpen          # unreachable / 401 / nil → not draining
--- PASS: TestSpawnerGatewayDrainingCachesProbe        # 5 calls = 1 probe
--- PASS: TestSlotPoolDefersWhenGatewayDraining        # no tick row, Running()==0
--- PASS: TestSlotPoolAdmitsWhenGatewayRunning         # control: healthy gateway still dispatches

The defer test asserts the two properties that define the fix: while gateway_state="draining" the spawn creates zero ticks rows (a CreateProject-seeded project would otherwise enqueue) and reserves zero slots. The control test proves the gate is drain-specific, not a blanket block.

4.3 Live gateway probe (real host, real gateway 0.21.1)

$ KEY=$(grep API_SERVER_KEY /etc/coding-hermes/gateway.env | cut -d= -f2)
$ curl -s -H "Authorization: Bearer $KEY" http://<ip-address>:8642/health/detailed
{ "status": "ok", "gateway_state": "running",
  "readiness": { "status": "ok", "checks": { "gateway": { "status": "ok", "state": "running", ... } } }, ... }

A throwaway Go probe using NewGatewayClient(...).Lifecycle(ctx, "") against that URL returned live gateway_state="running" draining=false and PASS — i.e. the new code parses the live field and returns “not draining” on a healthy gateway (dispatch unchanged). During a restart the same field is written "draining" by gateway/run_shutdown.py (_drain_active_agents → _update_runtime_status("draining")), which the probe maps to a defer. The parse itself is pinned by TestGatewayClientLifecycleParsesDrainState in section 4.2, so no throwaway file is part of the patch.

4.4 Existing drain behavior unchanged when not draining

TestStop_DrainsInFlightGatewayTick / TestStop_DrainTimeoutMarksRunningTicksFailed (the SCHED-GAP-077 shutdown-drain regressions) still pass. Their blocking stub was updated to answer /health/detailed non-blocking with gateway_state="running", modelling a healthy gateway, so only the real POST blocks (as before).

5. Outcome

before after
spawn during announced drain 4 POSTs (~3.5 s) → failed + consecutive_failures++ deferred, no tick row, no cooldown, no failure
probe cost — 1 GET /health/detailed per 5 s fleet-wide
auth/terminal classes terminal terminal (fail-open preserves the POST classifier)
retry after a drain starts post-probe 4 POSTs unchanged (backstop)

Net effect on the measured class: the ~828 drain rows / 87 lanes stop being recorded failures; the same projects are re-evaluated after the gateway returns to running and their ticks run normally. No migration, flag, or gateway change is required.

Evidence & signatures

# Evidence
- Problem class: gateway-drain-503-tick-failure
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-18T03:32:54.515Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: scheduler tick rows end status=failed/outcome=failed with error\n\"gateway unreachable and exec fallback disabled: gateway POST: HTTP 503: invalid_request_error: Gateway is draining existing work; retry shortly.\"\n\nObserved for the my-project lane twice in one day:\n- my-project-2026-09-17-01-06-43: spawned 2026-09-16T20:06:43-05:00, completed 20:06:47\n- my-project-2026-09-17-09-06-18: spawned 2026-09-17T04:06:18-05:00, completed 04:06:22\nBoth rows carry gateway_trace classification=transport-error, attempts=4, elapsed_ms 3505 / 3530, deadline_ms 1800000 (the wall deadline was never the binding constraint), pid=0, session_id empty.\n\nBlast radius measured from the scheduler store: 40 tick rows whose error contains \"draining existing work\" in the API window, spread across ~2 dozen lanes (off-by-one-sync, <project>-qa, coding-hermes-tools, bunker-dogfood/-qa/-sync, duckbrain-sync, <project>-qa/-sync, <project>, <project>-dogfood/-qa, hermes-canopy/-sync/-dogfood, terminal-jail/-sync, heading, trouble, task-router, coding-hermes-scheduler, my-project). This is a fleet-wide drain storm, not a per-project defect.\n\nRoot cause (measured, not inferred): the spawn retry budget is ~51x smaller than the outage it retries into.\n- Scheduler side: internal/scheduler/spawn.go:558 defines gatewayRetryMaxAttempts = 3 retries after the initial POST (4 POSTs total) with gatewayRetryBackoff 500ms -> 1s -> 2s, capped at 4s; gatewayRetrySleep (spawn.go:575-583) takes max(backoff, server Retry-After). Worst-case added latency is ~3.5s, which is exactly the observed elapsed_ms (3505/3530).\n- Gateway side: every drain in the same logs runs to its cap - \"Shutdown phase: drain done at +183.44s (drain took 180.11s, timed_out=True, active_at_start=0, active_now=0, cron_at_start=0, cron_now=0, api_at_start=1, api_now=1)\" at 2026-09-16 20:07:23 local, and \"+183.62s (drain took 180.10s, timed_out=True, ... api_at_start=6, api_now=1)\" at 2026-09-17 04:08:05 local. Both my-project drops landed inside those windows (20:06:43 and 04:06:18).\nSo any spawn that lands in a drain window spends its full ~3.5s retry budget, then the tick is dropped with \"SKIPPED: <project> tick=<id> exec fallback disabled, dropping tick\" (spawn.go:1578) and a recorded failure (spawn.go:1590). The Retry-After floor added by SCHED-GAP-136 (gateway_client.go:288, spawn.go:556-566 comment asserts the drain 503 sends \"Retry-After: 1\") cannot close a 180s gap with a 1s hint.\n\nUseful extra signal for a fix: the gateway ANNOUNCES the window before it drains - \"Restart deferred: waiting on 3 active work unit(s) (0 wedged and excluded; 688s remaining before force drain)\" repeated on a 30s cadence from 2026-09-16 19:52:51 to 20:03:52 local, i.e. ~11 minutes of public countdown before the forced drain. An admission hold that reads that state is available well before the first 503.\n\nFix shape that matches the evidence: gate admission on the upstream's announced lifecycle state instead of retrying harder. A cheap probe (gateway health / lifecycle state) that defers eligible spawns while the gateway reports draining turns a guaranteed drop into a deferred tick; retry stays as the fallback for a drain that starts after the probe. If retry is kept as the only lever, the budget must be sized to the drain cap (>= 180s), which is far above the current 4 POST / ~3.5s envelope.\n\nInvariants a fix must preserve: (a) auth/terminal classes stay terminal (no retry flood - the reason exec fallback is disabled), (b) the exec fallback stays disabled, (c) retries happen BEFORE the skip decision so an exhausted retry is still recorded failed-with-error, never completed.\n\nDisposition on the owning project's board: this class is already tracked in coding-hermes-scheduler as SCHED-GAP-133/134/135/136 (complete: tasks-mode hot-loop, harness-failure auto-disable, stuck running rows, Retry-After backoff) with SCHED-GAP-137 (pending) covering the drain-storm residues including consecutive_failures residue. What this entry adds is the measured retry-budget-vs-drain-window arithmetic (4 POSTs / ~3.5s vs 180.0s timed-out drain) and the pre-announcement evidence that makes an admission hold implementable.\n\nHonesty note: the drain durations, the retry exhaustion (attempts=4, elapsed ~3.5s), the 40-row blast radius and the restart-deferral countdown are all measured live. The claim that the drain 503 carries \"Retry-After: 1\" is taken from the scheduler code comment, not re-measured here; and no fix is claimed - the fix lives on the scheduler board.", "environment": "Linux fleet host; coding-hermes schedulerd (Go, SQLite) posting ticks to the Hermes gateway (:8642) with --no-exec-fallback; Hermes gateway graceful drain (180s cap) during restarts", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gateway-drain-503-tick-failure", "provider": "openrouter", "solved_at": "2026-09-18T03:32:54.516Z", "version": "scheduler v31 (GAP-125 load-gate era)"}
Generated from the verified corpus · MIT licensedBack to the catalog