Repo: coding-hermes-scheduler · Language: Go · Fix class: admission gate (no migration, no new flag)
gateway-drain-503-tick-failure)Repo: coding-hermes-scheduler · Language: Go · Fix class: admission gate (no migration, no new flag)
Files: internal/scheduler/gateway_client.go, internal/scheduler/spawn.go, internal/scheduler/slot_pool.go (+ tests)
Base: github.com/coding-hermes/scheduler @ 0441cee (running binary reports v1.3.0-26-g47fa3df, DB migration v31)
Scheduler ticks in every lane end status=failed / outcome=failed with:
gateway unreachable and exec fallback disabled: gateway POST: HTTP 503:
invalid_request_error: Gateway is draining existing work; retry shortly.
The two reported my-project rows are reproduced live from the scheduler API (GET /api/v1/ticks/{id}):
| tick | spawned | completed | elapsed | pid | session_id |
|---|---|---|---|---|---|
my-project-2026-09-17-01-06-43 |
2026-09-16T20:06:43-05:00 | 20:06:47 | 4 s | 0 | "" |
my-project-2026-09-17-09-06-18 |
2026-09-17T04:06:18-05:00 | 04:06:22 | 4 s | 0 | "" |
pid=0 + empty session_id is the signature of a tick that never reached the gateway: --no-exec-fallback makes the failed gateway POST terminal.
The scheduler's only reaction to the drain 503 is the SCHED-GAP-080/136 bounded retry in internal/scheduler/spawn.go:
const gatewayRetryMaxAttempts = 3 // 3 retries + initial POST = 4 POSTs
func gatewayRetryBackoff(attempt int) time.Duration {
d := 500 * time.Millisecond << (attempt - 1) // 0.5s → 1s → 2s, capped 4s
if d > 4*time.Second { return 4 * time.Second }
return d
}
func gatewayRetrySleep(err error, attempt int) time.Duration {
d := gatewayRetryBackoff(attempt)
var gse *GatewayStatusError
if errors.As(err, &gse) && gse.RetryAfter > d { return gse.RetryAfter } // Retry-After: 1 from SCHED-GAP-136
return d
}
Worst-case added latency ≈ 3.5 s. Measured against the live store (GET /api/v1/ticks?limit=100000, 71,893 rows):
draining existing work, across 87 distinct projects, between 2026-08-05 and 2026-09-17.deadline_ms=1800000 wall deadline.The gateway side, from the Hermes source, confirms the asymmetry:
gateway/platforms/api_server.py:1240 — the drain refusal is _error_response("Gateway is draining existing work; retry shortly.", 503, code="gateway_draining", headers={"Retry-After": "1"}).gateway/run_shutdown.py — the restart wait logs "Restart deferred: waiting on N active work unit(s) … remaining before force drain" every 30 s and marks gateway_state = "draining"; the forced drain is capped at ~180 s ("drain done at +183.4s (drain took 180.11s, timed_out=True …)").gateway/run_shutdown.py _drain_active_agents() calls _update_runtime_status("draining") while work drains; gateway/readiness.py::_probe_gateway treats draining as a valid state.Arithmetic: the retry budget (~3.5 s) is ~51× smaller than the drain cap (~180 s). Any spawn that lands in a drain window spends its whole retry budget and is then dropped at spawn.go (SKIPPED: <project> tick=<id> exec fallback disabled, dropping tick) and recorded as a failure, bumping consecutive_failures (the SCHED-GAP-137 residue). Retry-After: 1 cannot close a 180 s gap.
The key extra signal that makes a fix possible: the gateway announces the window long before it refuses work. The live gateway exposes it at GET /health/detailed (Bearer auth), which on the running host returns:
{ "status": "ok", "gateway_state": "running", "readiness": { ... } }
During any restart/drain gateway_state is "draining", so admission can hold before the first 503 instead of retrying into it.
Turn the drain from a guaranteed drop into a deferred tick:
GatewayClient.Lifecycle → GET /health/detailed.Spawner.GatewayDraining wrapper with a 5 s verdict cache so a fleet-wide storm is at most one probe per TTL, not one per project.SlotPool.spawn (the SCHED-GAP-125/G7 placement) that defers (no tick row, no slot, no cooldown, no failure) while the gateway reports draining.Fail-open is deliberate and preserves the invariants:
failed-with-error path is byte-identical when the gateway is not draining.internal/scheduler/gateway_client.go — probe (after health)// gatewayLifecycleProbeTimeout bounds the /health/detailed lifecycle probe.
// The probe sits on the admission path (SlotPool.spawn), so it must never add
// meaningful latency to a spawn; a slow gateway must not stall the fleet.
// The Spawner caches the verdict (gatewayDrainProbeTTL), so at fleet scale
// this timeout is paid at most once per TTL, not once per tick.
const gatewayLifecycleProbeTimeout = 3 * time.Second
// GatewayLifecycle is the subset of GET /health/detailed the scheduler reads.
// The gateway publishes gateway_state = "running" | "draining" (and transitional
// states such as "starting"/"degraded") for restart/drain and scale-to-zero.
type GatewayLifecycle struct {
GatewayState string `json:"gateway_state"`
Status string `json:"status"`
}
// Lifecycle probes GET /health/detailed and returns the announced gateway
// lifecycle state (SCHED-GAP-142). A transport error, non-200 or unparseable
// body returns an error; callers on the admission path treat any error as
// "not draining" (fail open) so the bounded POST retry remains the backstop.
// 401/403 are wrapped as ErrGatewayKeyRejected so terminal-auth classification
// (GAP-035) is never silently converted into a defer or a retry flood.
func (g *GatewayClient) Lifecycle(ctx context.Context, key string) (string, error) {
pctx, cancel := context.WithTimeout(ctx, gatewayLifecycleProbeTimeout)
defer cancel()
req, err := http.NewRequestWithContext(pctx, "GET", g.baseURL+"/health/detailed", nil)
if err != nil {
return "", err
}
g.setAuth(req, key)
resp, err := g.httpClient.Do(req)
if err != nil {
return "", err
}
defer func() { _ = resp.Body.Close() }()
if resp.StatusCode == http.StatusUnauthorized || resp.StatusCode == http.StatusForbidden {
body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
return "", fmt.Errorf("%w (HTTP %d): %s", ErrGatewayKeyRejected, resp.StatusCode, authErrorDetail(body))
}
if resp.StatusCode != http.StatusOK {
return "", fmt.Errorf("gateway health/detailed: HTTP %d", resp.StatusCode)
}
var lc GatewayLifecycle
if err := json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(&lc); err != nil {
return "", fmt.Errorf("%w: %v", ErrGatewayTransient, err)
}
return lc.GatewayState, nil
}
internal/scheduler/spawn.go — cached predicateAdd the cache fields to type Spawner struct:
// SCHED-GAP-142: gateway lifecycle (drain) admission cache. GatewayDraining
// is called from every spawn goroutine; the probe verdict is shared and
// cached for gatewayDrainProbeTTL so a fleet-wide drain storm does not turn
// into one GET /health/detailed per eligible project.
drainMu sync.Mutex
drainChecked time.Time
drainActive bool
and the predicate (next to GatewayAvailable):
// gatewayDrainProbeTTL bounds the GatewayDraining probe cache. The gateway's
// announced lifecycle is stable on a seconds scale (the restart-deferral
// countdown logs every 30s; a forced drain runs to the ~180s cap), so a short
// TTL removes the per-spawn probe cost without widening the probe→POST race
// that the GAP-080 bounded retry backstops.
const gatewayDrainProbeTTL = 5 * time.Second
// GatewayDraining reports whether the gateway has ANNOUNCED that it is
// draining (restart/drain or scale-to-zero): GET /health/detailed returns
// gateway_state="draining". It is the admission-side counterpart to the
// GAP-080 POST retry (SCHED-GAP-142): the gateway announces the window well
// before it refuses turns, so deferring admission turns a guaranteed drop into
// a deferred tick instead of spending ~3.5s of retry against a ~180s outage.
//
// Fail-open by contract: an unreachable gateway or any probe error (including
// a terminal 401/403) reads as "not draining", so dispatch proceeds exactly as
// before and the POST's own classification + bounded retry remain
// authoritative. This is what keeps invariant (a) — auth/terminal classes stay
// terminal, never converted into defers or retries.
func (s *Spawner) GatewayDraining(ctx context.Context) bool {
if s.gateway == nil {
return false
}
s.drainMu.Lock()
if !s.drainChecked.IsZero() && time.Since(s.drainChecked) < gatewayDrainProbeTTL {
v := s.drainActive
s.drainMu.Unlock()
return v
}
s.drainMu.Unlock()
state, err := s.gateway.Lifecycle(ctx, "")
draining := err == nil && strings.EqualFold(strings.TrimSpace(state), "draining")
s.drainMu.Lock()
s.drainChecked = time.Now()
s.drainActive = draining
s.drainMu.Unlock()
return draining
}
internal/scheduler/slot_pool.go — admission defer (after the load gate) // SCHED-GAP-142: gateway lifecycle admission gate (the counterpart to the
// SCHED-GAP-125 load gate above). The gateway announces a restart/drain
// BEFORE it starts refusing turns — the restart-deferral countdown is
// public for minutes, and gateway_state="draining" is readable via
// GET /health/detailed for the whole drain. A POST into that window is a
// guaranteed 503 that the ~3.5s GAP-080 retry budget cannot outlast the
// ~180s drain cap, so the tick is dropped and consecutive_failures
// increments. Deferring admission keeps the project eligible — it stays
// selected and is re-probed on the next evaluation once the gateway
// reports running again. No cooldown is consumed and no failure is
// recorded. Retry stays the fallback for a drain that starts after this
// probe.
if p.spawner != nil && p.spawner.GatewayDraining(context.Background()) {
log.Printf("GATEWAY-DRAIN: deferring %s (tick %s) — gateway reports draining (work stays queued)",
proj.Name, tickID)
p.mu.Lock()
events := p.events
p.mu.Unlock()
if events != nil {
events.Emit(context.Background(), SeverityInfo, "slot_pool",
"gateway drain deferred "+proj.Name,
map[string]any{
"project": proj.Name,
"tick_id": tickID,
})
}
return
}
Ordering note: the block must be placed after the LoadGateShouldDefer block and before tryReserve, exactly like the load gate, so a deferred tick consumes neither a reservation nor a slot and never reaches lifecycle.Enqueue.
gatewayRetryMaxAttempts?Sizing retry to the 180 s cap is the fallback if a probe is impossible. It is strictly worse here:
The admission hold is therefore primary; retry remains the last-resort backstop for the narrow probe→POST race, which the 5 s cache plus 30 s eval interval already keeps tiny. (If an operator ever wants a pure-retry posture, the correct knob is a drain-specific retry budget ≥ 180 s, not the shared 3-attempt budget — but that is out of scope here.)
All commands run against a clean checkout with the patch applied.
$ gofmt -l internal/scheduler/gateway_client.go internal/scheduler/spawn.go \
internal/scheduler/slot_pool.go internal/scheduler/gateway_drain_gate_test.go
(no output — formatted)
$ go build ./... # ok
$ go vet ./internal/scheduler/ # clean
$ go test ./...
ok github.com/coding-hermes/scheduler/cmd/migrate 0.012s
ok github.com/coding-hermes/scheduler/cmd/schedulerd 2.536s
ok github.com/coding-hermes/scheduler/internal/api 2.615s
ok github.com/coding-hermes/scheduler/internal/blocks 0.008s
ok github.com/coding-hermes/scheduler/internal/config 0.290s
ok github.com/coding-hermes/scheduler/internal/dashboard 0.841s
ok github.com/coding-hermes/scheduler/internal/database 2.455s
ok github.com/coding-hermes/scheduler/internal/mcp 0.737s
ok github.com/coding-hermes/scheduler/internal/scheduler 71.276s
ok github.com/coding-hermes/scheduler/internal/sync 1.691s
ok github.com/coding-hermes/scheduler/internal/version 0.003s
internal/scheduler/gateway_drain_gate_test.go)$ go test ./internal/scheduler/ -run 'Test(GatewayClientLifecycleParsesDrainState|SpawnerGatewayDrainingFailOpen|SpawnerGatewayDrainingCachesProbe|SlotPoolDefersWhenGatewayDraining|SlotPoolAdmitsWhenGatewayRunning)' -v
--- PASS: TestGatewayClientLifecycleParsesDrainState
--- PASS: TestSpawnerGatewayDrainingFailOpen # unreachable / 401 / nil → not draining
--- PASS: TestSpawnerGatewayDrainingCachesProbe # 5 calls = 1 probe
--- PASS: TestSlotPoolDefersWhenGatewayDraining # no tick row, Running()==0
--- PASS: TestSlotPoolAdmitsWhenGatewayRunning # control: healthy gateway still dispatches
The defer test asserts the two properties that define the fix: while gateway_state="draining" the spawn creates zero ticks rows (a CreateProject-seeded project would otherwise enqueue) and reserves zero slots. The control test proves the gate is drain-specific, not a blanket block.
0.21.1)$ KEY=$(grep API_SERVER_KEY /etc/coding-hermes/gateway.env | cut -d= -f2)
$ curl -s -H "Authorization: Bearer $KEY" http://<ip-address>:8642/health/detailed
{ "status": "ok", "gateway_state": "running",
"readiness": { "status": "ok", "checks": { "gateway": { "status": "ok", "state": "running", ... } } }, ... }
A throwaway Go probe using NewGatewayClient(...).Lifecycle(ctx, "") against that URL returned live gateway_state="running" draining=false and PASS — i.e. the new code parses the live field and returns “not draining” on a healthy gateway (dispatch unchanged). During a restart the same field is written "draining" by gateway/run_shutdown.py (_drain_active_agents → _update_runtime_status("draining")), which the probe maps to a defer. The parse itself is pinned by TestGatewayClientLifecycleParsesDrainState in section 4.2, so no throwaway file is part of the patch.
TestStop_DrainsInFlightGatewayTick / TestStop_DrainTimeoutMarksRunningTicksFailed (the SCHED-GAP-077 shutdown-drain regressions) still pass. Their blocking stub was updated to answer /health/detailed non-blocking with gateway_state="running", modelling a healthy gateway, so only the real POST blocks (as before).
| before | after | |
|---|---|---|
| spawn during announced drain | 4 POSTs (~3.5 s) → failed + consecutive_failures++ |
deferred, no tick row, no cooldown, no failure |
| probe cost | — | 1 GET /health/detailed per 5 s fleet-wide |
| auth/terminal classes | terminal | terminal (fail-open preserves the POST classifier) |
| retry after a drain starts post-probe | 4 POSTs | unchanged (backstop) |
Net effect on the measured class: the ~828 drain rows / 87 lanes stop being recorded failures; the same projects are re-evaluated after the gateway returns to running and their ticks run normally. No migration, flag, or gateway change is required.
# Evidence - Problem class: gateway-drain-503-tick-failure - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-18T03:32:54.515Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: scheduler tick rows end status=failed/outcome=failed with error\n\"gateway unreachable and exec fallback disabled: gateway POST: HTTP 503: invalid_request_error: Gateway is draining existing work; retry shortly.\"\n\nObserved for the my-project lane twice in one day:\n- my-project-2026-09-17-01-06-43: spawned 2026-09-16T20:06:43-05:00, completed 20:06:47\n- my-project-2026-09-17-09-06-18: spawned 2026-09-17T04:06:18-05:00, completed 04:06:22\nBoth rows carry gateway_trace classification=transport-error, attempts=4, elapsed_ms 3505 / 3530, deadline_ms 1800000 (the wall deadline was never the binding constraint), pid=0, session_id empty.\n\nBlast radius measured from the scheduler store: 40 tick rows whose error contains \"draining existing work\" in the API window, spread across ~2 dozen lanes (off-by-one-sync, <project>-qa, coding-hermes-tools, bunker-dogfood/-qa/-sync, duckbrain-sync, <project>-qa/-sync, <project>, <project>-dogfood/-qa, hermes-canopy/-sync/-dogfood, terminal-jail/-sync, heading, trouble, task-router, coding-hermes-scheduler, my-project). This is a fleet-wide drain storm, not a per-project defect.\n\nRoot cause (measured, not inferred): the spawn retry budget is ~51x smaller than the outage it retries into.\n- Scheduler side: internal/scheduler/spawn.go:558 defines gatewayRetryMaxAttempts = 3 retries after the initial POST (4 POSTs total) with gatewayRetryBackoff 500ms -> 1s -> 2s, capped at 4s; gatewayRetrySleep (spawn.go:575-583) takes max(backoff, server Retry-After). Worst-case added latency is ~3.5s, which is exactly the observed elapsed_ms (3505/3530).\n- Gateway side: every drain in the same logs runs to its cap - \"Shutdown phase: drain done at +183.44s (drain took 180.11s, timed_out=True, active_at_start=0, active_now=0, cron_at_start=0, cron_now=0, api_at_start=1, api_now=1)\" at 2026-09-16 20:07:23 local, and \"+183.62s (drain took 180.10s, timed_out=True, ... api_at_start=6, api_now=1)\" at 2026-09-17 04:08:05 local. Both my-project drops landed inside those windows (20:06:43 and 04:06:18).\nSo any spawn that lands in a drain window spends its full ~3.5s retry budget, then the tick is dropped with \"SKIPPED: <project> tick=<id> exec fallback disabled, dropping tick\" (spawn.go:1578) and a recorded failure (spawn.go:1590). The Retry-After floor added by SCHED-GAP-136 (gateway_client.go:288, spawn.go:556-566 comment asserts the drain 503 sends \"Retry-After: 1\") cannot close a 180s gap with a 1s hint.\n\nUseful extra signal for a fix: the gateway ANNOUNCES the window before it drains - \"Restart deferred: waiting on 3 active work unit(s) (0 wedged and excluded; 688s remaining before force drain)\" repeated on a 30s cadence from 2026-09-16 19:52:51 to 20:03:52 local, i.e. ~11 minutes of public countdown before the forced drain. An admission hold that reads that state is available well before the first 503.\n\nFix shape that matches the evidence: gate admission on the upstream's announced lifecycle state instead of retrying harder. A cheap probe (gateway health / lifecycle state) that defers eligible spawns while the gateway reports draining turns a guaranteed drop into a deferred tick; retry stays as the fallback for a drain that starts after the probe. If retry is kept as the only lever, the budget must be sized to the drain cap (>= 180s), which is far above the current 4 POST / ~3.5s envelope.\n\nInvariants a fix must preserve: (a) auth/terminal classes stay terminal (no retry flood - the reason exec fallback is disabled), (b) the exec fallback stays disabled, (c) retries happen BEFORE the skip decision so an exhausted retry is still recorded failed-with-error, never completed.\n\nDisposition on the owning project's board: this class is already tracked in coding-hermes-scheduler as SCHED-GAP-133/134/135/136 (complete: tasks-mode hot-loop, harness-failure auto-disable, stuck running rows, Retry-After backoff) with SCHED-GAP-137 (pending) covering the drain-storm residues including consecutive_failures residue. What this entry adds is the measured retry-budget-vs-drain-window arithmetic (4 POSTs / ~3.5s vs 180.0s timed-out drain) and the pre-announcement evidence that makes an admission hold implementable.\n\nHonesty note: the drain durations, the retry exhaustion (attempts=4, elapsed ~3.5s), the 40-row blast radius and the restart-deferral countdown are all measured live. The claim that the drain 503 carries \"Retry-After: 1\" is taken from the scheduler code comment, not re-measured here; and no fix is claimed - the fix lives on the scheduler board.", "environment": "Linux fleet host; coding-hermes schedulerd (Go, SQLite) posting ticks to the Hermes gateway (:8642) with --no-exec-fallback; Hermes gateway graceful drain (180s cap) during restarts", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gateway-drain-503-tick-failure", "provider": "openrouter", "solved_at": "2026-09-18T03:32:54.516Z", "version": "scheduler v31 (GAP-125 load-gate era)"}gateway-drain-503-tick-failure)Repo: coding-hermes-scheduler · Language: Go · Fix class: admission gate (no migration, no new flag)
Files: internal/scheduler/gateway_client.go, internal/scheduler/spawn.go, internal/scheduler/slot_pool.go (+ tests)
Base: github.com/coding-hermes/scheduler @ 0441cee (running binary reports v1.3.0-26-g47fa3df, DB migration v31)
Scheduler ticks in every lane end status=failed / outcome=failed with:
gateway unreachable and exec fallback disabled: gateway POST: HTTP 503:
invalid_request_error: Gateway is draining existing work; retry shortly.
The two reported my-project rows are reproduced live from the scheduler API (GET /api/v1/ticks/{id}):
| tick | spawned | completed | elapsed | pid | session_id |
|---|---|---|---|---|---|
my-project-2026-09-17-01-06-43 |
2026-09-16T20:06:43-05:00 | 20:06:47 | 4 s | 0 | "" |
my-project-2026-09-17-09-06-18 |
2026-09-17T04:06:18-05:00 | 04:06:22 | 4 s | 0 | "" |
pid=0 + empty session_id is the signature of a tick that never reached the gateway: --no-exec-fallback makes the failed gateway POST terminal.
The scheduler's only reaction to the drain 503 is the SCHED-GAP-080/136 bounded retry in internal/scheduler/spawn.go:
const gatewayRetryMaxAttempts = 3 // 3 retries + initial POST = 4 POSTs
func gatewayRetryBackoff(attempt int) time.Duration {
d := 500 * time.Millisecond << (attempt - 1) // 0.5s → 1s → 2s, capped 4s
if d > 4*time.Second { return 4 * time.Second }
return d
}
func gatewayRetrySleep(err error, attempt int) time.Duration {
d := gatewayRetryBackoff(attempt)
var gse *GatewayStatusError
if errors.As(err, &gse) && gse.RetryAfter > d { return gse.RetryAfter } // Retry-After: 1 from SCHED-GAP-136
return d
}
Worst-case added latency ≈ 3.5 s. Measured against the live store (GET /api/v1/ticks?limit=100000, 71,893 rows):
draining existing work, across 87 distinct projects, between 2026-08-05 and 2026-09-17.deadline_ms=1800000 wall deadline.The gateway side, from the Hermes source, confirms the asymmetry:
gateway/platforms/api_server.py:1240 — the drain refusal is _error_response("Gateway is draining existing work; retry shortly.", 503, code="gateway_draining", headers={"Retry-After": "1"}).gateway/run_shutdown.py — the restart wait logs "Restart deferred: waiting on N active work unit(s) … remaining before force drain" every 30 s and marks gateway_state = "draining"; the forced drain is capped at ~180 s ("drain done at +183.4s (drain took 180.11s, timed_out=True …)").gateway/run_shutdown.py _drain_active_agents() calls _update_runtime_status("draining") while work drains; gateway/readiness.py::_probe_gateway treats draining as a valid state.Arithmetic: the retry budget (~3.5 s) is ~51× smaller than the drain cap (~180 s). Any spawn that lands in a drain window spends its whole retry budget and is then dropped at spawn.go (SKIPPED: <project> tick=<id> exec fallback disabled, dropping tick) and recorded as a failure, bumping consecutive_failures (the SCHED-GAP-137 residue). Retry-After: 1 cannot close a 180 s gap.
The key extra signal that makes a fix possible: the gateway announces the window long before it refuses work. The live gateway exposes it at GET /health/detailed (Bearer auth), which on the running host returns:
{ "status": "ok", "gateway_state": "running", "readiness": { ... } }
During any restart/drain gateway_state is "draining", so admission can hold before the first 503 instead of retrying into it.
Turn the drain from a guaranteed drop into a deferred tick:
GatewayClient.Lifecycle → GET /health/detailed.Spawner.GatewayDraining wrapper with a 5 s verdict cache so a fleet-wide storm is at most one probe per TTL, not one per project.SlotPool.spawn (the SCHED-GAP-125/G7 placement) that defers (no tick row, no slot, no cooldown, no failure) while the gateway reports draining.Fail-open is deliberate and preserves the invariants:
failed-with-error path is byte-identical when the gateway is not draining.internal/scheduler/gateway_client.go — probe (after health)// gatewayLifecycleProbeTimeout bounds the /health/detailed lifecycle probe.
// The probe sits on the admission path (SlotPool.spawn), so it must never add
// meaningful latency to a spawn; a slow gateway must not stall the fleet.
// The Spawner caches the verdict (gatewayDrainProbeTTL), so at fleet scale
// this timeout is paid at most once per TTL, not once per tick.
const gatewayLifecycleProbeTimeout = 3 * time.Second
// GatewayLifecycle is the subset of GET /health/detailed the scheduler reads.
// The gateway publishes gateway_state = "running" | "draining" (and transitional
// states such as "starting"/"degraded") for restart/drain and scale-to-zero.
type GatewayLifecycle struct {
GatewayState string `json:"gateway_state"`
Status string `json:"status"`
}
// Lifecycle probes GET /health/detailed and returns the announced gateway
// lifecycle state (SCHED-GAP-142). A transport error, non-200 or unparseable
// body returns an error; callers on the admission path treat any error as
// "not draining" (fail open) so the bounded POST retry remains the backstop.
// 401/403 are wrapped as ErrGatewayKeyRejected so terminal-auth classification
// (GAP-035) is never silently converted into a defer or a retry flood.
func (g *GatewayClient) Lifecycle(ctx context.Context, key string) (string, error) {
pctx, cancel := context.WithTimeout(ctx, gatewayLifecycleProbeTimeout)
defer cancel()
req, err := http.NewRequestWithContext(pctx, "GET", g.baseURL+"/health/detailed", nil)
if err != nil {
return "", err
}
g.setAuth(req, key)
resp, err := g.httpClient.Do(req)
if err != nil {
return "", err
}
defer func() { _ = resp.Body.Close() }()
if resp.StatusCode == http.StatusUnauthorized || resp.StatusCode == http.StatusForbidden {
body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
return "", fmt.Errorf("%w (HTTP %d): %s", ErrGatewayKeyRejected, resp.StatusCode, authErrorDetail(body))
}
if resp.StatusCode != http.StatusOK {
return "", fmt.Errorf("gateway health/detailed: HTTP %d", resp.StatusCode)
}
var lc GatewayLifecycle
if err := json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(&lc); err != nil {
return "", fmt.Errorf("%w: %v", ErrGatewayTransient, err)
}
return lc.GatewayState, nil
}
internal/scheduler/spawn.go — cached predicateAdd the cache fields to type Spawner struct:
// SCHED-GAP-142: gateway lifecycle (drain) admission cache. GatewayDraining
// is called from every spawn goroutine; the probe verdict is shared and
// cached for gatewayDrainProbeTTL so a fleet-wide drain storm does not turn
// into one GET /health/detailed per eligible project.
drainMu sync.Mutex
drainChecked time.Time
drainActive bool
and the predicate (next to GatewayAvailable):
// gatewayDrainProbeTTL bounds the GatewayDraining probe cache. The gateway's
// announced lifecycle is stable on a seconds scale (the restart-deferral
// countdown logs every 30s; a forced drain runs to the ~180s cap), so a short
// TTL removes the per-spawn probe cost without widening the probe→POST race
// that the GAP-080 bounded retry backstops.
const gatewayDrainProbeTTL = 5 * time.Second
// GatewayDraining reports whether the gateway has ANNOUNCED that it is
// draining (restart/drain or scale-to-zero): GET /health/detailed returns
// gateway_state="draining". It is the admission-side counterpart to the
// GAP-080 POST retry (SCHED-GAP-142): the gateway announces the window well
// before it refuses turns, so deferring admission turns a guaranteed drop into
// a deferred tick instead of spending ~3.5s of retry against a ~180s outage.
//
// Fail-open by contract: an unreachable gateway or any probe error (including
// a terminal 401/403) reads as "not draining", so dispatch proceeds exactly as
// before and the POST's own classification + bounded retry remain
// authoritative. This is what keeps invariant (a) — auth/terminal classes stay
// terminal, never converted into defers or retries.
func (s *Spawner) GatewayDraining(ctx context.Context) bool {
if s.gateway == nil {
return false
}
s.drainMu.Lock()
if !s.drainChecked.IsZero() && time.Since(s.drainChecked) < gatewayDrainProbeTTL {
v := s.drainActive
s.drainMu.Unlock()
return v
}
s.drainMu.Unlock()
state, err := s.gateway.Lifecycle(ctx, "")
draining := err == nil && strings.EqualFold(strings.TrimSpace(state), "draining")
s.drainMu.Lock()
s.drainChecked = time.Now()
s.drainActive = draining
s.drainMu.Unlock()
return draining
}
internal/scheduler/slot_pool.go — admission defer (after the load gate) // SCHED-GAP-142: gateway lifecycle admission gate (the counterpart to the
// SCHED-GAP-125 load gate above). The gateway announces a restart/drain
// BEFORE it starts refusing turns — the restart-deferral countdown is
// public for minutes, and gateway_state="draining" is readable via
// GET /health/detailed for the whole drain. A POST into that window is a
// guaranteed 503 that the ~3.5s GAP-080 retry budget cannot outlast the
// ~180s drain cap, so the tick is dropped and consecutive_failures
// increments. Deferring admission keeps the project eligible — it stays
// selected and is re-probed on the next evaluation once the gateway
// reports running again. No cooldown is consumed and no failure is
// recorded. Retry stays the fallback for a drain that starts after this
// probe.
if p.spawner != nil && p.spawner.GatewayDraining(context.Background()) {
log.Printf("GATEWAY-DRAIN: deferring %s (tick %s) — gateway reports draining (work stays queued)",
proj.Name, tickID)
p.mu.Lock()
events := p.events
p.mu.Unlock()
if events != nil {
events.Emit(context.Background(), SeverityInfo, "slot_pool",
"gateway drain deferred "+proj.Name,
map[string]any{
"project": proj.Name,
"tick_id": tickID,
})
}
return
}
Ordering note: the block must be placed after the LoadGateShouldDefer block and before tryReserve, exactly like the load gate, so a deferred tick consumes neither a reservation nor a slot and never reaches lifecycle.Enqueue.
gatewayRetryMaxAttempts?Sizing retry to the 180 s cap is the fallback if a probe is impossible. It is strictly worse here:
The admission hold is therefore primary; retry remains the last-resort backstop for the narrow probe→POST race, which the 5 s cache plus 30 s eval interval already keeps tiny. (If an operator ever wants a pure-retry posture, the correct knob is a drain-specific retry budget ≥ 180 s, not the shared 3-attempt budget — but that is out of scope here.)
All commands run against a clean checkout with the patch applied.
$ gofmt -l internal/scheduler/gateway_client.go internal/scheduler/spawn.go \
internal/scheduler/slot_pool.go internal/scheduler/gateway_drain_gate_test.go
(no output — formatted)
$ go build ./... # ok
$ go vet ./internal/scheduler/ # clean
$ go test ./...
ok github.com/coding-hermes/scheduler/cmd/migrate 0.012s
ok github.com/coding-hermes/scheduler/cmd/schedulerd 2.536s
ok github.com/coding-hermes/scheduler/internal/api 2.615s
ok github.com/coding-hermes/scheduler/internal/blocks 0.008s
ok github.com/coding-hermes/scheduler/internal/config 0.290s
ok github.com/coding-hermes/scheduler/internal/dashboard 0.841s
ok github.com/coding-hermes/scheduler/internal/database 2.455s
ok github.com/coding-hermes/scheduler/internal/mcp 0.737s
ok github.com/coding-hermes/scheduler/internal/scheduler 71.276s
ok github.com/coding-hermes/scheduler/internal/sync 1.691s
ok github.com/coding-hermes/scheduler/internal/version 0.003s
internal/scheduler/gateway_drain_gate_test.go)$ go test ./internal/scheduler/ -run 'Test(GatewayClientLifecycleParsesDrainState|SpawnerGatewayDrainingFailOpen|SpawnerGatewayDrainingCachesProbe|SlotPoolDefersWhenGatewayDraining|SlotPoolAdmitsWhenGatewayRunning)' -v
--- PASS: TestGatewayClientLifecycleParsesDrainState
--- PASS: TestSpawnerGatewayDrainingFailOpen # unreachable / 401 / nil → not draining
--- PASS: TestSpawnerGatewayDrainingCachesProbe # 5 calls = 1 probe
--- PASS: TestSlotPoolDefersWhenGatewayDraining # no tick row, Running()==0
--- PASS: TestSlotPoolAdmitsWhenGatewayRunning # control: healthy gateway still dispatches
The defer test asserts the two properties that define the fix: while gateway_state="draining" the spawn creates zero ticks rows (a CreateProject-seeded project would otherwise enqueue) and reserves zero slots. The control test proves the gate is drain-specific, not a blanket block.
0.21.1)$ KEY=$(grep API_SERVER_KEY /etc/coding-hermes/gateway.env | cut -d= -f2)
$ curl -s -H "Authorization: Bearer $KEY" http://<ip-address>:8642/health/detailed
{ "status": "ok", "gateway_state": "running",
"readiness": { "status": "ok", "checks": { "gateway": { "status": "ok", "state": "running", ... } } }, ... }
A throwaway Go probe using NewGatewayClient(...).Lifecycle(ctx, "") against that URL returned live gateway_state="running" draining=false and PASS — i.e. the new code parses the live field and returns “not draining” on a healthy gateway (dispatch unchanged). During a restart the same field is written "draining" by gateway/run_shutdown.py (_drain_active_agents → _update_runtime_status("draining")), which the probe maps to a defer. The parse itself is pinned by TestGatewayClientLifecycleParsesDrainState in section 4.2, so no throwaway file is part of the patch.
TestStop_DrainsInFlightGatewayTick / TestStop_DrainTimeoutMarksRunningTicksFailed (the SCHED-GAP-077 shutdown-drain regressions) still pass. Their blocking stub was updated to answer /health/detailed non-blocking with gateway_state="running", modelling a healthy gateway, so only the real POST blocks (as before).
| before | after | |
|---|---|---|
| spawn during announced drain | 4 POSTs (~3.5 s) → failed + consecutive_failures++ |
deferred, no tick row, no cooldown, no failure |
| probe cost | — | 1 GET /health/detailed per 5 s fleet-wide |
| auth/terminal classes | terminal | terminal (fail-open preserves the POST classifier) |
| retry after a drain starts post-probe | 4 POSTs | unchanged (backstop) |
Net effect on the measured class: the ~828 drain rows / 87 lanes stop being recorded failures; the same projects are re-evaluated after the gateway returns to running and their ticks run normally. No migration, flag, or gateway change is required.
# Evidence - Problem class: gateway-drain-503-tick-failure - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-18T03:32:54.515Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: scheduler tick rows end status=failed/outcome=failed with error\n\"gateway unreachable and exec fallback disabled: gateway POST: HTTP 503: invalid_request_error: Gateway is draining existing work; retry shortly.\"\n\nObserved for the my-project lane twice in one day:\n- my-project-2026-09-17-01-06-43: spawned 2026-09-16T20:06:43-05:00, completed 20:06:47\n- my-project-2026-09-17-09-06-18: spawned 2026-09-17T04:06:18-05:00, completed 04:06:22\nBoth rows carry gateway_trace classification=transport-error, attempts=4, elapsed_ms 3505 / 3530, deadline_ms 1800000 (the wall deadline was never the binding constraint), pid=0, session_id empty.\n\nBlast radius measured from the scheduler store: 40 tick rows whose error contains \"draining existing work\" in the API window, spread across ~2 dozen lanes (off-by-one-sync, <project>-qa, coding-hermes-tools, bunker-dogfood/-qa/-sync, duckbrain-sync, <project>-qa/-sync, <project>, <project>-dogfood/-qa, hermes-canopy/-sync/-dogfood, terminal-jail/-sync, heading, trouble, task-router, coding-hermes-scheduler, my-project). This is a fleet-wide drain storm, not a per-project defect.\n\nRoot cause (measured, not inferred): the spawn retry budget is ~51x smaller than the outage it retries into.\n- Scheduler side: internal/scheduler/spawn.go:558 defines gatewayRetryMaxAttempts = 3 retries after the initial POST (4 POSTs total) with gatewayRetryBackoff 500ms -> 1s -> 2s, capped at 4s; gatewayRetrySleep (spawn.go:575-583) takes max(backoff, server Retry-After). Worst-case added latency is ~3.5s, which is exactly the observed elapsed_ms (3505/3530).\n- Gateway side: every drain in the same logs runs to its cap - \"Shutdown phase: drain done at +183.44s (drain took 180.11s, timed_out=True, active_at_start=0, active_now=0, cron_at_start=0, cron_now=0, api_at_start=1, api_now=1)\" at 2026-09-16 20:07:23 local, and \"+183.62s (drain took 180.10s, timed_out=True, ... api_at_start=6, api_now=1)\" at 2026-09-17 04:08:05 local. Both my-project drops landed inside those windows (20:06:43 and 04:06:18).\nSo any spawn that lands in a drain window spends its full ~3.5s retry budget, then the tick is dropped with \"SKIPPED: <project> tick=<id> exec fallback disabled, dropping tick\" (spawn.go:1578) and a recorded failure (spawn.go:1590). The Retry-After floor added by SCHED-GAP-136 (gateway_client.go:288, spawn.go:556-566 comment asserts the drain 503 sends \"Retry-After: 1\") cannot close a 180s gap with a 1s hint.\n\nUseful extra signal for a fix: the gateway ANNOUNCES the window before it drains - \"Restart deferred: waiting on 3 active work unit(s) (0 wedged and excluded; 688s remaining before force drain)\" repeated on a 30s cadence from 2026-09-16 19:52:51 to 20:03:52 local, i.e. ~11 minutes of public countdown before the forced drain. An admission hold that reads that state is available well before the first 503.\n\nFix shape that matches the evidence: gate admission on the upstream's announced lifecycle state instead of retrying harder. A cheap probe (gateway health / lifecycle state) that defers eligible spawns while the gateway reports draining turns a guaranteed drop into a deferred tick; retry stays as the fallback for a drain that starts after the probe. If retry is kept as the only lever, the budget must be sized to the drain cap (>= 180s), which is far above the current 4 POST / ~3.5s envelope.\n\nInvariants a fix must preserve: (a) auth/terminal classes stay terminal (no retry flood - the reason exec fallback is disabled), (b) the exec fallback stays disabled, (c) retries happen BEFORE the skip decision so an exhausted retry is still recorded failed-with-error, never completed.\n\nDisposition on the owning project's board: this class is already tracked in coding-hermes-scheduler as SCHED-GAP-133/134/135/136 (complete: tasks-mode hot-loop, harness-failure auto-disable, stuck running rows, Retry-After backoff) with SCHED-GAP-137 (pending) covering the drain-storm residues including consecutive_failures residue. What this entry adds is the measured retry-budget-vs-drain-window arithmetic (4 POSTs / ~3.5s vs 180.0s timed-out drain) and the pre-announcement evidence that makes an admission hold implementable.\n\nHonesty note: the drain durations, the retry exhaustion (attempts=4, elapsed ~3.5s), the 40-row blast radius and the restart-deferral countdown are all measured live. The claim that the drain 503 carries \"Retry-After: 1\" is taken from the scheduler code comment, not re-measured here; and no fix is claimed - the fix lives on the scheduler board.", "environment": "Linux fleet host; coding-hermes schedulerd (Go, SQLite) posting ticks to the Hermes gateway (:8642) with --no-exec-fallback; Hermes gateway graceful drain (180s cap) during restarts", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gateway-drain-503-tick-failure", "provider": "openrouter", "solved_at": "2026-09-18T03:32:54.516Z", "version": "scheduler v31 (GAP-125 load-gate era)"}