◐ Off-By-One · answer catalog

scheduler-cooldown-pin-regression

1 answer(s)godocker

r = apiput(APIURL, {"cooldowns": 7200, "source": "operator", "pinned": True})

📦 Source in repository (JSON)

Answer

Root cause. A stand-in wake handler PUT cooldown_s=900 into the API during the tick storm (transient throttle). fleet-cooldown-policy.py then regenerated the durable fleet.toml pin from API state, so the transient 900 was written as if it were an operator pin. The scheduler loads fleet.toml at startup only (no hot-reload), so 900 survived daemon restarts and the operator pin (7200) was never reverted.

Fix — two actions, in this order:

  1. PUT 7200 via the API first — if fleet-cooldown-policy.py runs between our steps, it copies the corrected value, not the stale 900.
  2. Edit the fleet.toml pin to 7200 — the file is what survives daemon restarts.
  3. Verify file and API agree post-fix (both 7200).

fix_cooldown_pin.py (core):

# Step 1: PUT corrected value via API (FIRST — interleave-safe)
r = api_put(API_URL, {"cooldown_s": 7200, "source": "operator", "pinned": True})

# Step 2: pin fleet.toml so it survives daemon restart (no hot-reload)
set_pin("sim/fleet.toml", 7200)   # atomic: tmp file + os.replace

# Step 3: verify file + API agree post-fix
assert api_get()["cooldown_s"] == 7200 == read_pin("sim/fleet.toml")

The file edit is atomic (temp + os.replace), handles a missing key, and the PUT-first ordering is what protects against an interleaved policy run.

Recurrence guard (the actual root fix): the buggy policy copied any API value blindly. Hardened fleet-cooldown-policy.py only pins values marked source == "operator" / pinned == true; stand-in wake values are refused:

if not state.get("pinned") or state.get("source") != "operator":
    print("REFUSED transient value; fleet.toml untouched")
    sys.exit(0)

So a future storm's transient 900 can never masquerade as an operator pin again.

Evidence & signatures

The environment contained no fleet code, so I built a faithful simulation (`~/scheduler-cooldown-fix/`): a persisted HTTP API (`sim/api_server.py`), the regressed `fleet.toml`, the buggy + hardened policy scripts, the fix script, and a 19-test harness. All 19 pass (`python3 test_cooldown_fix.py`).

| Scenario | Result |
|---|---|
| Regression reproduced (stand-in PUT 900 → buggy policy pins 900, 7200 lost) | PASS |
| Fix exits 0; API = 7200, file pin = 7200, file+API agree | PASS |
| **Daemon restart** (server killed/restarted) — API and pin both still 7200 | PASS |
| New storm: transient 900 in API only; hardened policy refuses → pin stays 7200 | PASS |
| Idempotency — fix re-run twice (2nd and 3rd runs are no-op passes) | PASS |
| Missing `cooldown_s` key in fleet.toml — fix adds it = 7200 | PASS |
| **API down during fix** — file still pinned to 7200, script reports partial (`ERROR` in API result, exit 1) | PASS |
| Atomicity — no `.tmp` litter; fleet.toml parses as valid TOML | PASS |
| Interleave proof — policy running between PUT and file-edit copies the corrected 7200 | PASS |

Final verified state — file: `[scheduler]\ncooldown_s = 7200  # operator pin`; API: `{"cooldown_s": 7200, "source": "operator", "pinned": true}`. They agree.
{"model": "deepseek-v4-flash", "problem_class": "scheduler-cooldown-pin-regression", "result": "passed", "tests": 19}
Generated from the verified corpus · MIT licensedBack to the catalog