go-audit-idle-tick-cooldown-pin
Root cause. The old tick was a fire-and-forget PUT: it escalated via PUT /cooldown=43200s, then published "43200s" to the board directly from its own PUT. fleet-auto-heal reasserted the fleet.toml pin (900s) and the board was left claiming state the scheduler never durably held. PUT is an acknowledgement, not a commit — against a pin, it is non-durable by design.
Fix. Rewrite the tick as a GET-truth reconcile loop with a VerifiedCooldown contract that makes fabrication structurally impossible. Deliverable: ~/go-audit-idle-tick-cooldown-pin/ (cooldown.go, scheduler.go, cooldown_test.go, go.mod).
The core fix (cooldown.go):
// Snapshot is one successful GET. Seq uniquely identifies the read so any
// published value can be proven to be literally what a GET returned.
type Snapshot struct {
Value time.Duration
Seq int
}
type CooldownAPI interface {
GetCooldown() (Snapshot, error) // GET returns evidence
PutCooldown(time.Duration) error // PUT returns only an acknowledgement
}
// VerifiedCooldown is the ONLY state publishable to the board. Every field is
// derived from GET results; PUTs contribute nothing.
type VerifiedCooldown struct {
Value time.Duration
Seq int
Source string
Verified bool
Stale bool
LastErr error
}
func (r *Reconciler) Tick(ctx context.Context) VerifiedCooldown {
r.mu.Lock()
defer r.mu.Unlock()
// 1. GET is the only truth — every tick, unconditionally.
snap, err := r.api.GetCooldown()
if err != nil {
r.state.Stale = true
r.state.LastErr = fmt.Errorf("get: %w", err)
r.putFails++ // never act blind on a failed read
return r.state
}
r.state.Value, r.state.Seq, r.state.Verified = snap.Value, snap.Seq, true
r.state.Stale, r.state.LastErr = false, nil
want, err := r.desired() // pin (or durable escalation policy) = ground truth
if err != nil {
r.state.Source = "policy unavailable; reporting verified value"
return r.state
}
if snap.Value == want.Value { // no drift → no PUT (idempotent)
r.state.Source = "in-sync: " + want.Label
return r.state
}
// 2. Drift (e.g. auto-heal reasserted the pin over our escalation).
if !r.canPut() { // flapping guard: exponential backoff on reverts
r.state.Source = "drift detected; re-escalation throttled by backoff"
return r.state
}
if err := r.api.PutCooldown(want.Value); err != nil {
r.putFails++
r.lastPut = time.Now()
r.state.Source = "reconciliation failed (put error); reporting verified " + r.state.Value.String()
return r.state // board sees the verified ACTUAL, never the wish
}
putAt := time.Now()
// 3. Re-verify with a GET before claiming anything. This is the line the
// bug report was missing.
post, err := r.api.GetCooldown()
if err != nil {
r.state.Stale = true
r.putFails++
r.lastPut = putAt
return r.state // value stays at the PRE-PUT verified value
}
r.state.Value, r.state.Seq, r.state.Verified = post.Value, post.Seq, true
r.lastPut = putAt
if post.Value != want.Value {
r.putFails++
r.state.Source = "reverted by pin reassertion (put not durable); reporting verified " + post.Value.String()
} else {
r.putFails = 0
r.state.Source = "converged: " + want.Label + " (verified by GET)"
}
return r.state
}
// canPut: after N consecutive failed reconciliations, wait 2^N intervals so
// the tick never fights auto-heal in a tight loop yet still converges when
// the adversary stops reverting.
func (r *Reconciler) canPut() bool {
if r.lastPut.IsZero() {
return true
}
wait := r.interval
for i := 0; i < r.putFails && i < 8; i++ {
wait *= 2
}
return time.Since(r.lastPut) >= wait
}
Three rules encode the lesson:
1. GET every tick, treat pin as ground truth — the scheduler's GET result is the only thing that may become a board claim (r.state.Value is only ever assigned from a Snapshot).
2. PUT is best-effort — success is never recorded as state; the post-PUT re-verify GET decides.
3. Reverts are reported, not hidden — a reasserted pin surfaces as "reverted by pin reassertion" / "throttled by backoff" with the verified 900s value, so operators see the escalation is not holding instead of a fabricated 43200s.
Verified by a runnable Go test suite (`cooldown_test.go`, 9 tests) against a `FakeScheduler` that models the adversary exactly: PUTs are accepted but the fleet.toml pin is reasserted — synchronously (`healAfter=0`, the worst case from the bug) or after a healing window. Every successful GET is recorded by unique `Seq`, so tests can *prove* the board only ever published what a GET returned:
```go
func assertNoFabrication(t *testing.T, fake *FakeScheduler, states []VerifiedCooldown) {
for i, s := range states {
val, ok := fake.SnapshotBySeq(s.Seq)
if !ok || val != s.Value {
t.Fatalf("tick %d: FABRICATED state — board published %v but GET#%d returned %v", ...)
}
}
}
```
`go test -race -count=3` → **9/9 PASS, 3 consecutive runs**; timing-sensitive tests additionally `-count=10` → all pass, no flakes. `go vet` clean, `gofmt` clean.
| Test | Scenario | Proves |
|---|---|---|
| `TestPinReassertsImmediatelyReproducesBug` | healAfter=0, 50 ticks | **The reported bug**: board never claims 43200s; tick detects revert (`reverted by pin reassertion`), keeps re-escalating (≥3 PUTs) but bounded (≤10 → no tight loop); flapping guard engages |
| `TestHealingWindowConvergenceAndRevertDetection` | heal 80ms, tick 5ms, 300ms | Escalation genuinely converges (verified by GET), each heal is re-detected and re-converged (PUT oscillation ≥3) |
| `TestPolicyRevertToPin` | escalation window ends | Policy change converges scheduler back to the pin, verified by GET |
| `TestGetFailureNeverActsBlind` | GET outage | No PUT during outage; state flagged stale, never claims verification; recovers cleanly |
| `TestPutFailureReportsVerifiedActual` | PUT rejected then recovered | Board reports verified 900s + `put error`, never the desired 43200s; retries backed off; converges after recovery |
| `TestIdempotentNoOp` | no drift | Zero PUTs when actual == desired == pin |
| `TestReverifyGetFailureKeepsPrevVerifiedValue` | re-verify GET fails post-PUT | Board shows pre-PUT verified value (Seq matches), flagged stale — unverified PUT result never leaks |
| `TestConcurrentTicksAgainstAutoHeal` | 4 goroutines × 40 ticks, heal every 25ms, `-race` | Race detector clean; no-fabrication invariant holds under concurrency |
| `TestRunLoopPublishesUntilCancel` | production `Run` entry | Ticks publish until ctx cancel; all published states GET-derived |
Edge cases covered: immediate vs delayed pin reassertion, GET outage, PUT rejection, re-verify failure, policy switch mid-run, concurrent ticks + concurrent heals, no-drift no-op. The two failures I hit during development were over-strict test expectations (a revert landing *between* ticks shows as re-convergence, not a `reverted` label; backoff still throttles the first tick after PUT recovery) — corrected by asserting the actual behavioral contract, then everything passed.{"model": "deepseek-v4-flash", "problem_class": "go-audit-idle-tick-cooldown-pin", "result": "passed", "tests": 9}Root cause. The old tick was a fire-and-forget PUT: it escalated via PUT /cooldown=43200s, then published "43200s" to the board directly from its own PUT. fleet-auto-heal reasserted the fleet.toml pin (900s) and the board was left claiming state the scheduler never durably held. PUT is an acknowledgement, not a commit — against a pin, it is non-durable by design.
Fix. Rewrite the tick as a GET-truth reconcile loop with a VerifiedCooldown contract that makes fabrication structurally impossible. Deliverable: ~/go-audit-idle-tick-cooldown-pin/ (cooldown.go, scheduler.go, cooldown_test.go, go.mod).
The core fix (cooldown.go):
// Snapshot is one successful GET. Seq uniquely identifies the read so any
// published value can be proven to be literally what a GET returned.
type Snapshot struct {
Value time.Duration
Seq int
}
type CooldownAPI interface {
GetCooldown() (Snapshot, error) // GET returns evidence
PutCooldown(time.Duration) error // PUT returns only an acknowledgement
}
// VerifiedCooldown is the ONLY state publishable to the board. Every field is
// derived from GET results; PUTs contribute nothing.
type VerifiedCooldown struct {
Value time.Duration
Seq int
Source string
Verified bool
Stale bool
LastErr error
}
func (r *Reconciler) Tick(ctx context.Context) VerifiedCooldown {
r.mu.Lock()
defer r.mu.Unlock()
// 1. GET is the only truth — every tick, unconditionally.
snap, err := r.api.GetCooldown()
if err != nil {
r.state.Stale = true
r.state.LastErr = fmt.Errorf("get: %w", err)
r.putFails++ // never act blind on a failed read
return r.state
}
r.state.Value, r.state.Seq, r.state.Verified = snap.Value, snap.Seq, true
r.state.Stale, r.state.LastErr = false, nil
want, err := r.desired() // pin (or durable escalation policy) = ground truth
if err != nil {
r.state.Source = "policy unavailable; reporting verified value"
return r.state
}
if snap.Value == want.Value { // no drift → no PUT (idempotent)
r.state.Source = "in-sync: " + want.Label
return r.state
}
// 2. Drift (e.g. auto-heal reasserted the pin over our escalation).
if !r.canPut() { // flapping guard: exponential backoff on reverts
r.state.Source = "drift detected; re-escalation throttled by backoff"
return r.state
}
if err := r.api.PutCooldown(want.Value); err != nil {
r.putFails++
r.lastPut = time.Now()
r.state.Source = "reconciliation failed (put error); reporting verified " + r.state.Value.String()
return r.state // board sees the verified ACTUAL, never the wish
}
putAt := time.Now()
// 3. Re-verify with a GET before claiming anything. This is the line the
// bug report was missing.
post, err := r.api.GetCooldown()
if err != nil {
r.state.Stale = true
r.putFails++
r.lastPut = putAt
return r.state // value stays at the PRE-PUT verified value
}
r.state.Value, r.state.Seq, r.state.Verified = post.Value, post.Seq, true
r.lastPut = putAt
if post.Value != want.Value {
r.putFails++
r.state.Source = "reverted by pin reassertion (put not durable); reporting verified " + post.Value.String()
} else {
r.putFails = 0
r.state.Source = "converged: " + want.Label + " (verified by GET)"
}
return r.state
}
// canPut: after N consecutive failed reconciliations, wait 2^N intervals so
// the tick never fights auto-heal in a tight loop yet still converges when
// the adversary stops reverting.
func (r *Reconciler) canPut() bool {
if r.lastPut.IsZero() {
return true
}
wait := r.interval
for i := 0; i < r.putFails && i < 8; i++ {
wait *= 2
}
return time.Since(r.lastPut) >= wait
}
Three rules encode the lesson:
1. GET every tick, treat pin as ground truth — the scheduler's GET result is the only thing that may become a board claim (r.state.Value is only ever assigned from a Snapshot).
2. PUT is best-effort — success is never recorded as state; the post-PUT re-verify GET decides.
3. Reverts are reported, not hidden — a reasserted pin surfaces as "reverted by pin reassertion" / "throttled by backoff" with the verified 900s value, so operators see the escalation is not holding instead of a fabricated 43200s.
Verified by a runnable Go test suite (`cooldown_test.go`, 9 tests) against a `FakeScheduler` that models the adversary exactly: PUTs are accepted but the fleet.toml pin is reasserted — synchronously (`healAfter=0`, the worst case from the bug) or after a healing window. Every successful GET is recorded by unique `Seq`, so tests can *prove* the board only ever published what a GET returned:
```go
func assertNoFabrication(t *testing.T, fake *FakeScheduler, states []VerifiedCooldown) {
for i, s := range states {
val, ok := fake.SnapshotBySeq(s.Seq)
if !ok || val != s.Value {
t.Fatalf("tick %d: FABRICATED state — board published %v but GET#%d returned %v", ...)
}
}
}
```
`go test -race -count=3` → **9/9 PASS, 3 consecutive runs**; timing-sensitive tests additionally `-count=10` → all pass, no flakes. `go vet` clean, `gofmt` clean.
| Test | Scenario | Proves |
|---|---|---|
| `TestPinReassertsImmediatelyReproducesBug` | healAfter=0, 50 ticks | **The reported bug**: board never claims 43200s; tick detects revert (`reverted by pin reassertion`), keeps re-escalating (≥3 PUTs) but bounded (≤10 → no tight loop); flapping guard engages |
| `TestHealingWindowConvergenceAndRevertDetection` | heal 80ms, tick 5ms, 300ms | Escalation genuinely converges (verified by GET), each heal is re-detected and re-converged (PUT oscillation ≥3) |
| `TestPolicyRevertToPin` | escalation window ends | Policy change converges scheduler back to the pin, verified by GET |
| `TestGetFailureNeverActsBlind` | GET outage | No PUT during outage; state flagged stale, never claims verification; recovers cleanly |
| `TestPutFailureReportsVerifiedActual` | PUT rejected then recovered | Board reports verified 900s + `put error`, never the desired 43200s; retries backed off; converges after recovery |
| `TestIdempotentNoOp` | no drift | Zero PUTs when actual == desired == pin |
| `TestReverifyGetFailureKeepsPrevVerifiedValue` | re-verify GET fails post-PUT | Board shows pre-PUT verified value (Seq matches), flagged stale — unverified PUT result never leaks |
| `TestConcurrentTicksAgainstAutoHeal` | 4 goroutines × 40 ticks, heal every 25ms, `-race` | Race detector clean; no-fabrication invariant holds under concurrency |
| `TestRunLoopPublishesUntilCancel` | production `Run` entry | Ticks publish until ctx cancel; all published states GET-derived |
Edge cases covered: immediate vs delayed pin reassertion, GET outage, PUT rejection, re-verify failure, policy switch mid-run, concurrent ticks + concurrent heals, no-drift no-op. The two failures I hit during development were over-strict test expectations (a revert landing *between* ticks shows as re-convergence, not a `reverted` label; backoff still throttles the first tick after PUT recovery) — corrected by asserting the actual behavioral contract, then everything passed.{"model": "deepseek-v4-flash", "problem_class": "go-audit-idle-tick-cooldown-pin", "result": "passed", "tests": 9}