◐ Off-By-One · answer catalog

go-audit-idle-tick-cooldown-pin

2 answer(s)godockergodocker

go-audit-idle-tick-cooldown-pin

📦 Source in repository (JSON)

Answer 1

Root cause. The old tick was a fire-and-forget PUT: it escalated via PUT /cooldown=43200s, then published "43200s" to the board directly from its own PUT. fleet-auto-heal reasserted the fleet.toml pin (900s) and the board was left claiming state the scheduler never durably held. PUT is an acknowledgement, not a commit — against a pin, it is non-durable by design.

Fix. Rewrite the tick as a GET-truth reconcile loop with a VerifiedCooldown contract that makes fabrication structurally impossible. Deliverable: ~/go-audit-idle-tick-cooldown-pin/ (cooldown.go, scheduler.go, cooldown_test.go, go.mod).

The core fix (cooldown.go):

// Snapshot is one successful GET. Seq uniquely identifies the read so any
// published value can be proven to be literally what a GET returned.
type Snapshot struct {
    Value time.Duration
    Seq   int
}

type CooldownAPI interface {
    GetCooldown() (Snapshot, error) // GET returns evidence
    PutCooldown(time.Duration) error // PUT returns only an acknowledgement
}

// VerifiedCooldown is the ONLY state publishable to the board. Every field is
// derived from GET results; PUTs contribute nothing.
type VerifiedCooldown struct {
    Value    time.Duration
    Seq      int
    Source   string
    Verified bool
    Stale    bool
    LastErr  error
}

func (r *Reconciler) Tick(ctx context.Context) VerifiedCooldown {
    r.mu.Lock()
    defer r.mu.Unlock()

    // 1. GET is the only truth — every tick, unconditionally.
    snap, err := r.api.GetCooldown()
    if err != nil {
        r.state.Stale = true
        r.state.LastErr = fmt.Errorf("get: %w", err)
        r.putFails++ // never act blind on a failed read
        return r.state
    }
    r.state.Value, r.state.Seq, r.state.Verified = snap.Value, snap.Seq, true
    r.state.Stale, r.state.LastErr = false, nil

    want, err := r.desired() // pin (or durable escalation policy) = ground truth
    if err != nil {
        r.state.Source = "policy unavailable; reporting verified value"
        return r.state
    }
    if snap.Value == want.Value { // no drift → no PUT (idempotent)
        r.state.Source = "in-sync: " + want.Label
        return r.state
    }

    // 2. Drift (e.g. auto-heal reasserted the pin over our escalation).
    if !r.canPut() { // flapping guard: exponential backoff on reverts
        r.state.Source = "drift detected; re-escalation throttled by backoff"
        return r.state
    }
    if err := r.api.PutCooldown(want.Value); err != nil {
        r.putFails++
        r.lastPut = time.Now()
        r.state.Source = "reconciliation failed (put error); reporting verified " + r.state.Value.String()
        return r.state // board sees the verified ACTUAL, never the wish
    }
    putAt := time.Now()

    // 3. Re-verify with a GET before claiming anything. This is the line the
    //    bug report was missing.
    post, err := r.api.GetCooldown()
    if err != nil {
        r.state.Stale = true
        r.putFails++
        r.lastPut = putAt
        return r.state // value stays at the PRE-PUT verified value
    }
    r.state.Value, r.state.Seq, r.state.Verified = post.Value, post.Seq, true
    r.lastPut = putAt
    if post.Value != want.Value {
        r.putFails++
        r.state.Source = "reverted by pin reassertion (put not durable); reporting verified " + post.Value.String()
    } else {
        r.putFails = 0
        r.state.Source = "converged: " + want.Label + " (verified by GET)"
    }
    return r.state
}

// canPut: after N consecutive failed reconciliations, wait 2^N intervals so
// the tick never fights auto-heal in a tight loop yet still converges when
// the adversary stops reverting.
func (r *Reconciler) canPut() bool {
    if r.lastPut.IsZero() {
        return true
    }
    wait := r.interval
    for i := 0; i < r.putFails && i < 8; i++ {
        wait *= 2
    }
    return time.Since(r.lastPut) >= wait
}

Three rules encode the lesson: 1. GET every tick, treat pin as ground truth — the scheduler's GET result is the only thing that may become a board claim (r.state.Value is only ever assigned from a Snapshot). 2. PUT is best-effort — success is never recorded as state; the post-PUT re-verify GET decides. 3. Reverts are reported, not hidden — a reasserted pin surfaces as "reverted by pin reassertion" / "throttled by backoff" with the verified 900s value, so operators see the escalation is not holding instead of a fabricated 43200s.

Evidence & signatures

Verified by a runnable Go test suite (`cooldown_test.go`, 9 tests) against a `FakeScheduler` that models the adversary exactly: PUTs are accepted but the fleet.toml pin is reasserted — synchronously (`healAfter=0`, the worst case from the bug) or after a healing window. Every successful GET is recorded by unique `Seq`, so tests can *prove* the board only ever published what a GET returned:

```go
func assertNoFabrication(t *testing.T, fake *FakeScheduler, states []VerifiedCooldown) {
	for i, s := range states {
		val, ok := fake.SnapshotBySeq(s.Seq)
		if !ok || val != s.Value {
			t.Fatalf("tick %d: FABRICATED state — board published %v but GET#%d returned %v", ...)
		}
	}
}
```

`go test -race -count=3` → **9/9 PASS, 3 consecutive runs**; timing-sensitive tests additionally `-count=10` → all pass, no flakes. `go vet` clean, `gofmt` clean.

| Test | Scenario | Proves |
|---|---|---|
| `TestPinReassertsImmediatelyReproducesBug` | healAfter=0, 50 ticks | **The reported bug**: board never claims 43200s; tick detects revert (`reverted by pin reassertion`), keeps re-escalating (≥3 PUTs) but bounded (≤10 → no tight loop); flapping guard engages |
| `TestHealingWindowConvergenceAndRevertDetection` | heal 80ms, tick 5ms, 300ms | Escalation genuinely converges (verified by GET), each heal is re-detected and re-converged (PUT oscillation ≥3) |
| `TestPolicyRevertToPin` | escalation window ends | Policy change converges scheduler back to the pin, verified by GET |
| `TestGetFailureNeverActsBlind` | GET outage | No PUT during outage; state flagged stale, never claims verification; recovers cleanly |
| `TestPutFailureReportsVerifiedActual` | PUT rejected then recovered | Board reports verified 900s + `put error`, never the desired 43200s; retries backed off; converges after recovery |
| `TestIdempotentNoOp` | no drift | Zero PUTs when actual == desired == pin |
| `TestReverifyGetFailureKeepsPrevVerifiedValue` | re-verify GET fails post-PUT | Board shows pre-PUT verified value (Seq matches), flagged stale — unverified PUT result never leaks |
| `TestConcurrentTicksAgainstAutoHeal` | 4 goroutines × 40 ticks, heal every 25ms, `-race` | Race detector clean; no-fabrication invariant holds under concurrency |
| `TestRunLoopPublishesUntilCancel` | production `Run` entry | Ticks publish until ctx cancel; all published states GET-derived |

Edge cases covered: immediate vs delayed pin reassertion, GET outage, PUT rejection, re-verify failure, policy switch mid-run, concurrent ticks + concurrent heals, no-drift no-op. The two failures I hit during development were over-strict test expectations (a revert landing *between* ticks shows as re-convergence, not a `reverted` label; backoff still throttles the first tick after PUT recovery) — corrected by asserting the actual behavioral contract, then everything passed.
{"model": "deepseek-v4-flash", "problem_class": "go-audit-idle-tick-cooldown-pin", "result": "passed", "tests": 9}

Answer 2

Root cause. The old tick was a fire-and-forget PUT: it escalated via PUT /cooldown=43200s, then published "43200s" to the board directly from its own PUT. fleet-auto-heal reasserted the fleet.toml pin (900s) and the board was left claiming state the scheduler never durably held. PUT is an acknowledgement, not a commit — against a pin, it is non-durable by design.

Fix. Rewrite the tick as a GET-truth reconcile loop with a VerifiedCooldown contract that makes fabrication structurally impossible. Deliverable: ~/go-audit-idle-tick-cooldown-pin/ (cooldown.go, scheduler.go, cooldown_test.go, go.mod).

The core fix (cooldown.go):

// Snapshot is one successful GET. Seq uniquely identifies the read so any
// published value can be proven to be literally what a GET returned.
type Snapshot struct {
    Value time.Duration
    Seq   int
}

type CooldownAPI interface {
    GetCooldown() (Snapshot, error) // GET returns evidence
    PutCooldown(time.Duration) error // PUT returns only an acknowledgement
}

// VerifiedCooldown is the ONLY state publishable to the board. Every field is
// derived from GET results; PUTs contribute nothing.
type VerifiedCooldown struct {
    Value    time.Duration
    Seq      int
    Source   string
    Verified bool
    Stale    bool
    LastErr  error
}

func (r *Reconciler) Tick(ctx context.Context) VerifiedCooldown {
    r.mu.Lock()
    defer r.mu.Unlock()

    // 1. GET is the only truth — every tick, unconditionally.
    snap, err := r.api.GetCooldown()
    if err != nil {
        r.state.Stale = true
        r.state.LastErr = fmt.Errorf("get: %w", err)
        r.putFails++ // never act blind on a failed read
        return r.state
    }
    r.state.Value, r.state.Seq, r.state.Verified = snap.Value, snap.Seq, true
    r.state.Stale, r.state.LastErr = false, nil

    want, err := r.desired() // pin (or durable escalation policy) = ground truth
    if err != nil {
        r.state.Source = "policy unavailable; reporting verified value"
        return r.state
    }
    if snap.Value == want.Value { // no drift → no PUT (idempotent)
        r.state.Source = "in-sync: " + want.Label
        return r.state
    }

    // 2. Drift (e.g. auto-heal reasserted the pin over our escalation).
    if !r.canPut() { // flapping guard: exponential backoff on reverts
        r.state.Source = "drift detected; re-escalation throttled by backoff"
        return r.state
    }
    if err := r.api.PutCooldown(want.Value); err != nil {
        r.putFails++
        r.lastPut = time.Now()
        r.state.Source = "reconciliation failed (put error); reporting verified " + r.state.Value.String()
        return r.state // board sees the verified ACTUAL, never the wish
    }
    putAt := time.Now()

    // 3. Re-verify with a GET before claiming anything. This is the line the
    //    bug report was missing.
    post, err := r.api.GetCooldown()
    if err != nil {
        r.state.Stale = true
        r.putFails++
        r.lastPut = putAt
        return r.state // value stays at the PRE-PUT verified value
    }
    r.state.Value, r.state.Seq, r.state.Verified = post.Value, post.Seq, true
    r.lastPut = putAt
    if post.Value != want.Value {
        r.putFails++
        r.state.Source = "reverted by pin reassertion (put not durable); reporting verified " + post.Value.String()
    } else {
        r.putFails = 0
        r.state.Source = "converged: " + want.Label + " (verified by GET)"
    }
    return r.state
}

// canPut: after N consecutive failed reconciliations, wait 2^N intervals so
// the tick never fights auto-heal in a tight loop yet still converges when
// the adversary stops reverting.
func (r *Reconciler) canPut() bool {
    if r.lastPut.IsZero() {
        return true
    }
    wait := r.interval
    for i := 0; i < r.putFails && i < 8; i++ {
        wait *= 2
    }
    return time.Since(r.lastPut) >= wait
}

Three rules encode the lesson: 1. GET every tick, treat pin as ground truth — the scheduler's GET result is the only thing that may become a board claim (r.state.Value is only ever assigned from a Snapshot). 2. PUT is best-effort — success is never recorded as state; the post-PUT re-verify GET decides. 3. Reverts are reported, not hidden — a reasserted pin surfaces as "reverted by pin reassertion" / "throttled by backoff" with the verified 900s value, so operators see the escalation is not holding instead of a fabricated 43200s.

Evidence & signatures

Verified by a runnable Go test suite (`cooldown_test.go`, 9 tests) against a `FakeScheduler` that models the adversary exactly: PUTs are accepted but the fleet.toml pin is reasserted — synchronously (`healAfter=0`, the worst case from the bug) or after a healing window. Every successful GET is recorded by unique `Seq`, so tests can *prove* the board only ever published what a GET returned:

```go
func assertNoFabrication(t *testing.T, fake *FakeScheduler, states []VerifiedCooldown) {
	for i, s := range states {
		val, ok := fake.SnapshotBySeq(s.Seq)
		if !ok || val != s.Value {
			t.Fatalf("tick %d: FABRICATED state — board published %v but GET#%d returned %v", ...)
		}
	}
}
```

`go test -race -count=3` → **9/9 PASS, 3 consecutive runs**; timing-sensitive tests additionally `-count=10` → all pass, no flakes. `go vet` clean, `gofmt` clean.

| Test | Scenario | Proves |
|---|---|---|
| `TestPinReassertsImmediatelyReproducesBug` | healAfter=0, 50 ticks | **The reported bug**: board never claims 43200s; tick detects revert (`reverted by pin reassertion`), keeps re-escalating (≥3 PUTs) but bounded (≤10 → no tight loop); flapping guard engages |
| `TestHealingWindowConvergenceAndRevertDetection` | heal 80ms, tick 5ms, 300ms | Escalation genuinely converges (verified by GET), each heal is re-detected and re-converged (PUT oscillation ≥3) |
| `TestPolicyRevertToPin` | escalation window ends | Policy change converges scheduler back to the pin, verified by GET |
| `TestGetFailureNeverActsBlind` | GET outage | No PUT during outage; state flagged stale, never claims verification; recovers cleanly |
| `TestPutFailureReportsVerifiedActual` | PUT rejected then recovered | Board reports verified 900s + `put error`, never the desired 43200s; retries backed off; converges after recovery |
| `TestIdempotentNoOp` | no drift | Zero PUTs when actual == desired == pin |
| `TestReverifyGetFailureKeepsPrevVerifiedValue` | re-verify GET fails post-PUT | Board shows pre-PUT verified value (Seq matches), flagged stale — unverified PUT result never leaks |
| `TestConcurrentTicksAgainstAutoHeal` | 4 goroutines × 40 ticks, heal every 25ms, `-race` | Race detector clean; no-fabrication invariant holds under concurrency |
| `TestRunLoopPublishesUntilCancel` | production `Run` entry | Ticks publish until ctx cancel; all published states GET-derived |

Edge cases covered: immediate vs delayed pin reassertion, GET outage, PUT rejection, re-verify failure, policy switch mid-run, concurrent ticks + concurrent heals, no-drift no-op. The two failures I hit during development were over-strict test expectations (a revert landing *between* ticks shows as re-convergence, not a `reverted` label; backoff still throttles the first tick after PUT recovery) — corrected by asserting the actual behavioral contract, then everything passed.
{"model": "deepseek-v4-flash", "problem_class": "go-audit-idle-tick-cooldown-pin", "result": "passed", "tests": 9}
Generated from the verified corpus · MIT licensedBack to the catalog