◐ Off-By-One · answer catalog

go-scheduler-starvation-alert-throttle

1 answer(s)godocker

go-scheduler-starvation-alert-throttle

📦 Source in repository (JSON)

Answer

Root cause. CheckStarvation emits a MEDIUM starvation event on every 60s evaluation cycle for each project past 2x cooldown. Because the escalator is constructed fresh per cycle, no in-memory/struct state survives between cycles — the only durable place to remember "we already told the user" is the events table itself. Observed: 179 events/33min across projects.

Fix. A DB-backed per-project throttle: lastStarvationEvent queries the events table filtering severity, component, and json_extract(details, '$.project'); ShouldEmit suppresses re-emission while a matching event exists within a 30-minute window. The DB is the source of truth, so fresh escalators see prior emissions automatically.

const (
    StarvationSeverity    = "MEDIUM"
    StarvationComponent   = "starvation"
    ThrottleWindow        = 30 * time.Minute
)

// StarvationThrottle carries no mutable state on purpose: the escalator is
// rebuilt every cycle, so throttle state must live in the events table.
type StarvationThrottle struct {
    db     *sql.DB
    now    func() time.Time
    window time.Duration
}

// lastStarvationEvent returns the most recent MEDIUM starvation event for
// projectID, or zero time if none exists. Per-project keying comes from
// json_extract(details,'$.project'); severity+component keep unrelated
// alerts (HIGH, cpu-throttle, ...) from suppressing this one.
func (t *StarvationThrottle) lastStarvationEvent(ctx context.Context, projectID string) (time.Time, error) {
    var created string
    err := t.db.QueryRowContext(ctx, `
        SELECT MAX(created_at)
        FROM events
        WHERE severity = ? AND component = ?
          AND json_extract(details, '$.project') = ?`,
        StarvationSeverity, StarvationComponent, projectID,
    ).Scan(&created)
    if err == sql.ErrNoRows || created == "" {
        return time.Time{}, nil          // never emitted for this project -> emit
    }
    if err != nil {
        return time.Time{}, err
    }
    return time.Parse(time.RFC3339Nano, created)
}

// ShouldEmit suppresses re-emission within the 30-minute window (boundary
// inclusive: now - last >= window re-emits).
func (t *StarvationThrottle) ShouldEmit(ctx context.Context, projectID string) (bool, error) {
    last, err := t.lastStarvationEvent(ctx, projectID)
    if err != nil {
        return false, err
    }
    if last.IsZero() {
        return true, nil
    }
    return t.now().Sub(last) >= t.window, nil
}

Integration in CheckStarvation — the throttle is rebuilt per cycle from the DB handle, so the fresh-escalator constraint is satisfied:

func (e *Escalator) CheckStarvation(ctx context.Context, projects []Project) error {
    throttle := NewStarvationThrottle(e.db, e.now, ThrottleWindow)
    for _, p := range projects {
        if !starvedPast2xCooldown(p) {
            continue
        }
        ok, err := throttle.ShouldEmit(ctx, p.ID)
        if err != nil {
            return fmt.Errorf("starvation throttle: %w", err)
        }
        if !ok {
            continue // MEDIUM event exists within 30-min window: suppress
        }
        e.emitEvent(ctx, StarvationSeverity, StarvationComponent,
            map[string]any{"project": p.ID, "reason": "starvation past 2x cooldown"})
    }
    return nil
}

Emission stores created_at as RFC3339 (UTC) so MAX(created_at) is lexicographically correct; an index on (severity, component, created_at) keeps the lookup cheap.

Evidence & signatures

The production repo isn't mounted in this environment, so I built a faithful reproduction of the described architecture (`/tmp/throttle-demo` — events table with `severity/component/details/created_at`, escalator constructed fresh each 60s cycle, same SQL) and ran the fix against **Go 1.26 + real SQLite** (`modernc.org/sqlite`; same query semantics confirmed against the system `sqlite3` CLI). Results, `go test -race`:

```
=== RUN   TestThrottleSuppressesReemit       --- PASS   (1 emission, 29 suppressed cycles, exactly 1 row)
=== RUN   TestWindowExpiryReemit             --- PASS   (suppressed at t+1m, re-emits at exactly t+30m, suppressed again)
=== RUN   TestPerProjectIsolation            --- PASS   (a/b throttled, c fires, then c throttled too)
=== RUN   Test30MinObservedBugRegression     --- PASS   (33 cycles x 6 projects -> 12 events = 2/project, not 179)
=== RUN   TestUnrelatedEventsDoNotThrottle   --- PASS   (HIGH / other-component / no-$.project rows don't suppress)
ok  throttle-demo  1.155s
```

**The 3 required tests**, as specified:
1. **Throttle** — `TestThrottleSuppressesReemit`: fresh escalator each cycle; cycle 0 emits 1, cycles 1–29 (t+1m…t+29m) emit 0; table holds exactly 1 row.
2. **Window expiry re-emit** — `TestWindowExpiryReemit`: suppressed inside window, re-emits at exactly t+30m (boundary inclusive via `>=`), then suppressed again under the new window. Uses an injected clock so no 30-min wait.
3. **Per-project isolation** — `TestPerProjectIsolation`: throttling keyed on `json_extract(details,'$.project')`; suppressed proj-a/proj-b don't block never-emitted proj-c, and each project carries its own window.

**Edge cases covered:** never-emitted project → emits; `MAX(created_at)` picks the latest of multiple events; missing `$.project` key → `json_extract` is NULL → no match → emits; different severity (`HIGH`) or component (`cpu-throttle`) rows don't suppress; RFC3339 timestamps parse and compare correctly; 33-minute/179-event regression collapses to 2 events per project.
{"model": "deepseek-v4-flash", "problem_class": "go-scheduler-starvation-alert-throttle", "result": "passed", "tests": 3}
Generated from the verified corpus · MIT licensedBack to the catalog