go-scheduler-starvation-alert-throttle
Root cause. CheckStarvation emits a MEDIUM starvation event on every 60s evaluation cycle for each project past 2x cooldown. Because the escalator is constructed fresh per cycle, no in-memory/struct state survives between cycles — the only durable place to remember "we already told the user" is the events table itself. Observed: 179 events/33min across projects.
Fix. A DB-backed per-project throttle: lastStarvationEvent queries the events table filtering severity, component, and json_extract(details, '$.project'); ShouldEmit suppresses re-emission while a matching event exists within a 30-minute window. The DB is the source of truth, so fresh escalators see prior emissions automatically.
const (
StarvationSeverity = "MEDIUM"
StarvationComponent = "starvation"
ThrottleWindow = 30 * time.Minute
)
// StarvationThrottle carries no mutable state on purpose: the escalator is
// rebuilt every cycle, so throttle state must live in the events table.
type StarvationThrottle struct {
db *sql.DB
now func() time.Time
window time.Duration
}
// lastStarvationEvent returns the most recent MEDIUM starvation event for
// projectID, or zero time if none exists. Per-project keying comes from
// json_extract(details,'$.project'); severity+component keep unrelated
// alerts (HIGH, cpu-throttle, ...) from suppressing this one.
func (t *StarvationThrottle) lastStarvationEvent(ctx context.Context, projectID string) (time.Time, error) {
var created string
err := t.db.QueryRowContext(ctx, `
SELECT MAX(created_at)
FROM events
WHERE severity = ? AND component = ?
AND json_extract(details, '$.project') = ?`,
StarvationSeverity, StarvationComponent, projectID,
).Scan(&created)
if err == sql.ErrNoRows || created == "" {
return time.Time{}, nil // never emitted for this project -> emit
}
if err != nil {
return time.Time{}, err
}
return time.Parse(time.RFC3339Nano, created)
}
// ShouldEmit suppresses re-emission within the 30-minute window (boundary
// inclusive: now - last >= window re-emits).
func (t *StarvationThrottle) ShouldEmit(ctx context.Context, projectID string) (bool, error) {
last, err := t.lastStarvationEvent(ctx, projectID)
if err != nil {
return false, err
}
if last.IsZero() {
return true, nil
}
return t.now().Sub(last) >= t.window, nil
}
Integration in CheckStarvation — the throttle is rebuilt per cycle from the DB handle, so the fresh-escalator constraint is satisfied:
func (e *Escalator) CheckStarvation(ctx context.Context, projects []Project) error {
throttle := NewStarvationThrottle(e.db, e.now, ThrottleWindow)
for _, p := range projects {
if !starvedPast2xCooldown(p) {
continue
}
ok, err := throttle.ShouldEmit(ctx, p.ID)
if err != nil {
return fmt.Errorf("starvation throttle: %w", err)
}
if !ok {
continue // MEDIUM event exists within 30-min window: suppress
}
e.emitEvent(ctx, StarvationSeverity, StarvationComponent,
map[string]any{"project": p.ID, "reason": "starvation past 2x cooldown"})
}
return nil
}
Emission stores created_at as RFC3339 (UTC) so MAX(created_at) is lexicographically correct; an index on (severity, component, created_at) keeps the lookup cheap.
The production repo isn't mounted in this environment, so I built a faithful reproduction of the described architecture (`/tmp/throttle-demo` — events table with `severity/component/details/created_at`, escalator constructed fresh each 60s cycle, same SQL) and ran the fix against **Go 1.26 + real SQLite** (`modernc.org/sqlite`; same query semantics confirmed against the system `sqlite3` CLI). Results, `go test -race`: ``` === RUN TestThrottleSuppressesReemit --- PASS (1 emission, 29 suppressed cycles, exactly 1 row) === RUN TestWindowExpiryReemit --- PASS (suppressed at t+1m, re-emits at exactly t+30m, suppressed again) === RUN TestPerProjectIsolation --- PASS (a/b throttled, c fires, then c throttled too) === RUN Test30MinObservedBugRegression --- PASS (33 cycles x 6 projects -> 12 events = 2/project, not 179) === RUN TestUnrelatedEventsDoNotThrottle --- PASS (HIGH / other-component / no-$.project rows don't suppress) ok throttle-demo 1.155s ``` **The 3 required tests**, as specified: 1. **Throttle** — `TestThrottleSuppressesReemit`: fresh escalator each cycle; cycle 0 emits 1, cycles 1–29 (t+1m…t+29m) emit 0; table holds exactly 1 row. 2. **Window expiry re-emit** — `TestWindowExpiryReemit`: suppressed inside window, re-emits at exactly t+30m (boundary inclusive via `>=`), then suppressed again under the new window. Uses an injected clock so no 30-min wait. 3. **Per-project isolation** — `TestPerProjectIsolation`: throttling keyed on `json_extract(details,'$.project')`; suppressed proj-a/proj-b don't block never-emitted proj-c, and each project carries its own window. **Edge cases covered:** never-emitted project → emits; `MAX(created_at)` picks the latest of multiple events; missing `$.project` key → `json_extract` is NULL → no match → emits; different severity (`HIGH`) or component (`cpu-throttle`) rows don't suppress; RFC3339 timestamps parse and compare correctly; 33-minute/179-event regression collapses to 2 events per project.
{"model": "deepseek-v4-flash", "problem_class": "go-scheduler-starvation-alert-throttle", "result": "passed", "tests": 3}