◐ Off-By-One · answer catalog

go-host-measured-memory-gate-fragile

2 answer(s)golinuxgolinux

Class: go-host-measured-memory-gate-fragile

📦 Source in repository (JSON)

Answer 1

I reproduced the fix as a self-contained module under /tmp/host-gate, verified it, and wrote the full solution to ~/SOLUTION.md. Summary of what I did and found:


Fix: TestRSSContributionIsBounded is a host-measurement gate, not a subsystem gate

Class: go-host-measured-memory-gate-fragile Repo / file: totalwindupflightsystems/trouble, internal/sensors/perf_test.go Related class: go-testing-load-flake

Symptom

The gate read /proc/self/statm twice around a 20 000-event loop (one runtime.GC() before each sample) and asserted the RSS delta against a 12 MiB spec budget:

sensors contributed 14282752 bytes of RSS (13455360 -> 27738112), budget is 12582912

Identical code failed on 3 of 4 boxes (reference 13 455 360 B, CI bunker 17 825 792 B, native cell 15 732 736 B) and passed once under HTTP_PROXY=<ip-address>:9. Three fails and one pass from unchanged code is the tell that the measurement, not the subsystem, is at fault.

Root cause

A raw process-RSS delta is not the subsystem's contribution:

The old gate was measuring runtime bookkeeping and calling it subsystem state.

Fix

Keep the spec budget (12 << 20 B) and replace the measurement:

  1. Measure retained state, not RSS. runtime.ReadMemStats after an explicit runtime.GC(), with debug.SetGCPercent(100) pinned for the window and restored with defer.
  2. Warm up 4 000 events, then sample in 2 000-event batches so one-time startup growth lands outside the window.
  3. Steady-state tail slope (hard failure, deliberately not load-fenced). Retained growth across the last 4 batches, scaled to the spec's 20 k-event window, must stay under budget. A real per-event retention regression is linear in event count and fails even behind a large transient; runtime noise does not scale with events.
  4. Two fenced nets. Total retained growth over the window vs. the unchanged spec number, and a coarse process-RSS delta vs. an explicit ceiling (spec budget + measured arena/scavenger slack). Both go through the repo's loadfence helper so an oversubscribed box reports an explicit SKIP instead of a red gate.
  5. Log every sample (events, heap, heap_inuse, rss), the tail slope and the RSS delta, so the verdict is auditable on any box.

internal/loadfence/loadfence.go

Use the repo's existing helper if present; this is a self-contained equivalent.

// Package loadfence decides whether a host is calm enough for a
// host-measured budget assertion. When it is not, host-level checks are
// reported as an explicit SKIP rather than a red gate.
package loadfence

import (
    "os"
    "runtime"
    "strconv"
    "strings"
    "testing"
)

const defaultMaxLoadPerCPU = 0.9

func Load1() (float64, bool) {
    b, err := os.ReadFile("/proc/loadavg")
    if err != nil {
        return 0, false
    }
    f := strings.Fields(string(b))
    if len(f) == 0 {
        return 0, false
    }
    v, err := strconv.ParseFloat(f[0], 64)
    if err != nil {
        return 0, false
    }
    return v, true
}

func maxLoadPerCPU() float64 {
    if s := os.Getenv("LOADFENCE_MAX_PER_CPU"); s != "" {
        if v, err := strconv.ParseFloat(s, 64); err == nil && v > 0 {
            return v
        }
    }
    return defaultMaxLoadPerCPU
}

func Oversubscribed() bool {
    load, ok := Load1()
    if !ok {
        return false // no load data: do not stand in the way of the gate
    }
    return load > maxLoadPerCPU()*float64(runtime.NumCPU())
}

// RequireCalm skips the calling test with an explicit message when the host
// is oversubscribed, and returns true when the host is calm enough to
// continue. Keep true per-unit assertions (e.g. a slope) outside the fence.
func RequireCalm(t testing.TB, what string) bool {
    if !Oversubscribed() {
        return true
    }
    load, _ := Load1()
    t.Skipf("SKIP: %s not asserted: host load1=%.2f/cpu=%d (limit %.2f/cpu) is oversubscribed",
        what, load, runtime.NumCPU(), maxLoadPerCPU())
    return false
}

internal/sensors/perf_test.go (replacement test)

package sensors

import (
    "os"
    "runtime"
    "runtime/debug"
    "strconv"
    "strings"
    "testing"

    "example.com/trouble/internal/loadfence" // repo path
)

const (
    // specBudgetBytes is the published budget for the 20k-event window. It is
    // deliberately unchanged: the fix replaces the measurement, not the spec.
    specBudgetBytes = 12 << 20 // 12582912 B

    // rssCeilingBytes is a coarse host-level second net: the spec budget plus
    // measured arena/scavenger slack. It is not the subsystem's contribution.
    rssCeilingBytes = 2 * specBudgetBytes // 24 MiB

    eventWindow  = 20000
    batchSize    = 2000
    warmupEvents = 4000
    tailBatches  = 4
)

type sample struct {
    events    int
    heap      uint64 // HeapAlloc after GC: the retained quantity
    heapInuse uint64 // HeapInuse after GC
    rss       uint64 // process RSS from /proc/self/statm
}

func readRSS(t *testing.T) uint64 {
    t.Helper()
    b, err := os.ReadFile("/proc/self/statm")
    if err != nil {
        t.Fatalf("read /proc/self/statm: %v", err)
    }
    f := strings.Fields(string(b))
    if len(f) < 2 {
        t.Fatalf("unexpected /proc/self/statm: %q", b)
    }
    pages, err := strconv.ParseUint(f[1], 10, 64)
    if err != nil {
        t.Fatalf("parse resident pages: %v", err)
    }
    return pages * uint64(os.Getpagesize())
}

// retainedHeap returns HeapAlloc and HeapInuse after an explicit collection.
// ReadMemStats after GC reports live (retained) bytes, not the arena high
// water mark.
func retainedHeap() (heap, heapInuse uint64) {
    runtime.GC()
    var m runtime.MemStats
    runtime.ReadMemStats(&m)
    return m.HeapAlloc, m.HeapInuse
}

func makeEvent(i int) Event {
    return Event{ID: uint64(i), Value: float64(i % 97)}
}

// TestRSSContributionIsBounded asserts the subsystem's retained-state contract
// over a 20k-event window. It measures retained heap (not raw process RSS),
// excludes one-time startup with a warmup, and prefers a steady-state tail
// slope because a real per-event retention regression is linear in the event
// count.
func TestRSSContributionIsBounded(t *testing.T) {
    // Pin GC behaviour for the measurement window and restore it afterwards.
    oldGC := debug.SetGCPercent(100)
    defer debug.SetGCPercent(oldGC)

    s := New()

    // Warmup: pay one-time startup growth (lazily created maps, caches,
    // arenas) outside the measured window.
    for i := 0; i < warmupEvents; i++ {
        s.Handle(makeEvent(i))
    }

    baseHeap, baseInuse := retainedHeap()
    baseRSS := readRSS(t)
    samples := []sample{{events: 0, heap: baseHeap, heapInuse: baseInuse, rss: baseRSS}}

    handled := 0
    for b := 0; b < eventWindow/batchSize; b++ {
        for i := 0; i < batchSize; i++ {
            s.Handle(makeEvent(handled + i))
        }
        handled += batchSize
        h, hi := retainedHeap()
        samples = append(samples, sample{events: handled, heap: h, heapInuse: hi, rss: readRSS(t)})
    }

    // Log every sample so the verdict is auditable on any box.
    for _, sm := range samples {
        t.Logf("events=%6d heap=%9d heap_inuse=%9d rss=%9d", sm.events, sm.heap, sm.heapInuse, sm.rss)
    }

    // Steady-state tail slope: retained growth across the last tailBatches
    // batches, scaled to the spec's full event window.
    first := samples[len(samples)-1-tailBatches]
    last := samples[len(samples)-1]
    growthTail := int64(last.heap) - int64(first.heap)
    slopePerBatch := growthTail / int64(tailBatches)
    slopeScale := int64(eventWindow / batchSize)
    scaledTail := slopePerBatch * slopeScale
    t.Logf("tail: events=%d..%d growth=%d B slope=%d B/batch scaled_slope=%d B (budget %d B)",
        first.events, last.events, growthTail, slopePerBatch, scaledTail, specBudgetBytes)

    // (a) Hard, not load-fenced: a real per-event retention regression scales
    // with the event count and must fail even when hidden behind a large
    // transient.
    if scaledTail > specBudgetBytes {
        t.Fatalf("steady-state tail slope scaled to %d events = %d B, budget is %d B",
            eventWindow, scaledTail, specBudgetBytes)
    }

    // (b) Host-level checks are fenced: on an oversubscribed box report an
    // explicit SKIP instead of a red gate. The slope above was still asserted.
    if !loadfence.RequireCalm(t, "retained-heap total and process-RSS ceiling") {
        return
    }

    // Total retained growth over the window, against the unchanged spec number.
    retainedTotal := int64(last.heap) - int64(baseHeap)
    t.Logf("retained heap growth over %d events = %d B (budget %d B)",
        eventWindow, retainedTotal, specBudgetBytes)
    if retainedTotal > specBudgetBytes {
        t.Fatalf("retained heap growth over %d events = %d B, budget is %d B",
            eventWindow, retainedTotal, specBudgetBytes)
    }

    // Coarse process-RSS delta against an explicit ceiling.
    rssDelta := int64(last.rss) - int64(baseRSS)
    if rssDelta < 0 {
        rssDelta = 0
    }
    t.Logf("process RSS delta over %d events = %d B (ceiling %d B)",
        eventWindow, rssDelta, rssCeilingBytes)
    if rssDelta > rssCeilingBytes {
        t.Fatalf("process RSS delta %d B > ceiling %d B", rssDelta, rssCeilingBytes)
    }
}

The production contract this verifies is that Sensor.Handle keeps no per-event state. Any bounded aggregator (counters, fixed sketches) satisfies it; a slice/map keyed by event is a regression.

Verification

Environment: Go 1.26, Linux, 16 cores, host load ~8.4 (0.53/core).

gofmt -l .        # expect no output
go vet ./...      # expect no output
go build ./...    # expect no output
go test ./internal/sensors/ -run TestRSSContributionIsBounded -count=5 -v

Observed (5/5 reps PASS). Per-rep logged series, e.g.:

events=     0 heap=   407296 heap_inuse=   819200 rss=  4616192
events= 20000 heap=   376848 heap_inuse=   786432 rss=  5259264
tail: events=12000..20000 growth=256 B slope=64 B/batch scaled_slope=640 B (budget 12582912 B)
retained heap growth over 20000 events = -30448 B (budget 12582912 B)
process RSS delta over 20000 events = 643072 B (ceiling 25165824 B)
--- PASS: TestRSSContributionIsBounded

Across the 5 reps: scaled tail slope 640 B (budget 12 582 912 B), retained growth -41 376 .. -30 448 B, RSS delta 0 .. 643 072 B vs the 25 165 824 B ceiling. ok .../internal/sensors.

Oversubscription path (forced with a tiny threshold, proving an explicit SKIP rather than a red gate — note the slope is still evaluated first):

LOADFENCE_MAX_PER_CPU=0.0001 go test ./internal/sensors/ -run TestRSSContributionIsBounded -v
# loadfence.go: SKIP: retained-heap total and process-RSS ceiling not asserted:
#   host load1=12.70/cpu=16 (limit 0.00/cpu) is oversubscribed
# --- SKIP: TestRSSContributionIsBounded
# PASS

Falsification (run this before trusting any rewrite)

A gate that cannot be made to fail by a real retention regression is a phantom gate. Mutate the production path to retain every event:

// internal/sensors/sensors.go
var retained []Event // MUTATION

func (s *Sensor) Handle(e Event) {
    retained = append(retained, e) // MUTATION: per-event retention
    ...
}

Re-run the gate. It must fail loudly on the hard slope assertion:

tail: events=12000..20000 growth=15262208 B slope=3815552 B/batch
      scaled_slope=38155520 B (budget 12582912 B)
--- FAIL: TestRSSContributionIsBounded

The scaled slope (≈38 MB) exceeds the 12 MiB budget even though the transient looks like normal noise. Revert the mutation and confirm green again.

General rule

For any host-measured budget gate:

  1. Measure the quantity the budget is actually about — retained state for a memory budget, undisturbed cost for a timing budget (see go-testing-load-flake).
  2. Exclude one-time startup with a warmup.
  3. Prefer a slope over an absolute delta when the defect of interest is per-unit.
  4. Keep a coarse host-level ceiling as a second net.
  5. Fence or SKIP explicitly on an oversubscribed host instead of reporting red.
  6. Prove the gate still fails under a deliberate mutation.

Evidence & signatures

# Evidence
- Problem class: go-host-measured-memory-gate-fragile
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-19T07:12:31.076Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a Go perf gate that asserts a memory budget fails on 3 of 4 environments. internal/sensors TestRSSContributionIsBounded read /proc/self/statm twice (single sample, one runtime.GC() before each) around a 20000-event loop and asserted the RSS delta against a spec budget of 12 MiB: 13455360 B on the reference box, 17825792 B on a CI bunker, 15732736 B on a native cell, all FAIL; one run under HTTP_PROXY=<ip-address>:9 PASSED. Three fails and one pass from identical code is the tell that the measurement, not the code, is at fault.\n\nROOT CAUSE: a raw process-RSS delta is not the subsystem's contribution. Go grows the heap arena in chunks and the scavenger returns pages on its own schedule, so RSS moves by several MB even when the code retains almost nothing; the delta also carries the whole test binary's GC timing and any host/GOGC differences between environments. The old gate was therefore measuring runtime bookkeeping and calling it subsystem state.\n\nFIX (kept the spec budget, replaced the measurement): (1) retained heap - runtime.ReadMemStats after an explicit runtime.GC() with debug.SetGCPercent(100) pinned for the window and restored via defer; (2) a 4000-event warmup, then sample in 2000-event batches so one-time startup growth is outside the window; (3) a STEADY-STATE TAIL SLOPE assertion: the retained growth of the last 4 batches, scaled to the spec's 20k-event window, must stay under the budget - a real per-event retention regression is linear in event count and fails even when hidden behind a large transient, while runtime noise does not scale with events (this assertion is a hard failure, deliberately not load-fenced); (4) total retained growth over the window asserted against the unchanged spec number, and a coarse process-RSS delta asserted against an explicit ceiling (spec budget + measured arena/scavenger slack) - both routed through the repo's loadfence helper so an oversubscribed box reports an explicit SKIP instead of a red gate; (5) every sample (events, heap, heap_inuse, rss), the tail slope and the RSS delta are logged so the verdict is auditable on any box.\n\nVERIFICATION: 5/5 reps of -count=5 PASS with margin (retained heap growth 2.36-2.57 MB vs the 12 MiB budget; scaled tail slope -3.5 KB..256 KB vs 12 MiB; RSS delta 4.15-6.19 MB vs the 24 MiB ceiling; host load 4.26-5.80), whole package green in 47.9 s, build/vet/gofmt clean. FALSIFICATION, run twice independently (the worker's mutation and the foreman's own): retaining every event in a package-level slice makes the gate fail loudly - scaled tail slope 32137660 B vs the 12582912 budget, retained growth 35470240 B, RSS delta 73793536 B. A gate that cannot be made to fail by a real retention regression is a phantom gate, so run this mutation before trusting any rewrite of a memory budget test.\n\nGENERAL RULE: for any host-measured budget gate, (a) measure the quantity the budget is about (retained state for a memory budget, undisturbed cost for a timing budget - see the adjacent class go-testing-load-flake for the timing estimators), (b) exclude one-time startup via warmup, (c) prefer a SLOPE over an absolute delta when the defect of interest is per-unit retention, (d) keep a coarse host-level ceiling as a second net, (e) fence or skip explicitly on an oversubscribed host instead of reporting red, and (f) prove the gate still fails under a deliberate mutation.", "environment": "Go 1.x on Linux; measured on a 16-core control host at load_avg 4.3-5.8, a CI bunker and a native cell", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-host-measured-memory-gate-fragile", "provider": "openrouter", "solved_at": "2026-09-19T07:12:31.077Z", "version": ""}

Answer 2

I reproduced the fix as a self-contained module under /tmp/host-gate, verified it, and wrote the full solution to ~/SOLUTION.md. Summary of what I did and found:


Fix: TestRSSContributionIsBounded is a host-measurement gate, not a subsystem gate

Class: go-host-measured-memory-gate-fragile Repo / file: totalwindupflightsystems/trouble, internal/sensors/perf_test.go Related class: go-testing-load-flake

Symptom

The gate read /proc/self/statm twice around a 20 000-event loop (one runtime.GC() before each sample) and asserted the RSS delta against a 12 MiB spec budget:

sensors contributed 14282752 bytes of RSS (13455360 -> 27738112), budget is 12582912

Identical code failed on 3 of 4 boxes (reference 13 455 360 B, CI bunker 17 825 792 B, native cell 15 732 736 B) and passed once under HTTP_PROXY=<ip-address>:9. Three fails and one pass from unchanged code is the tell that the measurement, not the subsystem, is at fault.

Root cause

A raw process-RSS delta is not the subsystem's contribution:

The old gate was measuring runtime bookkeeping and calling it subsystem state.

Fix

Keep the spec budget (12 << 20 B) and replace the measurement:

  1. Measure retained state, not RSS. runtime.ReadMemStats after an explicit runtime.GC(), with debug.SetGCPercent(100) pinned for the window and restored with defer.
  2. Warm up 4 000 events, then sample in 2 000-event batches so one-time startup growth lands outside the window.
  3. Steady-state tail slope (hard failure, deliberately not load-fenced). Retained growth across the last 4 batches, scaled to the spec's 20 k-event window, must stay under budget. A real per-event retention regression is linear in event count and fails even behind a large transient; runtime noise does not scale with events.
  4. Two fenced nets. Total retained growth over the window vs. the unchanged spec number, and a coarse process-RSS delta vs. an explicit ceiling (spec budget + measured arena/scavenger slack). Both go through the repo's loadfence helper so an oversubscribed box reports an explicit SKIP instead of a red gate.
  5. Log every sample (events, heap, heap_inuse, rss), the tail slope and the RSS delta, so the verdict is auditable on any box.

internal/loadfence/loadfence.go

Use the repo's existing helper if present; this is a self-contained equivalent.

// Package loadfence decides whether a host is calm enough for a
// host-measured budget assertion. When it is not, host-level checks are
// reported as an explicit SKIP rather than a red gate.
package loadfence

import (
    "os"
    "runtime"
    "strconv"
    "strings"
    "testing"
)

const defaultMaxLoadPerCPU = 0.9

func Load1() (float64, bool) {
    b, err := os.ReadFile("/proc/loadavg")
    if err != nil {
        return 0, false
    }
    f := strings.Fields(string(b))
    if len(f) == 0 {
        return 0, false
    }
    v, err := strconv.ParseFloat(f[0], 64)
    if err != nil {
        return 0, false
    }
    return v, true
}

func maxLoadPerCPU() float64 {
    if s := os.Getenv("LOADFENCE_MAX_PER_CPU"); s != "" {
        if v, err := strconv.ParseFloat(s, 64); err == nil && v > 0 {
            return v
        }
    }
    return defaultMaxLoadPerCPU
}

func Oversubscribed() bool {
    load, ok := Load1()
    if !ok {
        return false // no load data: do not stand in the way of the gate
    }
    return load > maxLoadPerCPU()*float64(runtime.NumCPU())
}

// RequireCalm skips the calling test with an explicit message when the host
// is oversubscribed, and returns true when the host is calm enough to
// continue. Keep true per-unit assertions (e.g. a slope) outside the fence.
func RequireCalm(t testing.TB, what string) bool {
    if !Oversubscribed() {
        return true
    }
    load, _ := Load1()
    t.Skipf("SKIP: %s not asserted: host load1=%.2f/cpu=%d (limit %.2f/cpu) is oversubscribed",
        what, load, runtime.NumCPU(), maxLoadPerCPU())
    return false
}

internal/sensors/perf_test.go (replacement test)

package sensors

import (
    "os"
    "runtime"
    "runtime/debug"
    "strconv"
    "strings"
    "testing"

    "example.com/trouble/internal/loadfence" // repo path
)

const (
    // specBudgetBytes is the published budget for the 20k-event window. It is
    // deliberately unchanged: the fix replaces the measurement, not the spec.
    specBudgetBytes = 12 << 20 // 12582912 B

    // rssCeilingBytes is a coarse host-level second net: the spec budget plus
    // measured arena/scavenger slack. It is not the subsystem's contribution.
    rssCeilingBytes = 2 * specBudgetBytes // 24 MiB

    eventWindow  = 20000
    batchSize    = 2000
    warmupEvents = 4000
    tailBatches  = 4
)

type sample struct {
    events    int
    heap      uint64 // HeapAlloc after GC: the retained quantity
    heapInuse uint64 // HeapInuse after GC
    rss       uint64 // process RSS from /proc/self/statm
}

func readRSS(t *testing.T) uint64 {
    t.Helper()
    b, err := os.ReadFile("/proc/self/statm")
    if err != nil {
        t.Fatalf("read /proc/self/statm: %v", err)
    }
    f := strings.Fields(string(b))
    if len(f) < 2 {
        t.Fatalf("unexpected /proc/self/statm: %q", b)
    }
    pages, err := strconv.ParseUint(f[1], 10, 64)
    if err != nil {
        t.Fatalf("parse resident pages: %v", err)
    }
    return pages * uint64(os.Getpagesize())
}

// retainedHeap returns HeapAlloc and HeapInuse after an explicit collection.
// ReadMemStats after GC reports live (retained) bytes, not the arena high
// water mark.
func retainedHeap() (heap, heapInuse uint64) {
    runtime.GC()
    var m runtime.MemStats
    runtime.ReadMemStats(&m)
    return m.HeapAlloc, m.HeapInuse
}

func makeEvent(i int) Event {
    return Event{ID: uint64(i), Value: float64(i % 97)}
}

// TestRSSContributionIsBounded asserts the subsystem's retained-state contract
// over a 20k-event window. It measures retained heap (not raw process RSS),
// excludes one-time startup with a warmup, and prefers a steady-state tail
// slope because a real per-event retention regression is linear in the event
// count.
func TestRSSContributionIsBounded(t *testing.T) {
    // Pin GC behaviour for the measurement window and restore it afterwards.
    oldGC := debug.SetGCPercent(100)
    defer debug.SetGCPercent(oldGC)

    s := New()

    // Warmup: pay one-time startup growth (lazily created maps, caches,
    // arenas) outside the measured window.
    for i := 0; i < warmupEvents; i++ {
        s.Handle(makeEvent(i))
    }

    baseHeap, baseInuse := retainedHeap()
    baseRSS := readRSS(t)
    samples := []sample{{events: 0, heap: baseHeap, heapInuse: baseInuse, rss: baseRSS}}

    handled := 0
    for b := 0; b < eventWindow/batchSize; b++ {
        for i := 0; i < batchSize; i++ {
            s.Handle(makeEvent(handled + i))
        }
        handled += batchSize
        h, hi := retainedHeap()
        samples = append(samples, sample{events: handled, heap: h, heapInuse: hi, rss: readRSS(t)})
    }

    // Log every sample so the verdict is auditable on any box.
    for _, sm := range samples {
        t.Logf("events=%6d heap=%9d heap_inuse=%9d rss=%9d", sm.events, sm.heap, sm.heapInuse, sm.rss)
    }

    // Steady-state tail slope: retained growth across the last tailBatches
    // batches, scaled to the spec's full event window.
    first := samples[len(samples)-1-tailBatches]
    last := samples[len(samples)-1]
    growthTail := int64(last.heap) - int64(first.heap)
    slopePerBatch := growthTail / int64(tailBatches)
    slopeScale := int64(eventWindow / batchSize)
    scaledTail := slopePerBatch * slopeScale
    t.Logf("tail: events=%d..%d growth=%d B slope=%d B/batch scaled_slope=%d B (budget %d B)",
        first.events, last.events, growthTail, slopePerBatch, scaledTail, specBudgetBytes)

    // (a) Hard, not load-fenced: a real per-event retention regression scales
    // with the event count and must fail even when hidden behind a large
    // transient.
    if scaledTail > specBudgetBytes {
        t.Fatalf("steady-state tail slope scaled to %d events = %d B, budget is %d B",
            eventWindow, scaledTail, specBudgetBytes)
    }

    // (b) Host-level checks are fenced: on an oversubscribed box report an
    // explicit SKIP instead of a red gate. The slope above was still asserted.
    if !loadfence.RequireCalm(t, "retained-heap total and process-RSS ceiling") {
        return
    }

    // Total retained growth over the window, against the unchanged spec number.
    retainedTotal := int64(last.heap) - int64(baseHeap)
    t.Logf("retained heap growth over %d events = %d B (budget %d B)",
        eventWindow, retainedTotal, specBudgetBytes)
    if retainedTotal > specBudgetBytes {
        t.Fatalf("retained heap growth over %d events = %d B, budget is %d B",
            eventWindow, retainedTotal, specBudgetBytes)
    }

    // Coarse process-RSS delta against an explicit ceiling.
    rssDelta := int64(last.rss) - int64(baseRSS)
    if rssDelta < 0 {
        rssDelta = 0
    }
    t.Logf("process RSS delta over %d events = %d B (ceiling %d B)",
        eventWindow, rssDelta, rssCeilingBytes)
    if rssDelta > rssCeilingBytes {
        t.Fatalf("process RSS delta %d B > ceiling %d B", rssDelta, rssCeilingBytes)
    }
}

The production contract this verifies is that Sensor.Handle keeps no per-event state. Any bounded aggregator (counters, fixed sketches) satisfies it; a slice/map keyed by event is a regression.

Verification

Environment: Go 1.26, Linux, 16 cores, host load ~8.4 (0.53/core).

gofmt -l .        # expect no output
go vet ./...      # expect no output
go build ./...    # expect no output
go test ./internal/sensors/ -run TestRSSContributionIsBounded -count=5 -v

Observed (5/5 reps PASS). Per-rep logged series, e.g.:

events=     0 heap=   407296 heap_inuse=   819200 rss=  4616192
events= 20000 heap=   376848 heap_inuse=   786432 rss=  5259264
tail: events=12000..20000 growth=256 B slope=64 B/batch scaled_slope=640 B (budget 12582912 B)
retained heap growth over 20000 events = -30448 B (budget 12582912 B)
process RSS delta over 20000 events = 643072 B (ceiling 25165824 B)
--- PASS: TestRSSContributionIsBounded

Across the 5 reps: scaled tail slope 640 B (budget 12 582 912 B), retained growth -41 376 .. -30 448 B, RSS delta 0 .. 643 072 B vs the 25 165 824 B ceiling. ok .../internal/sensors.

Oversubscription path (forced with a tiny threshold, proving an explicit SKIP rather than a red gate — note the slope is still evaluated first):

LOADFENCE_MAX_PER_CPU=0.0001 go test ./internal/sensors/ -run TestRSSContributionIsBounded -v
# loadfence.go: SKIP: retained-heap total and process-RSS ceiling not asserted:
#   host load1=12.70/cpu=16 (limit 0.00/cpu) is oversubscribed
# --- SKIP: TestRSSContributionIsBounded
# PASS

Falsification (run this before trusting any rewrite)

A gate that cannot be made to fail by a real retention regression is a phantom gate. Mutate the production path to retain every event:

// internal/sensors/sensors.go
var retained []Event // MUTATION

func (s *Sensor) Handle(e Event) {
    retained = append(retained, e) // MUTATION: per-event retention
    ...
}

Re-run the gate. It must fail loudly on the hard slope assertion:

tail: events=12000..20000 growth=15262208 B slope=3815552 B/batch
      scaled_slope=38155520 B (budget 12582912 B)
--- FAIL: TestRSSContributionIsBounded

The scaled slope (≈38 MB) exceeds the 12 MiB budget even though the transient looks like normal noise. Revert the mutation and confirm green again.

General rule

For any host-measured budget gate:

  1. Measure the quantity the budget is actually about — retained state for a memory budget, undisturbed cost for a timing budget (see go-testing-load-flake).
  2. Exclude one-time startup with a warmup.
  3. Prefer a slope over an absolute delta when the defect of interest is per-unit.
  4. Keep a coarse host-level ceiling as a second net.
  5. Fence or SKIP explicitly on an oversubscribed host instead of reporting red.
  6. Prove the gate still fails under a deliberate mutation.

Evidence & signatures

# Evidence
- Problem class: go-host-measured-memory-gate-fragile
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-19T07:12:31.076Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a Go perf gate that asserts a memory budget fails on 3 of 4 environments. internal/sensors TestRSSContributionIsBounded read /proc/self/statm twice (single sample, one runtime.GC() before each) around a 20000-event loop and asserted the RSS delta against a spec budget of 12 MiB: 13455360 B on the reference box, 17825792 B on a CI bunker, 15732736 B on a native cell, all FAIL; one run under HTTP_PROXY=<ip-address>:9 PASSED. Three fails and one pass from identical code is the tell that the measurement, not the code, is at fault.\n\nROOT CAUSE: a raw process-RSS delta is not the subsystem's contribution. Go grows the heap arena in chunks and the scavenger returns pages on its own schedule, so RSS moves by several MB even when the code retains almost nothing; the delta also carries the whole test binary's GC timing and any host/GOGC differences between environments. The old gate was therefore measuring runtime bookkeeping and calling it subsystem state.\n\nFIX (kept the spec budget, replaced the measurement): (1) retained heap - runtime.ReadMemStats after an explicit runtime.GC() with debug.SetGCPercent(100) pinned for the window and restored via defer; (2) a 4000-event warmup, then sample in 2000-event batches so one-time startup growth is outside the window; (3) a STEADY-STATE TAIL SLOPE assertion: the retained growth of the last 4 batches, scaled to the spec's 20k-event window, must stay under the budget - a real per-event retention regression is linear in event count and fails even when hidden behind a large transient, while runtime noise does not scale with events (this assertion is a hard failure, deliberately not load-fenced); (4) total retained growth over the window asserted against the unchanged spec number, and a coarse process-RSS delta asserted against an explicit ceiling (spec budget + measured arena/scavenger slack) - both routed through the repo's loadfence helper so an oversubscribed box reports an explicit SKIP instead of a red gate; (5) every sample (events, heap, heap_inuse, rss), the tail slope and the RSS delta are logged so the verdict is auditable on any box.\n\nVERIFICATION: 5/5 reps of -count=5 PASS with margin (retained heap growth 2.36-2.57 MB vs the 12 MiB budget; scaled tail slope -3.5 KB..256 KB vs 12 MiB; RSS delta 4.15-6.19 MB vs the 24 MiB ceiling; host load 4.26-5.80), whole package green in 47.9 s, build/vet/gofmt clean. FALSIFICATION, run twice independently (the worker's mutation and the foreman's own): retaining every event in a package-level slice makes the gate fail loudly - scaled tail slope 32137660 B vs the 12582912 budget, retained growth 35470240 B, RSS delta 73793536 B. A gate that cannot be made to fail by a real retention regression is a phantom gate, so run this mutation before trusting any rewrite of a memory budget test.\n\nGENERAL RULE: for any host-measured budget gate, (a) measure the quantity the budget is about (retained state for a memory budget, undisturbed cost for a timing budget - see the adjacent class go-testing-load-flake for the timing estimators), (b) exclude one-time startup via warmup, (c) prefer a SLOPE over an absolute delta when the defect of interest is per-unit retention, (d) keep a coarse host-level ceiling as a second net, (e) fence or skip explicitly on an oversubscribed host instead of reporting red, and (f) prove the gate still fails under a deliberate mutation.", "environment": "Go 1.x on Linux; measured on a 16-core control host at load_avg 4.3-5.8, a CI bunker and a native cell", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-host-measured-memory-gate-fragile", "provider": "openrouter", "solved_at": "2026-09-19T07:12:31.077Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog