Class: go-host-measured-memory-gate-fragile
I reproduced the fix as a self-contained module under /tmp/host-gate, verified it, and wrote the full solution to ~/SOLUTION.md. Summary of what I did and found:
GOGC differences — not the subsystem's retained state. Single-sample, both samples straddling a 20k loop.ReadMemStats after explicit GC, SetGCPercent(100) pinned + defer restore), 4k warmup, 2k batches, hard steady-state tail-slope assertion (not load-fenced), plus load-fenced total-retained and RSS-ceiling nets, full per-sample logging.-count=5 all PASS; scaled tail slope 640 B vs 12 582 912 B; retained growth -41 376..-30 448 B; RSS delta 0..643 072 B vs 25 165 824 B ceiling; gofmt/vet/build clean. Forced-oversubscription run reports SKIP (after the slope assertion).38 155 520 B > 12 582 912 B.TestRSSContributionIsBounded is a host-measurement gate, not a subsystem gateClass: go-host-measured-memory-gate-fragile
Repo / file: totalwindupflightsystems/trouble, internal/sensors/perf_test.go
Related class: go-testing-load-flake
The gate read /proc/self/statm twice around a 20 000-event loop (one runtime.GC() before each sample) and asserted the RSS delta against a 12 MiB spec budget:
sensors contributed 14282752 bytes of RSS (13455360 -> 27738112), budget is 12582912
Identical code failed on 3 of 4 boxes (reference 13 455 360 B, CI bunker 17 825 792 B, native cell 15 732 736 B) and passed once under HTTP_PROXY=<ip-address>:9. Three fails and one pass from unchanged code is the tell that the measurement, not the subsystem, is at fault.
A raw process-RSS delta is not the subsystem's contribution:
GOGC differences between environments.The old gate was measuring runtime bookkeeping and calling it subsystem state.
Keep the spec budget (12 << 20 B) and replace the measurement:
runtime.ReadMemStats after an explicit runtime.GC(), with debug.SetGCPercent(100) pinned for the window and restored with defer.loadfence helper so an oversubscribed box reports an explicit SKIP instead of a red gate.events, heap, heap_inuse, rss), the tail slope and the RSS delta, so the verdict is auditable on any box.internal/loadfence/loadfence.goUse the repo's existing helper if present; this is a self-contained equivalent.
// Package loadfence decides whether a host is calm enough for a
// host-measured budget assertion. When it is not, host-level checks are
// reported as an explicit SKIP rather than a red gate.
package loadfence
import (
"os"
"runtime"
"strconv"
"strings"
"testing"
)
const defaultMaxLoadPerCPU = 0.9
func Load1() (float64, bool) {
b, err := os.ReadFile("/proc/loadavg")
if err != nil {
return 0, false
}
f := strings.Fields(string(b))
if len(f) == 0 {
return 0, false
}
v, err := strconv.ParseFloat(f[0], 64)
if err != nil {
return 0, false
}
return v, true
}
func maxLoadPerCPU() float64 {
if s := os.Getenv("LOADFENCE_MAX_PER_CPU"); s != "" {
if v, err := strconv.ParseFloat(s, 64); err == nil && v > 0 {
return v
}
}
return defaultMaxLoadPerCPU
}
func Oversubscribed() bool {
load, ok := Load1()
if !ok {
return false // no load data: do not stand in the way of the gate
}
return load > maxLoadPerCPU()*float64(runtime.NumCPU())
}
// RequireCalm skips the calling test with an explicit message when the host
// is oversubscribed, and returns true when the host is calm enough to
// continue. Keep true per-unit assertions (e.g. a slope) outside the fence.
func RequireCalm(t testing.TB, what string) bool {
if !Oversubscribed() {
return true
}
load, _ := Load1()
t.Skipf("SKIP: %s not asserted: host load1=%.2f/cpu=%d (limit %.2f/cpu) is oversubscribed",
what, load, runtime.NumCPU(), maxLoadPerCPU())
return false
}
internal/sensors/perf_test.go (replacement test)package sensors
import (
"os"
"runtime"
"runtime/debug"
"strconv"
"strings"
"testing"
"example.com/trouble/internal/loadfence" // repo path
)
const (
// specBudgetBytes is the published budget for the 20k-event window. It is
// deliberately unchanged: the fix replaces the measurement, not the spec.
specBudgetBytes = 12 << 20 // 12582912 B
// rssCeilingBytes is a coarse host-level second net: the spec budget plus
// measured arena/scavenger slack. It is not the subsystem's contribution.
rssCeilingBytes = 2 * specBudgetBytes // 24 MiB
eventWindow = 20000
batchSize = 2000
warmupEvents = 4000
tailBatches = 4
)
type sample struct {
events int
heap uint64 // HeapAlloc after GC: the retained quantity
heapInuse uint64 // HeapInuse after GC
rss uint64 // process RSS from /proc/self/statm
}
func readRSS(t *testing.T) uint64 {
t.Helper()
b, err := os.ReadFile("/proc/self/statm")
if err != nil {
t.Fatalf("read /proc/self/statm: %v", err)
}
f := strings.Fields(string(b))
if len(f) < 2 {
t.Fatalf("unexpected /proc/self/statm: %q", b)
}
pages, err := strconv.ParseUint(f[1], 10, 64)
if err != nil {
t.Fatalf("parse resident pages: %v", err)
}
return pages * uint64(os.Getpagesize())
}
// retainedHeap returns HeapAlloc and HeapInuse after an explicit collection.
// ReadMemStats after GC reports live (retained) bytes, not the arena high
// water mark.
func retainedHeap() (heap, heapInuse uint64) {
runtime.GC()
var m runtime.MemStats
runtime.ReadMemStats(&m)
return m.HeapAlloc, m.HeapInuse
}
func makeEvent(i int) Event {
return Event{ID: uint64(i), Value: float64(i % 97)}
}
// TestRSSContributionIsBounded asserts the subsystem's retained-state contract
// over a 20k-event window. It measures retained heap (not raw process RSS),
// excludes one-time startup with a warmup, and prefers a steady-state tail
// slope because a real per-event retention regression is linear in the event
// count.
func TestRSSContributionIsBounded(t *testing.T) {
// Pin GC behaviour for the measurement window and restore it afterwards.
oldGC := debug.SetGCPercent(100)
defer debug.SetGCPercent(oldGC)
s := New()
// Warmup: pay one-time startup growth (lazily created maps, caches,
// arenas) outside the measured window.
for i := 0; i < warmupEvents; i++ {
s.Handle(makeEvent(i))
}
baseHeap, baseInuse := retainedHeap()
baseRSS := readRSS(t)
samples := []sample{{events: 0, heap: baseHeap, heapInuse: baseInuse, rss: baseRSS}}
handled := 0
for b := 0; b < eventWindow/batchSize; b++ {
for i := 0; i < batchSize; i++ {
s.Handle(makeEvent(handled + i))
}
handled += batchSize
h, hi := retainedHeap()
samples = append(samples, sample{events: handled, heap: h, heapInuse: hi, rss: readRSS(t)})
}
// Log every sample so the verdict is auditable on any box.
for _, sm := range samples {
t.Logf("events=%6d heap=%9d heap_inuse=%9d rss=%9d", sm.events, sm.heap, sm.heapInuse, sm.rss)
}
// Steady-state tail slope: retained growth across the last tailBatches
// batches, scaled to the spec's full event window.
first := samples[len(samples)-1-tailBatches]
last := samples[len(samples)-1]
growthTail := int64(last.heap) - int64(first.heap)
slopePerBatch := growthTail / int64(tailBatches)
slopeScale := int64(eventWindow / batchSize)
scaledTail := slopePerBatch * slopeScale
t.Logf("tail: events=%d..%d growth=%d B slope=%d B/batch scaled_slope=%d B (budget %d B)",
first.events, last.events, growthTail, slopePerBatch, scaledTail, specBudgetBytes)
// (a) Hard, not load-fenced: a real per-event retention regression scales
// with the event count and must fail even when hidden behind a large
// transient.
if scaledTail > specBudgetBytes {
t.Fatalf("steady-state tail slope scaled to %d events = %d B, budget is %d B",
eventWindow, scaledTail, specBudgetBytes)
}
// (b) Host-level checks are fenced: on an oversubscribed box report an
// explicit SKIP instead of a red gate. The slope above was still asserted.
if !loadfence.RequireCalm(t, "retained-heap total and process-RSS ceiling") {
return
}
// Total retained growth over the window, against the unchanged spec number.
retainedTotal := int64(last.heap) - int64(baseHeap)
t.Logf("retained heap growth over %d events = %d B (budget %d B)",
eventWindow, retainedTotal, specBudgetBytes)
if retainedTotal > specBudgetBytes {
t.Fatalf("retained heap growth over %d events = %d B, budget is %d B",
eventWindow, retainedTotal, specBudgetBytes)
}
// Coarse process-RSS delta against an explicit ceiling.
rssDelta := int64(last.rss) - int64(baseRSS)
if rssDelta < 0 {
rssDelta = 0
}
t.Logf("process RSS delta over %d events = %d B (ceiling %d B)",
eventWindow, rssDelta, rssCeilingBytes)
if rssDelta > rssCeilingBytes {
t.Fatalf("process RSS delta %d B > ceiling %d B", rssDelta, rssCeilingBytes)
}
}
The production contract this verifies is that Sensor.Handle keeps no per-event state. Any bounded aggregator (counters, fixed sketches) satisfies it; a slice/map keyed by event is a regression.
Environment: Go 1.26, Linux, 16 cores, host load ~8.4 (0.53/core).
gofmt -l . # expect no output
go vet ./... # expect no output
go build ./... # expect no output
go test ./internal/sensors/ -run TestRSSContributionIsBounded -count=5 -v
Observed (5/5 reps PASS). Per-rep logged series, e.g.:
events= 0 heap= 407296 heap_inuse= 819200 rss= 4616192
events= 20000 heap= 376848 heap_inuse= 786432 rss= 5259264
tail: events=12000..20000 growth=256 B slope=64 B/batch scaled_slope=640 B (budget 12582912 B)
retained heap growth over 20000 events = -30448 B (budget 12582912 B)
process RSS delta over 20000 events = 643072 B (ceiling 25165824 B)
--- PASS: TestRSSContributionIsBounded
Across the 5 reps: scaled tail slope 640 B (budget 12 582 912 B), retained growth -41 376 .. -30 448 B, RSS delta 0 .. 643 072 B vs the 25 165 824 B ceiling. ok .../internal/sensors.
Oversubscription path (forced with a tiny threshold, proving an explicit SKIP rather than a red gate — note the slope is still evaluated first):
LOADFENCE_MAX_PER_CPU=0.0001 go test ./internal/sensors/ -run TestRSSContributionIsBounded -v
# loadfence.go: SKIP: retained-heap total and process-RSS ceiling not asserted:
# host load1=12.70/cpu=16 (limit 0.00/cpu) is oversubscribed
# --- SKIP: TestRSSContributionIsBounded
# PASS
A gate that cannot be made to fail by a real retention regression is a phantom gate. Mutate the production path to retain every event:
// internal/sensors/sensors.go
var retained []Event // MUTATION
func (s *Sensor) Handle(e Event) {
retained = append(retained, e) // MUTATION: per-event retention
...
}
Re-run the gate. It must fail loudly on the hard slope assertion:
tail: events=12000..20000 growth=15262208 B slope=3815552 B/batch
scaled_slope=38155520 B (budget 12582912 B)
--- FAIL: TestRSSContributionIsBounded
The scaled slope (≈38 MB) exceeds the 12 MiB budget even though the transient looks like normal noise. Revert the mutation and confirm green again.
For any host-measured budget gate:
go-testing-load-flake).# Evidence - Problem class: go-host-measured-memory-gate-fragile - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T07:12:31.076Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a Go perf gate that asserts a memory budget fails on 3 of 4 environments. internal/sensors TestRSSContributionIsBounded read /proc/self/statm twice (single sample, one runtime.GC() before each) around a 20000-event loop and asserted the RSS delta against a spec budget of 12 MiB: 13455360 B on the reference box, 17825792 B on a CI bunker, 15732736 B on a native cell, all FAIL; one run under HTTP_PROXY=<ip-address>:9 PASSED. Three fails and one pass from identical code is the tell that the measurement, not the code, is at fault.\n\nROOT CAUSE: a raw process-RSS delta is not the subsystem's contribution. Go grows the heap arena in chunks and the scavenger returns pages on its own schedule, so RSS moves by several MB even when the code retains almost nothing; the delta also carries the whole test binary's GC timing and any host/GOGC differences between environments. The old gate was therefore measuring runtime bookkeeping and calling it subsystem state.\n\nFIX (kept the spec budget, replaced the measurement): (1) retained heap - runtime.ReadMemStats after an explicit runtime.GC() with debug.SetGCPercent(100) pinned for the window and restored via defer; (2) a 4000-event warmup, then sample in 2000-event batches so one-time startup growth is outside the window; (3) a STEADY-STATE TAIL SLOPE assertion: the retained growth of the last 4 batches, scaled to the spec's 20k-event window, must stay under the budget - a real per-event retention regression is linear in event count and fails even when hidden behind a large transient, while runtime noise does not scale with events (this assertion is a hard failure, deliberately not load-fenced); (4) total retained growth over the window asserted against the unchanged spec number, and a coarse process-RSS delta asserted against an explicit ceiling (spec budget + measured arena/scavenger slack) - both routed through the repo's loadfence helper so an oversubscribed box reports an explicit SKIP instead of a red gate; (5) every sample (events, heap, heap_inuse, rss), the tail slope and the RSS delta are logged so the verdict is auditable on any box.\n\nVERIFICATION: 5/5 reps of -count=5 PASS with margin (retained heap growth 2.36-2.57 MB vs the 12 MiB budget; scaled tail slope -3.5 KB..256 KB vs 12 MiB; RSS delta 4.15-6.19 MB vs the 24 MiB ceiling; host load 4.26-5.80), whole package green in 47.9 s, build/vet/gofmt clean. FALSIFICATION, run twice independently (the worker's mutation and the foreman's own): retaining every event in a package-level slice makes the gate fail loudly - scaled tail slope 32137660 B vs the 12582912 budget, retained growth 35470240 B, RSS delta 73793536 B. A gate that cannot be made to fail by a real retention regression is a phantom gate, so run this mutation before trusting any rewrite of a memory budget test.\n\nGENERAL RULE: for any host-measured budget gate, (a) measure the quantity the budget is about (retained state for a memory budget, undisturbed cost for a timing budget - see the adjacent class go-testing-load-flake for the timing estimators), (b) exclude one-time startup via warmup, (c) prefer a SLOPE over an absolute delta when the defect of interest is per-unit retention, (d) keep a coarse host-level ceiling as a second net, (e) fence or skip explicitly on an oversubscribed host instead of reporting red, and (f) prove the gate still fails under a deliberate mutation.", "environment": "Go 1.x on Linux; measured on a 16-core control host at load_avg 4.3-5.8, a CI bunker and a native cell", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-host-measured-memory-gate-fragile", "provider": "openrouter", "solved_at": "2026-09-19T07:12:31.077Z", "version": ""}I reproduced the fix as a self-contained module under /tmp/host-gate, verified it, and wrote the full solution to ~/SOLUTION.md. Summary of what I did and found:
GOGC differences — not the subsystem's retained state. Single-sample, both samples straddling a 20k loop.ReadMemStats after explicit GC, SetGCPercent(100) pinned + defer restore), 4k warmup, 2k batches, hard steady-state tail-slope assertion (not load-fenced), plus load-fenced total-retained and RSS-ceiling nets, full per-sample logging.-count=5 all PASS; scaled tail slope 640 B vs 12 582 912 B; retained growth -41 376..-30 448 B; RSS delta 0..643 072 B vs 25 165 824 B ceiling; gofmt/vet/build clean. Forced-oversubscription run reports SKIP (after the slope assertion).38 155 520 B > 12 582 912 B.TestRSSContributionIsBounded is a host-measurement gate, not a subsystem gateClass: go-host-measured-memory-gate-fragile
Repo / file: totalwindupflightsystems/trouble, internal/sensors/perf_test.go
Related class: go-testing-load-flake
The gate read /proc/self/statm twice around a 20 000-event loop (one runtime.GC() before each sample) and asserted the RSS delta against a 12 MiB spec budget:
sensors contributed 14282752 bytes of RSS (13455360 -> 27738112), budget is 12582912
Identical code failed on 3 of 4 boxes (reference 13 455 360 B, CI bunker 17 825 792 B, native cell 15 732 736 B) and passed once under HTTP_PROXY=<ip-address>:9. Three fails and one pass from unchanged code is the tell that the measurement, not the subsystem, is at fault.
A raw process-RSS delta is not the subsystem's contribution:
GOGC differences between environments.The old gate was measuring runtime bookkeeping and calling it subsystem state.
Keep the spec budget (12 << 20 B) and replace the measurement:
runtime.ReadMemStats after an explicit runtime.GC(), with debug.SetGCPercent(100) pinned for the window and restored with defer.loadfence helper so an oversubscribed box reports an explicit SKIP instead of a red gate.events, heap, heap_inuse, rss), the tail slope and the RSS delta, so the verdict is auditable on any box.internal/loadfence/loadfence.goUse the repo's existing helper if present; this is a self-contained equivalent.
// Package loadfence decides whether a host is calm enough for a
// host-measured budget assertion. When it is not, host-level checks are
// reported as an explicit SKIP rather than a red gate.
package loadfence
import (
"os"
"runtime"
"strconv"
"strings"
"testing"
)
const defaultMaxLoadPerCPU = 0.9
func Load1() (float64, bool) {
b, err := os.ReadFile("/proc/loadavg")
if err != nil {
return 0, false
}
f := strings.Fields(string(b))
if len(f) == 0 {
return 0, false
}
v, err := strconv.ParseFloat(f[0], 64)
if err != nil {
return 0, false
}
return v, true
}
func maxLoadPerCPU() float64 {
if s := os.Getenv("LOADFENCE_MAX_PER_CPU"); s != "" {
if v, err := strconv.ParseFloat(s, 64); err == nil && v > 0 {
return v
}
}
return defaultMaxLoadPerCPU
}
func Oversubscribed() bool {
load, ok := Load1()
if !ok {
return false // no load data: do not stand in the way of the gate
}
return load > maxLoadPerCPU()*float64(runtime.NumCPU())
}
// RequireCalm skips the calling test with an explicit message when the host
// is oversubscribed, and returns true when the host is calm enough to
// continue. Keep true per-unit assertions (e.g. a slope) outside the fence.
func RequireCalm(t testing.TB, what string) bool {
if !Oversubscribed() {
return true
}
load, _ := Load1()
t.Skipf("SKIP: %s not asserted: host load1=%.2f/cpu=%d (limit %.2f/cpu) is oversubscribed",
what, load, runtime.NumCPU(), maxLoadPerCPU())
return false
}
internal/sensors/perf_test.go (replacement test)package sensors
import (
"os"
"runtime"
"runtime/debug"
"strconv"
"strings"
"testing"
"example.com/trouble/internal/loadfence" // repo path
)
const (
// specBudgetBytes is the published budget for the 20k-event window. It is
// deliberately unchanged: the fix replaces the measurement, not the spec.
specBudgetBytes = 12 << 20 // 12582912 B
// rssCeilingBytes is a coarse host-level second net: the spec budget plus
// measured arena/scavenger slack. It is not the subsystem's contribution.
rssCeilingBytes = 2 * specBudgetBytes // 24 MiB
eventWindow = 20000
batchSize = 2000
warmupEvents = 4000
tailBatches = 4
)
type sample struct {
events int
heap uint64 // HeapAlloc after GC: the retained quantity
heapInuse uint64 // HeapInuse after GC
rss uint64 // process RSS from /proc/self/statm
}
func readRSS(t *testing.T) uint64 {
t.Helper()
b, err := os.ReadFile("/proc/self/statm")
if err != nil {
t.Fatalf("read /proc/self/statm: %v", err)
}
f := strings.Fields(string(b))
if len(f) < 2 {
t.Fatalf("unexpected /proc/self/statm: %q", b)
}
pages, err := strconv.ParseUint(f[1], 10, 64)
if err != nil {
t.Fatalf("parse resident pages: %v", err)
}
return pages * uint64(os.Getpagesize())
}
// retainedHeap returns HeapAlloc and HeapInuse after an explicit collection.
// ReadMemStats after GC reports live (retained) bytes, not the arena high
// water mark.
func retainedHeap() (heap, heapInuse uint64) {
runtime.GC()
var m runtime.MemStats
runtime.ReadMemStats(&m)
return m.HeapAlloc, m.HeapInuse
}
func makeEvent(i int) Event {
return Event{ID: uint64(i), Value: float64(i % 97)}
}
// TestRSSContributionIsBounded asserts the subsystem's retained-state contract
// over a 20k-event window. It measures retained heap (not raw process RSS),
// excludes one-time startup with a warmup, and prefers a steady-state tail
// slope because a real per-event retention regression is linear in the event
// count.
func TestRSSContributionIsBounded(t *testing.T) {
// Pin GC behaviour for the measurement window and restore it afterwards.
oldGC := debug.SetGCPercent(100)
defer debug.SetGCPercent(oldGC)
s := New()
// Warmup: pay one-time startup growth (lazily created maps, caches,
// arenas) outside the measured window.
for i := 0; i < warmupEvents; i++ {
s.Handle(makeEvent(i))
}
baseHeap, baseInuse := retainedHeap()
baseRSS := readRSS(t)
samples := []sample{{events: 0, heap: baseHeap, heapInuse: baseInuse, rss: baseRSS}}
handled := 0
for b := 0; b < eventWindow/batchSize; b++ {
for i := 0; i < batchSize; i++ {
s.Handle(makeEvent(handled + i))
}
handled += batchSize
h, hi := retainedHeap()
samples = append(samples, sample{events: handled, heap: h, heapInuse: hi, rss: readRSS(t)})
}
// Log every sample so the verdict is auditable on any box.
for _, sm := range samples {
t.Logf("events=%6d heap=%9d heap_inuse=%9d rss=%9d", sm.events, sm.heap, sm.heapInuse, sm.rss)
}
// Steady-state tail slope: retained growth across the last tailBatches
// batches, scaled to the spec's full event window.
first := samples[len(samples)-1-tailBatches]
last := samples[len(samples)-1]
growthTail := int64(last.heap) - int64(first.heap)
slopePerBatch := growthTail / int64(tailBatches)
slopeScale := int64(eventWindow / batchSize)
scaledTail := slopePerBatch * slopeScale
t.Logf("tail: events=%d..%d growth=%d B slope=%d B/batch scaled_slope=%d B (budget %d B)",
first.events, last.events, growthTail, slopePerBatch, scaledTail, specBudgetBytes)
// (a) Hard, not load-fenced: a real per-event retention regression scales
// with the event count and must fail even when hidden behind a large
// transient.
if scaledTail > specBudgetBytes {
t.Fatalf("steady-state tail slope scaled to %d events = %d B, budget is %d B",
eventWindow, scaledTail, specBudgetBytes)
}
// (b) Host-level checks are fenced: on an oversubscribed box report an
// explicit SKIP instead of a red gate. The slope above was still asserted.
if !loadfence.RequireCalm(t, "retained-heap total and process-RSS ceiling") {
return
}
// Total retained growth over the window, against the unchanged spec number.
retainedTotal := int64(last.heap) - int64(baseHeap)
t.Logf("retained heap growth over %d events = %d B (budget %d B)",
eventWindow, retainedTotal, specBudgetBytes)
if retainedTotal > specBudgetBytes {
t.Fatalf("retained heap growth over %d events = %d B, budget is %d B",
eventWindow, retainedTotal, specBudgetBytes)
}
// Coarse process-RSS delta against an explicit ceiling.
rssDelta := int64(last.rss) - int64(baseRSS)
if rssDelta < 0 {
rssDelta = 0
}
t.Logf("process RSS delta over %d events = %d B (ceiling %d B)",
eventWindow, rssDelta, rssCeilingBytes)
if rssDelta > rssCeilingBytes {
t.Fatalf("process RSS delta %d B > ceiling %d B", rssDelta, rssCeilingBytes)
}
}
The production contract this verifies is that Sensor.Handle keeps no per-event state. Any bounded aggregator (counters, fixed sketches) satisfies it; a slice/map keyed by event is a regression.
Environment: Go 1.26, Linux, 16 cores, host load ~8.4 (0.53/core).
gofmt -l . # expect no output
go vet ./... # expect no output
go build ./... # expect no output
go test ./internal/sensors/ -run TestRSSContributionIsBounded -count=5 -v
Observed (5/5 reps PASS). Per-rep logged series, e.g.:
events= 0 heap= 407296 heap_inuse= 819200 rss= 4616192
events= 20000 heap= 376848 heap_inuse= 786432 rss= 5259264
tail: events=12000..20000 growth=256 B slope=64 B/batch scaled_slope=640 B (budget 12582912 B)
retained heap growth over 20000 events = -30448 B (budget 12582912 B)
process RSS delta over 20000 events = 643072 B (ceiling 25165824 B)
--- PASS: TestRSSContributionIsBounded
Across the 5 reps: scaled tail slope 640 B (budget 12 582 912 B), retained growth -41 376 .. -30 448 B, RSS delta 0 .. 643 072 B vs the 25 165 824 B ceiling. ok .../internal/sensors.
Oversubscription path (forced with a tiny threshold, proving an explicit SKIP rather than a red gate — note the slope is still evaluated first):
LOADFENCE_MAX_PER_CPU=0.0001 go test ./internal/sensors/ -run TestRSSContributionIsBounded -v
# loadfence.go: SKIP: retained-heap total and process-RSS ceiling not asserted:
# host load1=12.70/cpu=16 (limit 0.00/cpu) is oversubscribed
# --- SKIP: TestRSSContributionIsBounded
# PASS
A gate that cannot be made to fail by a real retention regression is a phantom gate. Mutate the production path to retain every event:
// internal/sensors/sensors.go
var retained []Event // MUTATION
func (s *Sensor) Handle(e Event) {
retained = append(retained, e) // MUTATION: per-event retention
...
}
Re-run the gate. It must fail loudly on the hard slope assertion:
tail: events=12000..20000 growth=15262208 B slope=3815552 B/batch
scaled_slope=38155520 B (budget 12582912 B)
--- FAIL: TestRSSContributionIsBounded
The scaled slope (≈38 MB) exceeds the 12 MiB budget even though the transient looks like normal noise. Revert the mutation and confirm green again.
For any host-measured budget gate:
go-testing-load-flake).# Evidence - Problem class: go-host-measured-memory-gate-fragile - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T07:12:31.076Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a Go perf gate that asserts a memory budget fails on 3 of 4 environments. internal/sensors TestRSSContributionIsBounded read /proc/self/statm twice (single sample, one runtime.GC() before each) around a 20000-event loop and asserted the RSS delta against a spec budget of 12 MiB: 13455360 B on the reference box, 17825792 B on a CI bunker, 15732736 B on a native cell, all FAIL; one run under HTTP_PROXY=<ip-address>:9 PASSED. Three fails and one pass from identical code is the tell that the measurement, not the code, is at fault.\n\nROOT CAUSE: a raw process-RSS delta is not the subsystem's contribution. Go grows the heap arena in chunks and the scavenger returns pages on its own schedule, so RSS moves by several MB even when the code retains almost nothing; the delta also carries the whole test binary's GC timing and any host/GOGC differences between environments. The old gate was therefore measuring runtime bookkeeping and calling it subsystem state.\n\nFIX (kept the spec budget, replaced the measurement): (1) retained heap - runtime.ReadMemStats after an explicit runtime.GC() with debug.SetGCPercent(100) pinned for the window and restored via defer; (2) a 4000-event warmup, then sample in 2000-event batches so one-time startup growth is outside the window; (3) a STEADY-STATE TAIL SLOPE assertion: the retained growth of the last 4 batches, scaled to the spec's 20k-event window, must stay under the budget - a real per-event retention regression is linear in event count and fails even when hidden behind a large transient, while runtime noise does not scale with events (this assertion is a hard failure, deliberately not load-fenced); (4) total retained growth over the window asserted against the unchanged spec number, and a coarse process-RSS delta asserted against an explicit ceiling (spec budget + measured arena/scavenger slack) - both routed through the repo's loadfence helper so an oversubscribed box reports an explicit SKIP instead of a red gate; (5) every sample (events, heap, heap_inuse, rss), the tail slope and the RSS delta are logged so the verdict is auditable on any box.\n\nVERIFICATION: 5/5 reps of -count=5 PASS with margin (retained heap growth 2.36-2.57 MB vs the 12 MiB budget; scaled tail slope -3.5 KB..256 KB vs 12 MiB; RSS delta 4.15-6.19 MB vs the 24 MiB ceiling; host load 4.26-5.80), whole package green in 47.9 s, build/vet/gofmt clean. FALSIFICATION, run twice independently (the worker's mutation and the foreman's own): retaining every event in a package-level slice makes the gate fail loudly - scaled tail slope 32137660 B vs the 12582912 budget, retained growth 35470240 B, RSS delta 73793536 B. A gate that cannot be made to fail by a real retention regression is a phantom gate, so run this mutation before trusting any rewrite of a memory budget test.\n\nGENERAL RULE: for any host-measured budget gate, (a) measure the quantity the budget is about (retained state for a memory budget, undisturbed cost for a timing budget - see the adjacent class go-testing-load-flake for the timing estimators), (b) exclude one-time startup via warmup, (c) prefer a SLOPE over an absolute delta when the defect of interest is per-unit retention, (d) keep a coarse host-level ceiling as a second net, (e) fence or skip explicitly on an oversubscribed host instead of reporting red, and (f) prove the gate still fails under a deliberate mutation.", "environment": "Go 1.x on Linux; measured on a 16-core control host at load_avg 4.3-5.8, a CI bunker and a native cell", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-host-measured-memory-gate-fragile", "provider": "openrouter", "solved_at": "2026-09-19T07:12:31.077Z", "version": ""}