ts-e2e-fair-window-verification
Root cause: failures under fleet load (Chrome SIGABRT via exit 134, perf p95 33–66ms vs 16.8ms solo) are environment artifacts, but a naive exit != 0 → red gate reports them as regressions. The fix is a fair-window verification gate: never judge a result from a contended host; quarantine artifacts, re-run in a fair window, and only flag a regression after ≥2 confirmations in fair windows. Implemented in ~/e2e-fair-window/fair-window-runner.mjs (with sim-suite.mjs + sim-driver.mjs proving it).
1. Fair-window definition — two tiers, from a host probe (kernel sources by default: os.loadavg(), /proc/pressure/cpu PSI, /proc/meminfo, cgroup cpu.max quota; or an injected fleet-telemetry snapshot):
export function isFairWindow(snap, cfg, forPerf = false) {
const loadCap = (forPerf ? cfg.perfLoadCapPerCore : cfg.functionalLoadCapPerCore) * snap.cores;
const psiCap = forPerf ? cfg.psiPerfCapPct : cfg.psiFuncCapPct;
const reasons = [];
if (snap.l1 > loadCap) reasons.push(`load1=${snap.l1.toFixed(2)}>${loadCap.toFixed(2)}`);
else if (snap.l5 != null && snap.l5 > loadCap * 1.5) reasons.push(`load5=...`);
if (snap.psi && snap.psi.someAvg10 > psiCap) reasons.push(`psiCpuSome=${...}%>${psiCap}%`);
if (snap.memFreePct != null && snap.memFreePct < 10) reasons.push(`memFree=...%`);
return { fair: reasons.length === 0, reasons };
}
Functional tier: load1 ≤ 1.0×cores, PSI ≤ 20%. Perf tier (stricter, for gate judgement): load1 ≤ 0.5×cores, PSI ≤ 5%.
2. Per-attempt classification — SIGABRT and perf spikes are quarantined when the window is unfair; a perf spike is checked against the perf tier, so a p95 spike in a functionally-fair but perf-unfair window is still an artifact:
export function classifyAttempt(attempt, cfg) {
const { crashed, result } = attempt.run;
const perfGate = cfg.perfBaselineMs * cfg.perfFailFactor; // 16.8 × 1.5 = 25.2ms
const perfBad = result.p95Ms != null && result.p95Ms > perfGate;
if (!crashed && result.failed === 0 && !perfBad) return { kind: 'pass' };
const fair = isFairWindow(attempt.before, cfg, perfBad);
if (!fair.fair) return { kind: 'env-artifact', reasons: fair.reasons, perfBad, crashed };
return { kind: 'regression-suspect', perfBad, crashed }; // failed while FAIR
}
3. Orchestration — the core policy loop (quarantine → wait for fair window → re-run; 2 consecutive fair-window failures ⇒ regression; no fair window ⇒ INCONCLUSIVE, never false red):
for (let i = 0; i < cfg.maxAttempts; i++) {
const before = readHostSnapshot(cfg.probeFile);
const run = await runSuiteCommand(cfg);
const attempt = { index: i + 1, run, before, after: readHostSnapshot(cfg.probeFile) };
const cls = classifyAttempt(attempt, cfg);
if (cls.kind === 'pass') { report.verdict = 'PASS'; break; }
if (cls.kind === 'env-artifact') {
report.quarantined++; fairFailStreak = 0;
if (await waitForFairWindow(cfg, cls.perfBad)) continue;
report.verdict = 'INCONCLUSIVE'; report.reason = 'no fair window within budget'; break;
}
if (++fairFailStreak >= cfg.confirmations) { // confirmations = 2
report.verdict = cls.perfBad ? 'PERF_REGRESSION' : 'REGRESSION'; break;
}
if (!await waitForFairWindow(cfg, cls.perfBad)) { report.verdict = 'INCONCLUSIVE'; break; }
}
waitForFairWindow polls the probe every settleDelayMs up to waitBudgetMs. Failure signals detected in parseSuiteOutput: exit 134 / SIGABRT / "core dumped" / FATAL ERROR → crashed, and a E2E_RESULT {total,passed,failed,p95Ms,failures[]} marker is parsed for structured results. CI integration: run E2E on an idle-label host with no other heavy jobs; PASS/REGRESSION/PERF_REGRESSION gate the pipeline, INCONCLUSIVE aborts and reschedules rather than going red.
Built a fake WebGL suite reproducing the real failure modes (SIGABRT + p95 33–66ms only under contention; pixel-diff assertion; perf regression; fair-window noise) and ran the gate against controlled telemetry feeds. `node sim-driver.mjs` — **7/7 phases pass**, including a case that caught a real bug in my own probe normalization: | Phase | Host state | Suite behavior | Expected | Got | |---|---|---|---|---| | contention-flaky | load1 10.0 / PSI 85% → quiet at 8s | SIGABRT + p95 38.8ms under load | PASS | **PASS** (1 artifact quarantined, fair re-run 12/12, p95 18.2ms) | | func-regression | quiet | visual-proof-07 pixel diff | REGRESSION | **REGRESSION** (2× fair confirm) | | perf-regression | quiet | p95 60.0ms > gate 25.2ms | PERF_REGRESSION | **PERF_REGRESSION** (2× confirm) | | perf-noise-not-regression | quiet | 60ms then 16.8ms (alternating) | PASS | **PASS** (no false positive) | | never-quiet | load1 10.0 always | SIGABRT every run | INCONCLUSIVE | **INCONCLUSIVE** (never REGRESSION) | | crash-in-fair-window | quiet | SIGABRT always | REGRESSION | **REGRESSION** (real bug flagged) | | real-probe-clean | **real** kernel probe (16-core host) | clean | PASS | **PASS** (12/12, p95 17.7ms) | Edge cases verified: SIGABRT under load is never a regression; a single noisy fair-window p95 sample is not flagged (2-consecutive-confirm rule); persistent contention yields abort/reschedule, not a false red; real `/proc` PSI path works alongside the telemetry feed. The verification run itself caught the key bug — the telemetry-path snapshot wasn't normalizing `load1/psiCpuSomeAvg10` to the `l1/psi` shape the classifier reads, which would have made every contended run look "fair" and flipped the verdict to `PERF_REGRESSION`; after fixing the normalization, all phases pass. Standalone CLI also verified (`verdict: PASS | attempts: pass | 12/0`).
{"model": "deepseek-v4-flash", "problem_class": "ts-e2e-fair-window-verification", "result": "passed", "tests": 12}Root cause: failures under fleet load (Chrome SIGABRT via exit 134, perf p95 33–66ms vs 16.8ms solo) are environment artifacts, but a naive exit != 0 → red gate reports them as regressions. The fix is a fair-window verification gate: never judge a result from a contended host; quarantine artifacts, re-run in a fair window, and only flag a regression after ≥2 confirmations in fair windows. Implemented in ~/e2e-fair-window/fair-window-runner.mjs (with sim-suite.mjs + sim-driver.mjs proving it).
1. Fair-window definition — two tiers, from a host probe (kernel sources by default: os.loadavg(), /proc/pressure/cpu PSI, /proc/meminfo, cgroup cpu.max quota; or an injected fleet-telemetry snapshot):
export function isFairWindow(snap, cfg, forPerf = false) {
const loadCap = (forPerf ? cfg.perfLoadCapPerCore : cfg.functionalLoadCapPerCore) * snap.cores;
const psiCap = forPerf ? cfg.psiPerfCapPct : cfg.psiFuncCapPct;
const reasons = [];
if (snap.l1 > loadCap) reasons.push(`load1=${snap.l1.toFixed(2)}>${loadCap.toFixed(2)}`);
else if (snap.l5 != null && snap.l5 > loadCap * 1.5) reasons.push(`load5=...`);
if (snap.psi && snap.psi.someAvg10 > psiCap) reasons.push(`psiCpuSome=${...}%>${psiCap}%`);
if (snap.memFreePct != null && snap.memFreePct < 10) reasons.push(`memFree=...%`);
return { fair: reasons.length === 0, reasons };
}
Functional tier: load1 ≤ 1.0×cores, PSI ≤ 20%. Perf tier (stricter, for gate judgement): load1 ≤ 0.5×cores, PSI ≤ 5%.
2. Per-attempt classification — SIGABRT and perf spikes are quarantined when the window is unfair; a perf spike is checked against the perf tier, so a p95 spike in a functionally-fair but perf-unfair window is still an artifact:
export function classifyAttempt(attempt, cfg) {
const { crashed, result } = attempt.run;
const perfGate = cfg.perfBaselineMs * cfg.perfFailFactor; // 16.8 × 1.5 = 25.2ms
const perfBad = result.p95Ms != null && result.p95Ms > perfGate;
if (!crashed && result.failed === 0 && !perfBad) return { kind: 'pass' };
const fair = isFairWindow(attempt.before, cfg, perfBad);
if (!fair.fair) return { kind: 'env-artifact', reasons: fair.reasons, perfBad, crashed };
return { kind: 'regression-suspect', perfBad, crashed }; // failed while FAIR
}
3. Orchestration — the core policy loop (quarantine → wait for fair window → re-run; 2 consecutive fair-window failures ⇒ regression; no fair window ⇒ INCONCLUSIVE, never false red):
for (let i = 0; i < cfg.maxAttempts; i++) {
const before = readHostSnapshot(cfg.probeFile);
const run = await runSuiteCommand(cfg);
const attempt = { index: i + 1, run, before, after: readHostSnapshot(cfg.probeFile) };
const cls = classifyAttempt(attempt, cfg);
if (cls.kind === 'pass') { report.verdict = 'PASS'; break; }
if (cls.kind === 'env-artifact') {
report.quarantined++; fairFailStreak = 0;
if (await waitForFairWindow(cfg, cls.perfBad)) continue;
report.verdict = 'INCONCLUSIVE'; report.reason = 'no fair window within budget'; break;
}
if (++fairFailStreak >= cfg.confirmations) { // confirmations = 2
report.verdict = cls.perfBad ? 'PERF_REGRESSION' : 'REGRESSION'; break;
}
if (!await waitForFairWindow(cfg, cls.perfBad)) { report.verdict = 'INCONCLUSIVE'; break; }
}
waitForFairWindow polls the probe every settleDelayMs up to waitBudgetMs. Failure signals detected in parseSuiteOutput: exit 134 / SIGABRT / "core dumped" / FATAL ERROR → crashed, and a E2E_RESULT {total,passed,failed,p95Ms,failures[]} marker is parsed for structured results. CI integration: run E2E on an idle-label host with no other heavy jobs; PASS/REGRESSION/PERF_REGRESSION gate the pipeline, INCONCLUSIVE aborts and reschedules rather than going red.
Built a fake WebGL suite reproducing the real failure modes (SIGABRT + p95 33–66ms only under contention; pixel-diff assertion; perf regression; fair-window noise) and ran the gate against controlled telemetry feeds. `node sim-driver.mjs` — **7/7 phases pass**, including a case that caught a real bug in my own probe normalization: | Phase | Host state | Suite behavior | Expected | Got | |---|---|---|---|---| | contention-flaky | load1 10.0 / PSI 85% → quiet at 8s | SIGABRT + p95 38.8ms under load | PASS | **PASS** (1 artifact quarantined, fair re-run 12/12, p95 18.2ms) | | func-regression | quiet | visual-proof-07 pixel diff | REGRESSION | **REGRESSION** (2× fair confirm) | | perf-regression | quiet | p95 60.0ms > gate 25.2ms | PERF_REGRESSION | **PERF_REGRESSION** (2× confirm) | | perf-noise-not-regression | quiet | 60ms then 16.8ms (alternating) | PASS | **PASS** (no false positive) | | never-quiet | load1 10.0 always | SIGABRT every run | INCONCLUSIVE | **INCONCLUSIVE** (never REGRESSION) | | crash-in-fair-window | quiet | SIGABRT always | REGRESSION | **REGRESSION** (real bug flagged) | | real-probe-clean | **real** kernel probe (16-core host) | clean | PASS | **PASS** (12/12, p95 17.7ms) | Edge cases verified: SIGABRT under load is never a regression; a single noisy fair-window p95 sample is not flagged (2-consecutive-confirm rule); persistent contention yields abort/reschedule, not a false red; real `/proc` PSI path works alongside the telemetry feed. The verification run itself caught the key bug — the telemetry-path snapshot wasn't normalizing `load1/psiCpuSomeAvg10` to the `l1/psi` shape the classifier reads, which would have made every contended run look "fair" and flipped the verdict to `PERF_REGRESSION`; after fixing the normalization, all phases pass. Standalone CLI also verified (`verdict: PASS | attempts: pass | 12/0`).
{"model": "deepseek-v4-flash", "problem_class": "ts-e2e-fair-window-verification", "result": "passed", "tests": 12}