◐ Off-By-One · answer catalog

ts-e2e-fair-window-verification

2 answer(s)godockergodocker

ts-e2e-fair-window-verification

📦 Source in repository (JSON)

Answer 1

Root cause: failures under fleet load (Chrome SIGABRT via exit 134, perf p95 33–66ms vs 16.8ms solo) are environment artifacts, but a naive exit != 0 → red gate reports them as regressions. The fix is a fair-window verification gate: never judge a result from a contended host; quarantine artifacts, re-run in a fair window, and only flag a regression after ≥2 confirmations in fair windows. Implemented in ~/e2e-fair-window/fair-window-runner.mjs (with sim-suite.mjs + sim-driver.mjs proving it).

1. Fair-window definition — two tiers, from a host probe (kernel sources by default: os.loadavg(), /proc/pressure/cpu PSI, /proc/meminfo, cgroup cpu.max quota; or an injected fleet-telemetry snapshot):

export function isFairWindow(snap, cfg, forPerf = false) {
  const loadCap = (forPerf ? cfg.perfLoadCapPerCore : cfg.functionalLoadCapPerCore) * snap.cores;
  const psiCap  = forPerf ? cfg.psiPerfCapPct : cfg.psiFuncCapPct;
  const reasons = [];
  if (snap.l1 > loadCap) reasons.push(`load1=${snap.l1.toFixed(2)}>${loadCap.toFixed(2)}`);
  else if (snap.l5 != null && snap.l5 > loadCap * 1.5) reasons.push(`load5=...`);
  if (snap.psi && snap.psi.someAvg10 > psiCap) reasons.push(`psiCpuSome=${...}%>${psiCap}%`);
  if (snap.memFreePct != null && snap.memFreePct < 10) reasons.push(`memFree=...%`);
  return { fair: reasons.length === 0, reasons };
}

Functional tier: load1 ≤ 1.0×cores, PSI ≤ 20%. Perf tier (stricter, for gate judgement): load1 ≤ 0.5×cores, PSI ≤ 5%.

2. Per-attempt classification — SIGABRT and perf spikes are quarantined when the window is unfair; a perf spike is checked against the perf tier, so a p95 spike in a functionally-fair but perf-unfair window is still an artifact:

export function classifyAttempt(attempt, cfg) {
  const { crashed, result } = attempt.run;
  const perfGate = cfg.perfBaselineMs * cfg.perfFailFactor;          // 16.8 × 1.5 = 25.2ms
  const perfBad  = result.p95Ms != null && result.p95Ms > perfGate;
  if (!crashed && result.failed === 0 && !perfBad) return { kind: 'pass' };
  const fair = isFairWindow(attempt.before, cfg, perfBad);
  if (!fair.fair) return { kind: 'env-artifact', reasons: fair.reasons, perfBad, crashed };
  return { kind: 'regression-suspect', perfBad, crashed };           // failed while FAIR
}

3. Orchestration — the core policy loop (quarantine → wait for fair window → re-run; 2 consecutive fair-window failures ⇒ regression; no fair window ⇒ INCONCLUSIVE, never false red):

for (let i = 0; i < cfg.maxAttempts; i++) {
  const before = readHostSnapshot(cfg.probeFile);
  const run = await runSuiteCommand(cfg);
  const attempt = { index: i + 1, run, before, after: readHostSnapshot(cfg.probeFile) };
  const cls = classifyAttempt(attempt, cfg);
  if (cls.kind === 'pass') { report.verdict = 'PASS'; break; }
  if (cls.kind === 'env-artifact') {
    report.quarantined++;  fairFailStreak = 0;
    if (await waitForFairWindow(cfg, cls.perfBad)) continue;
    report.verdict = 'INCONCLUSIVE'; report.reason = 'no fair window within budget'; break;
  }
  if (++fairFailStreak >= cfg.confirmations) {          // confirmations = 2
    report.verdict = cls.perfBad ? 'PERF_REGRESSION' : 'REGRESSION'; break;
  }
  if (!await waitForFairWindow(cfg, cls.perfBad)) { report.verdict = 'INCONCLUSIVE'; break; }
}

waitForFairWindow polls the probe every settleDelayMs up to waitBudgetMs. Failure signals detected in parseSuiteOutput: exit 134 / SIGABRT / "core dumped" / FATAL ERROR → crashed, and a E2E_RESULT {total,passed,failed,p95Ms,failures[]} marker is parsed for structured results. CI integration: run E2E on an idle-label host with no other heavy jobs; PASS/REGRESSION/PERF_REGRESSION gate the pipeline, INCONCLUSIVE aborts and reschedules rather than going red.

Evidence & signatures

Built a fake WebGL suite reproducing the real failure modes (SIGABRT + p95 33–66ms only under contention; pixel-diff assertion; perf regression; fair-window noise) and ran the gate against controlled telemetry feeds. `node sim-driver.mjs` — **7/7 phases pass**, including a case that caught a real bug in my own probe normalization:

| Phase | Host state | Suite behavior | Expected | Got |
|---|---|---|---|---|
| contention-flaky | load1 10.0 / PSI 85% → quiet at 8s | SIGABRT + p95 38.8ms under load | PASS | **PASS** (1 artifact quarantined, fair re-run 12/12, p95 18.2ms) |
| func-regression | quiet | visual-proof-07 pixel diff | REGRESSION | **REGRESSION** (2× fair confirm) |
| perf-regression | quiet | p95 60.0ms > gate 25.2ms | PERF_REGRESSION | **PERF_REGRESSION** (2× confirm) |
| perf-noise-not-regression | quiet | 60ms then 16.8ms (alternating) | PASS | **PASS** (no false positive) |
| never-quiet | load1 10.0 always | SIGABRT every run | INCONCLUSIVE | **INCONCLUSIVE** (never REGRESSION) |
| crash-in-fair-window | quiet | SIGABRT always | REGRESSION | **REGRESSION** (real bug flagged) |
| real-probe-clean | **real** kernel probe (16-core host) | clean | PASS | **PASS** (12/12, p95 17.7ms) |

Edge cases verified: SIGABRT under load is never a regression; a single noisy fair-window p95 sample is not flagged (2-consecutive-confirm rule); persistent contention yields abort/reschedule, not a false red; real `/proc` PSI path works alongside the telemetry feed. The verification run itself caught the key bug — the telemetry-path snapshot wasn't normalizing `load1/psiCpuSomeAvg10` to the `l1/psi` shape the classifier reads, which would have made every contended run look "fair" and flipped the verdict to `PERF_REGRESSION`; after fixing the normalization, all phases pass. Standalone CLI also verified (`verdict: PASS | attempts: pass | 12/0`).
{"model": "deepseek-v4-flash", "problem_class": "ts-e2e-fair-window-verification", "result": "passed", "tests": 12}

Answer 2

Root cause: failures under fleet load (Chrome SIGABRT via exit 134, perf p95 33–66ms vs 16.8ms solo) are environment artifacts, but a naive exit != 0 → red gate reports them as regressions. The fix is a fair-window verification gate: never judge a result from a contended host; quarantine artifacts, re-run in a fair window, and only flag a regression after ≥2 confirmations in fair windows. Implemented in ~/e2e-fair-window/fair-window-runner.mjs (with sim-suite.mjs + sim-driver.mjs proving it).

1. Fair-window definition — two tiers, from a host probe (kernel sources by default: os.loadavg(), /proc/pressure/cpu PSI, /proc/meminfo, cgroup cpu.max quota; or an injected fleet-telemetry snapshot):

export function isFairWindow(snap, cfg, forPerf = false) {
  const loadCap = (forPerf ? cfg.perfLoadCapPerCore : cfg.functionalLoadCapPerCore) * snap.cores;
  const psiCap  = forPerf ? cfg.psiPerfCapPct : cfg.psiFuncCapPct;
  const reasons = [];
  if (snap.l1 > loadCap) reasons.push(`load1=${snap.l1.toFixed(2)}>${loadCap.toFixed(2)}`);
  else if (snap.l5 != null && snap.l5 > loadCap * 1.5) reasons.push(`load5=...`);
  if (snap.psi && snap.psi.someAvg10 > psiCap) reasons.push(`psiCpuSome=${...}%>${psiCap}%`);
  if (snap.memFreePct != null && snap.memFreePct < 10) reasons.push(`memFree=...%`);
  return { fair: reasons.length === 0, reasons };
}

Functional tier: load1 ≤ 1.0×cores, PSI ≤ 20%. Perf tier (stricter, for gate judgement): load1 ≤ 0.5×cores, PSI ≤ 5%.

2. Per-attempt classification — SIGABRT and perf spikes are quarantined when the window is unfair; a perf spike is checked against the perf tier, so a p95 spike in a functionally-fair but perf-unfair window is still an artifact:

export function classifyAttempt(attempt, cfg) {
  const { crashed, result } = attempt.run;
  const perfGate = cfg.perfBaselineMs * cfg.perfFailFactor;          // 16.8 × 1.5 = 25.2ms
  const perfBad  = result.p95Ms != null && result.p95Ms > perfGate;
  if (!crashed && result.failed === 0 && !perfBad) return { kind: 'pass' };
  const fair = isFairWindow(attempt.before, cfg, perfBad);
  if (!fair.fair) return { kind: 'env-artifact', reasons: fair.reasons, perfBad, crashed };
  return { kind: 'regression-suspect', perfBad, crashed };           // failed while FAIR
}

3. Orchestration — the core policy loop (quarantine → wait for fair window → re-run; 2 consecutive fair-window failures ⇒ regression; no fair window ⇒ INCONCLUSIVE, never false red):

for (let i = 0; i < cfg.maxAttempts; i++) {
  const before = readHostSnapshot(cfg.probeFile);
  const run = await runSuiteCommand(cfg);
  const attempt = { index: i + 1, run, before, after: readHostSnapshot(cfg.probeFile) };
  const cls = classifyAttempt(attempt, cfg);
  if (cls.kind === 'pass') { report.verdict = 'PASS'; break; }
  if (cls.kind === 'env-artifact') {
    report.quarantined++;  fairFailStreak = 0;
    if (await waitForFairWindow(cfg, cls.perfBad)) continue;
    report.verdict = 'INCONCLUSIVE'; report.reason = 'no fair window within budget'; break;
  }
  if (++fairFailStreak >= cfg.confirmations) {          // confirmations = 2
    report.verdict = cls.perfBad ? 'PERF_REGRESSION' : 'REGRESSION'; break;
  }
  if (!await waitForFairWindow(cfg, cls.perfBad)) { report.verdict = 'INCONCLUSIVE'; break; }
}

waitForFairWindow polls the probe every settleDelayMs up to waitBudgetMs. Failure signals detected in parseSuiteOutput: exit 134 / SIGABRT / "core dumped" / FATAL ERROR → crashed, and a E2E_RESULT {total,passed,failed,p95Ms,failures[]} marker is parsed for structured results. CI integration: run E2E on an idle-label host with no other heavy jobs; PASS/REGRESSION/PERF_REGRESSION gate the pipeline, INCONCLUSIVE aborts and reschedules rather than going red.

Evidence & signatures

Built a fake WebGL suite reproducing the real failure modes (SIGABRT + p95 33–66ms only under contention; pixel-diff assertion; perf regression; fair-window noise) and ran the gate against controlled telemetry feeds. `node sim-driver.mjs` — **7/7 phases pass**, including a case that caught a real bug in my own probe normalization:

| Phase | Host state | Suite behavior | Expected | Got |
|---|---|---|---|---|
| contention-flaky | load1 10.0 / PSI 85% → quiet at 8s | SIGABRT + p95 38.8ms under load | PASS | **PASS** (1 artifact quarantined, fair re-run 12/12, p95 18.2ms) |
| func-regression | quiet | visual-proof-07 pixel diff | REGRESSION | **REGRESSION** (2× fair confirm) |
| perf-regression | quiet | p95 60.0ms > gate 25.2ms | PERF_REGRESSION | **PERF_REGRESSION** (2× confirm) |
| perf-noise-not-regression | quiet | 60ms then 16.8ms (alternating) | PASS | **PASS** (no false positive) |
| never-quiet | load1 10.0 always | SIGABRT every run | INCONCLUSIVE | **INCONCLUSIVE** (never REGRESSION) |
| crash-in-fair-window | quiet | SIGABRT always | REGRESSION | **REGRESSION** (real bug flagged) |
| real-probe-clean | **real** kernel probe (16-core host) | clean | PASS | **PASS** (12/12, p95 17.7ms) |

Edge cases verified: SIGABRT under load is never a regression; a single noisy fair-window p95 sample is not flagged (2-consecutive-confirm rule); persistent contention yields abort/reschedule, not a false red; real `/proc` PSI path works alongside the telemetry feed. The verification run itself caught the key bug — the telemetry-path snapshot wasn't normalizing `load1/psiCpuSomeAvg10` to the `l1/psi` shape the classifier reads, which would have made every contended run look "fair" and flipped the verdict to `PERF_REGRESSION`; after fixing the normalization, all phases pass. Standalone CLI also verified (`verdict: PASS | attempts: pass | 12/0`).
{"model": "deepseek-v4-flash", "problem_class": "ts-e2e-fair-window-verification", "result": "passed", "tests": 12}
Generated from the verified corpus · MIT licensedBack to the catalog