playwright-slide-sustain-load-flake
Done. The solution is at ~/work/slide-flake/SOLUTION.md, backed by a self-contained, runnable verification harness (game.html sim + qa.spec.js + playwright.config.js).
Not a game bug — slide code is correct (DOGFOOD-001 clamp). It's a test-architecture race: the test gates a live in-page measurement (200 ms sample) on a live CDP read (waitForFunction(slideT > 0, 15s)). Under load-59 both legs of that handoff degrade: the 0.5 s one-shot slide collapses to 2–3 rAF frames (250 ms+ apart, dt-clamped burst stepping), while the CDP→Node dispatch is delayed hundreds of ms by host contention. The wait can therefore resolve after the slide finished, so the sample starts at slideT = 0 → expect(sustained).toBe(true) (qa.spec.js:175) gets false.
sustained | dropped | missed) instead of live slideT > 0; re-assert immediately; retry only environmental misses (bounded 3×). A genuine premature drop (dropped) still hard-fails, so DOGFOOD-001 regressions aren't masked.keyboard.down, riding the game's own rAF, recording engagedAt/lastSeen/zeroAt; the record is immutable, so a late fetch is harmless.window.__bank.SLIDE_TIME 0.5 s → 3.0 s under test so the sustain window is always observable even at worst-case frame gaps (the clamp under test is duration-invariant).Harness validation also caught a subtle Playwright API bug in an early revision of the patch: waitForFunction resolves with a JSHandle, not the value — unwrap with handle.jsonValue().
| Case | Result |
|---|---|
| ORIGINAL pattern, ticket-like load (250 ms bursts·p0.65 · 450 ms dispatch delay, ×10) | 6/10 failed — flake reproduced (profile matches "2 fails then 1 pass 5/5"; 350/450/550 ms delay sweep gave 1/8–4/8–2/8) |
| FIXED, identical settings ×10 | 10/10 passed |
| FIXED, forced starvation p=1.0 + 1000 ms delay ×5 | 5/5 passed |
| FIXED matrix (125/500 ms bursts · 0.65/1.0 · 450–1200 ms) | 9/9 passed |
| Regression guard (broken clamp) | fails as designed (dropped:true, sustained:false) |
| Zero contention sanity | passes, no behavioral change |
All reproduction commands are in the solution doc's Verification section, so it's directly usable against the real temple-runner repo (the patch touches exactly two game interfaces — window.slideT reads and the __bank.SLIDE_TIME write — both named by the ticket).
# Evidence - Problem class: playwright-slide-sustain-load-flake - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-09-01T18:22:11.462Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "slide-sustain test (qa.spec.js:175) intermittently fails under extreme host load (load avg 48-60): expect(sustained).toBe(true) receives false. The test holds ArrowDown, waits waitForFunction(slideT > 0, 15s), then samples in-page for 200ms requiring slideT > 0 throughout. Under heavy CPU contention (mafia-fresh vitest + ollama llama-server pegging cores), the 15s waitForFunction can resolve AFTER the 0.5s slide already completed (rAF frames ~250ms+ apart, sim steps in bursts), so the 200ms sample starts with slideT already at 0. Not a game bug: slide code is correct (Math.max(0, slideT-dt) clamp, DOGFOOD-001); 3rd run passed 5/5 under same load. Pre-existing host-contention flake (tick 29 characterized: old code flaked 1/5 same load).", "environment": "temple-runner NEON RUSH single-file three.js game, Playwright QA battery on GPU-less SwiftShader host", "language": "javascript", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "playwright-slide-sustain-load-flake", "provider": "openrouter", "solved_at": "2026-09-01T18:22:11.463Z", "version": "0afc4a3"}Done. The solution is at ~/work/slide-flake/SOLUTION.md, backed by a self-contained, runnable verification harness (game.html sim + qa.spec.js + playwright.config.js).
Not a game bug — slide code is correct (DOGFOOD-001 clamp). It's a test-architecture race: the test gates a live in-page measurement (200 ms sample) on a live CDP read (waitForFunction(slideT > 0, 15s)). Under load-59 both legs of that handoff degrade: the 0.5 s one-shot slide collapses to 2–3 rAF frames (250 ms+ apart, dt-clamped burst stepping), while the CDP→Node dispatch is delayed hundreds of ms by host contention. The wait can therefore resolve after the slide finished, so the sample starts at slideT = 0 → expect(sustained).toBe(true) (qa.spec.js:175) gets false.
sustained | dropped | missed) instead of live slideT > 0; re-assert immediately; retry only environmental misses (bounded 3×). A genuine premature drop (dropped) still hard-fails, so DOGFOOD-001 regressions aren't masked.keyboard.down, riding the game's own rAF, recording engagedAt/lastSeen/zeroAt; the record is immutable, so a late fetch is harmless.window.__bank.SLIDE_TIME 0.5 s → 3.0 s under test so the sustain window is always observable even at worst-case frame gaps (the clamp under test is duration-invariant).Harness validation also caught a subtle Playwright API bug in an early revision of the patch: waitForFunction resolves with a JSHandle, not the value — unwrap with handle.jsonValue().
| Case | Result |
|---|---|
| ORIGINAL pattern, ticket-like load (250 ms bursts·p0.65 · 450 ms dispatch delay, ×10) | 6/10 failed — flake reproduced (profile matches "2 fails then 1 pass 5/5"; 350/450/550 ms delay sweep gave 1/8–4/8–2/8) |
| FIXED, identical settings ×10 | 10/10 passed |
| FIXED, forced starvation p=1.0 + 1000 ms delay ×5 | 5/5 passed |
| FIXED matrix (125/500 ms bursts · 0.65/1.0 · 450–1200 ms) | 9/9 passed |
| Regression guard (broken clamp) | fails as designed (dropped:true, sustained:false) |
| Zero contention sanity | passes, no behavioral change |
All reproduction commands are in the solution doc's Verification section, so it's directly usable against the real temple-runner repo (the patch touches exactly two game interfaces — window.slideT reads and the __bank.SLIDE_TIME write — both named by the ticket).
# Evidence - Problem class: playwright-slide-sustain-load-flake - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-09-01T18:22:11.462Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "slide-sustain test (qa.spec.js:175) intermittently fails under extreme host load (load avg 48-60): expect(sustained).toBe(true) receives false. The test holds ArrowDown, waits waitForFunction(slideT > 0, 15s), then samples in-page for 200ms requiring slideT > 0 throughout. Under heavy CPU contention (mafia-fresh vitest + ollama llama-server pegging cores), the 15s waitForFunction can resolve AFTER the 0.5s slide already completed (rAF frames ~250ms+ apart, sim steps in bursts), so the 200ms sample starts with slideT already at 0. Not a game bug: slide code is correct (Math.max(0, slideT-dt) clamp, DOGFOOD-001); 3rd run passed 5/5 under same load. Pre-existing host-contention flake (tick 29 characterized: old code flaked 1/5 same load).", "environment": "temple-runner NEON RUSH single-file three.js game, Playwright QA battery on GPU-less SwiftShader host", "language": "javascript", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "playwright-slide-sustain-load-flake", "provider": "openrouter", "solved_at": "2026-09-01T18:22:11.463Z", "version": "0afc4a3"}