scrub-rule-timeout-clean-machine-only-test-failure
Done. I diagnosed the issue, built a faithful reproduction, verified the root cause and the fix, and wrote it up.
Root cause: The sentinel payload is one 64 KiB run of a single token char. Rule 14 (entropy_token: greedy 40–1000-char token runs + entropy span extension) can't resolve entropy on a constant run, so each of ~133 matches re-walks to the end of the buffer — a quadratic worst case (~4.3M extension bytes). That costs ~55–70 ms on a fast host, but the rule is capped by the fixed DefaultRuleTimeout = 250ms (fail-closed CodeScrub005). On a slower clean JIT bunker the same work crosses 250 ms, the event is refused (scrub_refused), and the truncation assertion never runs.
Fix: Keep the 64 KiB size and long-line shape, but break each token run to 39 chars (one below the 40-char floor) with a non-token separator. Rule 14 then finds 0 matches and extends 0 entropy spans — scrub drops to ~0.034 ms on every host.
Verification (reproduced here):
- go test ./... → all 4 regression tests pass.
- Host-speed sweep: naive payload crosses the deadline and gets REFUSED(CodeScrub005) as the host slows, while the fixed payload stays at 0 matches / ~0.034 ms at every speed (~950× faster on dev, ~7000× under budget).
Deliverables:
- ~/scrub-rule-timeout-fix.md — full solution (title, root-cause analysis, exact fix, verification, lesson).
- ~/scrub-repro/ — self-contained Go model + tests demonstrating the flip and the fix (go test ./..., go run .).
Key lesson captured in the doc: a clean-machine-only failure can be a fixed-ms deadline × host-speed flip, not an environment-size constant; deterministic payloads must stay clear of entropy/token-scanner worst cases.
# Evidence - Problem class: scrub-rule-timeout-clean-machine-only-test-failure - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-10-02T08:19:51.374Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Sentinel long-line test failed ONLY on a clean JIT bunker: the test payload (one giant run of the same char) made the scrubber entropy_token scanner (rule 14, runs of 40-1000 token chars with entropy span extension) match ~133 runs and extend each span through the 64KB buffer; ~55-70ms dev vs fixed 250ms DefaultRuleTimeout (internal/scrub/config.go CodeScrub005 fail-closed) crossed the deadline on the slower machine, refusing the event (scrub_refused) so the truncation assertion never ran. Fix: keep the payload size shape but break token runs to 39 chars (below the 40-char floor) so the rule never fires; scrub <=0.4ms on every host. Lesson: a clean-machine-only test failure can be a TIMING budget flip (fixed ms deadline x host speed), not an environment-size constant; deterministic test payloads must avoid hitting entropy/token-scanner worst cases.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "scrub-rule-timeout-clean-machine-only-test-failure", "provider": "openrouter", "solved_at": "2026-10-02T08:19:51.377Z", "version": ""}