typescript-benchmark-model-spec-separator-mismatch
I diagnosed the bug, traced it through the code, verified the fix at the fix commit 55d3340 (which is present in the checked-out repo), ran the regression test suite, and wrote a self-contained solution document.
The benchmark CLI documents and passes model specs in slash form (openai/gpt-4o), but the runner's parseModel() only understood the colon form (provider:model). A slash spec fell through to ['CUSTOM', 'openai/gpt-4o'] and was re-joined as 'CUSTOM/openai/gpt-4o' by launchLegacyGame(). Every downstream .split('/') destructure in game-engine.js, legacy-game-adapter.ts, and legacy-usage-collector.js then truncated it to provider='CUSTOM', model='openai' — producing the phantom bare openai row (~810K avg tokens) while the real model showed avgTokens=0, plus the invalid wire id CUSTOM/openai that triggered OpenRouter 400s.
benchmark-runner.ts — parseModel now accepts both separators (first-separator wins); new toLegacyModelSpec passes slash specs through verbatim instead of re-joining CUSTOM-prefixed parts.game-engine.js — role-model resolution derives provider from the first segment but keeps the full spec as the model string.legacy-game-adapter.ts — persistRoleModelAssignments stores the full spec and no longer skips slash-less names.legacy-usage-collector.js — mirrors the same lossless resolution.roleModels verbatim passthrough, no model === 'openai' rows, bare-name persistence, and avgTokens > 0 for both models.tsc --noEmit and the report-time normalization (models.ts normalizeModelKey) collapse prefixed-vs-bare keys so legacy rows stay correct.avgTokens > 0.Full write-up (title, root-cause, exact code, verification) saved to docs/MAF-GAP-057-solution.md.
Note: the repo you cloned already contains the applied fix (HEAD 024c8ff is the closing commit for this issue), so I validated the post-fix state rather than reimplementing it.
# Evidence - Problem class: typescript-benchmark-model-spec-separator-mismatch - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-08-25T11:18:28.536Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Benchmark usage mis-attribution: a benchmark runner's model-spec parser accepted only 'provider:model' (colon) while the CLI documented and passed 'provider/model' (slash). A slash spec ('openai/gpt-4o') fell through to ['CUSTOM', 'openai/gpt-4o'] and was re-joined as 'CUSTOM/openai/gpt-4o'; every downstream .split('/') destructure then truncated it to provider='CUSTOM', model='openai', so token_usage/api_calls rows for the second pairing's players recorded a phantom bare 'openai' model. Report showed the real model with avgTokens=0 and a fake model with ~810K avg tokens; the invalid wire id 'CUSTOM/openai' also caused OpenRouter 400s and estimated-token mock degradation. Fix at the source: parseModel accepts both separators (first-separator wins); toLegacyModelSpec passes slash specs through verbatim; the engine's role-model resolution keeps the full spec as the model string (provider = first segment) instead of truncating; assignment persistence stores the full spec and no longer skips slash-less names. Regression tests across runner/adapter/usage-collector; verified live: 0 phantom rows after container rebuild, real model rows carry avgTokens>0.", "environment": "pnpm monorepo; Express+WS server; legacy 5.3K-line game-engine.js; better-sqlite3; vitest; node 22", "language": "typescript", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "typescript-benchmark-model-spec-separator-mismatch", "provider": "openrouter", "solved_at": "2026-08-25T11:18:28.536Z", "version": ""}I diagnosed the bug, traced it through the code, verified the fix at the fix commit 55d3340 (which is present in the checked-out repo), ran the regression test suite, and wrote a self-contained solution document.
The benchmark CLI documents and passes model specs in slash form (openai/gpt-4o), but the runner's parseModel() only understood the colon form (provider:model). A slash spec fell through to ['CUSTOM', 'openai/gpt-4o'] and was re-joined as 'CUSTOM/openai/gpt-4o' by launchLegacyGame(). Every downstream .split('/') destructure in game-engine.js, legacy-game-adapter.ts, and legacy-usage-collector.js then truncated it to provider='CUSTOM', model='openai' — producing the phantom bare openai row (~810K avg tokens) while the real model showed avgTokens=0, plus the invalid wire id CUSTOM/openai that triggered OpenRouter 400s.
benchmark-runner.ts — parseModel now accepts both separators (first-separator wins); new toLegacyModelSpec passes slash specs through verbatim instead of re-joining CUSTOM-prefixed parts.game-engine.js — role-model resolution derives provider from the first segment but keeps the full spec as the model string.legacy-game-adapter.ts — persistRoleModelAssignments stores the full spec and no longer skips slash-less names.legacy-usage-collector.js — mirrors the same lossless resolution.roleModels verbatim passthrough, no model === 'openai' rows, bare-name persistence, and avgTokens > 0 for both models.tsc --noEmit and the report-time normalization (models.ts normalizeModelKey) collapse prefixed-vs-bare keys so legacy rows stay correct.avgTokens > 0.Full write-up (title, root-cause, exact code, verification) saved to docs/MAF-GAP-057-solution.md.
Note: the repo you cloned already contains the applied fix (HEAD 024c8ff is the closing commit for this issue), so I validated the post-fix state rather than reimplementing it.
# Evidence - Problem class: typescript-benchmark-model-spec-separator-mismatch - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-08-25T11:18:28.536Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Benchmark usage mis-attribution: a benchmark runner's model-spec parser accepted only 'provider:model' (colon) while the CLI documented and passed 'provider/model' (slash). A slash spec ('openai/gpt-4o') fell through to ['CUSTOM', 'openai/gpt-4o'] and was re-joined as 'CUSTOM/openai/gpt-4o'; every downstream .split('/') destructure then truncated it to provider='CUSTOM', model='openai', so token_usage/api_calls rows for the second pairing's players recorded a phantom bare 'openai' model. Report showed the real model with avgTokens=0 and a fake model with ~810K avg tokens; the invalid wire id 'CUSTOM/openai' also caused OpenRouter 400s and estimated-token mock degradation. Fix at the source: parseModel accepts both separators (first-separator wins); toLegacyModelSpec passes slash specs through verbatim; the engine's role-model resolution keeps the full spec as the model string (provider = first segment) instead of truncating; assignment persistence stores the full spec and no longer skips slash-less names. Regression tests across runner/adapter/usage-collector; verified live: 0 phantom rows after container rebuild, real model rows carry avgTokens>0.", "environment": "pnpm monorepo; Express+WS server; legacy 5.3K-line game-engine.js; better-sqlite3; vitest; node 22", "language": "typescript", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "typescript-benchmark-model-spec-separator-mismatch", "provider": "openrouter", "solved_at": "2026-08-25T11:18:28.536Z", "version": ""}