typescript-benchmark-report-dup-rows
Root cause (two independent defects):
model column is sometimes prefixed (openai/gpt-4o-mini) and sometimes bare (gpt-4o-mini). Grouping on the raw string treats them as two different models, so a single model appears as duplicate rows with contradictory totals.winRate=1.0 rows. When a game had a winner, the old code credited that win to every assigned model. Over 127 games that produced CUSTOM/openai 127/127 (games=127, wins=127, winRate=1.0) for models that never won, contradicting the real aggregated rates.The fix (src/benchmark-report.ts):
- Normalize before grouping — every model id is parsed into a canonical provider/name key (normalizeModel). openai/gpt-4o-mini and gpt-4o-mini (with provider: "openai") both collapse to key openai/gpt-4o-mini, so aggregation merges them into a single row. Record-level provider takes precedence over a prefix.
- Delete the blanket win attribution — only the model whose id equals the game's normalized winner gets a wins++. Ties/no-winner games contribute games but zero wins. Wins are also never credited to a model that wasn't assigned.
// src/benchmark-report.ts (core of the fix)
/** "openai/gpt-4o-mini" -> { provider: "openai", name: "gpt-4o-mini", key: "openai/gpt-4o-mini" } */
export function normalizeModel(raw: string, recordProvider?: string | null): NormalizedModel {
const trimmed = raw.trim();
let provider = recordProvider?.trim() || null;
let name = trimmed;
const slash = trimmed.indexOf("/");
if (slash > 0) {
provider ??= trimmed.slice(0, slash); // prefix is a provider only when unknown
name = trimmed.slice(slash + 1);
}
return { provider, name, key: provider ? `${provider}/${name}` : name };
}
// inside buildReport(records):
for (const rec of records) {
if (seenIds.has(rec.id)) continue; // dedupe duplicate game rows
seenIds.add(rec.id);
const assigned = new Map<string, NormalizedModel>();
for (const m of rec.models ?? []) {
const norm = normalizeModel(m, rec.provider); // canonical key -> no split rows
if (!assigned.has(norm.key)) assigned.set(norm.key, norm);
}
if (assigned.size === 0) continue;
let winnerKey: string | null = null;
if (rec.winner?.trim()) winnerKey = normalizeModel(rec.winner, rec.provider).key;
for (const [key, norm] of assigned) {
// games++ once per (game, model)
// wins++ ONLY when winnerKey === key <-- legacy fallback removed
}
}
Side-by-side on the problem scenario (127 games, models ["openai/gpt-4o-mini","gpt-4o-mini","claude-3-5-sonnet"], winner always gpt-4o-mini):
| BUGGY (before) | FIXED (after) |
|---|---|
openai/gpt-4o-mini 127/127 winRate 1.0 |
gpt-4o-mini (openai) 127/127 winRate 1.0 ← real winner |
gpt-4o-mini 127/127 winRate 1.0 (split dup row) |
merged into the row above |
claude-3-5-sonnet 127/127 winRate 1.0 (fabricated) |
claude-3-5-sonnet (openai) 127 games, 0 wins, winRate 0 |
Verified by running the suite under the target runtime (`node v22.22.3` + `tsx v4.23.12`, matching the docker-compose node22 / typescript-node22-tsx environment): ``` # tests 10 # pass 10 # fail 0 ``` Files: `src/benchmark-report.ts` (fix), `test/benchmark-report.test.ts` (10 tests via `node:test`). Edge cases tested: - **Provider-prefix merge** — prefixed and bare spellings of the same model collapse into exactly one row with merged games/wins. - **No fabricated 127/127** — the exact problem scenario: the non-winning assigned model reports 127 games / 0 wins / winRate 0; exactly 2 rows, no split rows. - **Consistency invariant** — one win per decisive game, never more; total row wins === number of decisive games. - **Record-level provider precedence** — `provider:"CUSTOM"` + `model:"openai/gpt-4o-mini"` yields `CUSTOM/gpt-4o-mini`; the winner normalized the same way still matches. - **Duplicate game ids** in the input counted once; **duplicate entries within one game's model list** counted once. - **Winner not among assigned models** → no random model is credited (0 wins, defensive). - **Tie / no winner** → zero wins, winRate 0; **empty input** → empty report with `totalGames: 0`. The demo run above prints the buggy vs fixed rows for the 127-game scenario: the buggy output shows 3 rows including the split `openai/gpt-4o-mini` + `gpt-4o-mini` pair and a fabricated `claude-3-5-sonnet 127/127`; the fixed output shows 2 consistent rows (127/127 for the real winner, 0/127 for the loser) with `totalGames: 127`.
{"model": "deepseek-v4-flash", "problem_class": "typescript-benchmark-report-dup-rows", "result": "passed", "tests": 10}Root cause (two independent defects):
model column is sometimes prefixed (openai/gpt-4o-mini) and sometimes bare (gpt-4o-mini). Grouping on the raw string treats them as two different models, so a single model appears as duplicate rows with contradictory totals.winRate=1.0 rows. When a game had a winner, the old code credited that win to every assigned model. Over 127 games that produced CUSTOM/openai 127/127 (games=127, wins=127, winRate=1.0) for models that never won, contradicting the real aggregated rates.The fix (src/benchmark-report.ts):
- Normalize before grouping — every model id is parsed into a canonical provider/name key (normalizeModel). openai/gpt-4o-mini and gpt-4o-mini (with provider: "openai") both collapse to key openai/gpt-4o-mini, so aggregation merges them into a single row. Record-level provider takes precedence over a prefix.
- Delete the blanket win attribution — only the model whose id equals the game's normalized winner gets a wins++. Ties/no-winner games contribute games but zero wins. Wins are also never credited to a model that wasn't assigned.
// src/benchmark-report.ts (core of the fix)
/** "openai/gpt-4o-mini" -> { provider: "openai", name: "gpt-4o-mini", key: "openai/gpt-4o-mini" } */
export function normalizeModel(raw: string, recordProvider?: string | null): NormalizedModel {
const trimmed = raw.trim();
let provider = recordProvider?.trim() || null;
let name = trimmed;
const slash = trimmed.indexOf("/");
if (slash > 0) {
provider ??= trimmed.slice(0, slash); // prefix is a provider only when unknown
name = trimmed.slice(slash + 1);
}
return { provider, name, key: provider ? `${provider}/${name}` : name };
}
// inside buildReport(records):
for (const rec of records) {
if (seenIds.has(rec.id)) continue; // dedupe duplicate game rows
seenIds.add(rec.id);
const assigned = new Map<string, NormalizedModel>();
for (const m of rec.models ?? []) {
const norm = normalizeModel(m, rec.provider); // canonical key -> no split rows
if (!assigned.has(norm.key)) assigned.set(norm.key, norm);
}
if (assigned.size === 0) continue;
let winnerKey: string | null = null;
if (rec.winner?.trim()) winnerKey = normalizeModel(rec.winner, rec.provider).key;
for (const [key, norm] of assigned) {
// games++ once per (game, model)
// wins++ ONLY when winnerKey === key <-- legacy fallback removed
}
}
Side-by-side on the problem scenario (127 games, models ["openai/gpt-4o-mini","gpt-4o-mini","claude-3-5-sonnet"], winner always gpt-4o-mini):
| BUGGY (before) | FIXED (after) |
|---|---|
openai/gpt-4o-mini 127/127 winRate 1.0 |
gpt-4o-mini (openai) 127/127 winRate 1.0 ← real winner |
gpt-4o-mini 127/127 winRate 1.0 (split dup row) |
merged into the row above |
claude-3-5-sonnet 127/127 winRate 1.0 (fabricated) |
claude-3-5-sonnet (openai) 127 games, 0 wins, winRate 0 |
Verified by running the suite under the target runtime (`node v22.22.3` + `tsx v4.23.12`, matching the docker-compose node22 / typescript-node22-tsx environment): ``` # tests 10 # pass 10 # fail 0 ``` Files: `src/benchmark-report.ts` (fix), `test/benchmark-report.test.ts` (10 tests via `node:test`). Edge cases tested: - **Provider-prefix merge** — prefixed and bare spellings of the same model collapse into exactly one row with merged games/wins. - **No fabricated 127/127** — the exact problem scenario: the non-winning assigned model reports 127 games / 0 wins / winRate 0; exactly 2 rows, no split rows. - **Consistency invariant** — one win per decisive game, never more; total row wins === number of decisive games. - **Record-level provider precedence** — `provider:"CUSTOM"` + `model:"openai/gpt-4o-mini"` yields `CUSTOM/gpt-4o-mini`; the winner normalized the same way still matches. - **Duplicate game ids** in the input counted once; **duplicate entries within one game's model list** counted once. - **Winner not among assigned models** → no random model is credited (0 wins, defensive). - **Tie / no winner** → zero wins, winRate 0; **empty input** → empty report with `totalGames: 0`. The demo run above prints the buggy vs fixed rows for the 127-game scenario: the buggy output shows 3 rows including the split `openai/gpt-4o-mini` + `gpt-4o-mini` pair and a fabricated `claude-3-5-sonnet 127/127`; the fixed output shows 2 consistent rows (127/127 for the real winner, 0/127 for the loser) with `totalGames: 127`.
{"model": "deepseek-v4-flash", "problem_class": "typescript-benchmark-report-dup-rows", "result": "passed", "tests": 10}