◐ Off-By-One · answer catalog

typescript-benchmark-report-dup-rows

2 answer(s)typescriptdockertypescriptdocker

typescript-benchmark-report-dup-rows

📦 Source in repository (JSON)

Answer 1

Root cause (two independent defects):

  1. Raw-key aggregation splits one model into multiple rows. The model column is sometimes prefixed (openai/gpt-4o-mini) and sometimes bare (gpt-4o-mini). Grouping on the raw string treats them as two different models, so a single model appears as duplicate rows with contradictory totals.
  2. Legacy fallback fabricates winRate=1.0 rows. When a game had a winner, the old code credited that win to every assigned model. Over 127 games that produced CUSTOM/openai 127/127 (games=127, wins=127, winRate=1.0) for models that never won, contradicting the real aggregated rates.

The fix (src/benchmark-report.ts): - Normalize before grouping — every model id is parsed into a canonical provider/name key (normalizeModel). openai/gpt-4o-mini and gpt-4o-mini (with provider: "openai") both collapse to key openai/gpt-4o-mini, so aggregation merges them into a single row. Record-level provider takes precedence over a prefix. - Delete the blanket win attribution — only the model whose id equals the game's normalized winner gets a wins++. Ties/no-winner games contribute games but zero wins. Wins are also never credited to a model that wasn't assigned.

// src/benchmark-report.ts (core of the fix)

/** "openai/gpt-4o-mini" -> { provider: "openai", name: "gpt-4o-mini", key: "openai/gpt-4o-mini" } */
export function normalizeModel(raw: string, recordProvider?: string | null): NormalizedModel {
  const trimmed = raw.trim();
  let provider = recordProvider?.trim() || null;
  let name = trimmed;
  const slash = trimmed.indexOf("/");
  if (slash > 0) {
    provider ??= trimmed.slice(0, slash);   // prefix is a provider only when unknown
    name = trimmed.slice(slash + 1);
  }
  return { provider, name, key: provider ? `${provider}/${name}` : name };
}

// inside buildReport(records):
for (const rec of records) {
  if (seenIds.has(rec.id)) continue;              // dedupe duplicate game rows
  seenIds.add(rec.id);
  const assigned = new Map<string, NormalizedModel>();
  for (const m of rec.models ?? []) {
    const norm = normalizeModel(m, rec.provider); // canonical key -> no split rows
    if (!assigned.has(norm.key)) assigned.set(norm.key, norm);
  }
  if (assigned.size === 0) continue;
  let winnerKey: string | null = null;
  if (rec.winner?.trim()) winnerKey = normalizeModel(rec.winner, rec.provider).key;
  for (const [key, norm] of assigned) {
    // games++ once per (game, model)
    // wins++ ONLY when winnerKey === key  <-- legacy fallback removed
  }
}

Side-by-side on the problem scenario (127 games, models ["openai/gpt-4o-mini","gpt-4o-mini","claude-3-5-sonnet"], winner always gpt-4o-mini):

BUGGY (before) FIXED (after)
openai/gpt-4o-mini 127/127 winRate 1.0 gpt-4o-mini (openai) 127/127 winRate 1.0 ← real winner
gpt-4o-mini 127/127 winRate 1.0 (split dup row) merged into the row above
claude-3-5-sonnet 127/127 winRate 1.0 (fabricated) claude-3-5-sonnet (openai) 127 games, 0 wins, winRate 0

Evidence & signatures

Verified by running the suite under the target runtime (`node v22.22.3` + `tsx v4.23.12`, matching the docker-compose node22 / typescript-node22-tsx environment):

```
# tests 10   # pass 10   # fail 0
```

Files: `src/benchmark-report.ts` (fix), `test/benchmark-report.test.ts` (10 tests via `node:test`).

Edge cases tested:
- **Provider-prefix merge** — prefixed and bare spellings of the same model collapse into exactly one row with merged games/wins.
- **No fabricated 127/127** — the exact problem scenario: the non-winning assigned model reports 127 games / 0 wins / winRate 0; exactly 2 rows, no split rows.
- **Consistency invariant** — one win per decisive game, never more; total row wins === number of decisive games.
- **Record-level provider precedence** — `provider:"CUSTOM"` + `model:"openai/gpt-4o-mini"` yields `CUSTOM/gpt-4o-mini`; the winner normalized the same way still matches.
- **Duplicate game ids** in the input counted once; **duplicate entries within one game's model list** counted once.
- **Winner not among assigned models** → no random model is credited (0 wins, defensive).
- **Tie / no winner** → zero wins, winRate 0; **empty input** → empty report with `totalGames: 0`.

The demo run above prints the buggy vs fixed rows for the 127-game scenario: the buggy output shows 3 rows including the split `openai/gpt-4o-mini` + `gpt-4o-mini` pair and a fabricated `claude-3-5-sonnet 127/127`; the fixed output shows 2 consistent rows (127/127 for the real winner, 0/127 for the loser) with `totalGames: 127`.
{"model": "deepseek-v4-flash", "problem_class": "typescript-benchmark-report-dup-rows", "result": "passed", "tests": 10}

Answer 2

Root cause (two independent defects):

  1. Raw-key aggregation splits one model into multiple rows. The model column is sometimes prefixed (openai/gpt-4o-mini) and sometimes bare (gpt-4o-mini). Grouping on the raw string treats them as two different models, so a single model appears as duplicate rows with contradictory totals.
  2. Legacy fallback fabricates winRate=1.0 rows. When a game had a winner, the old code credited that win to every assigned model. Over 127 games that produced CUSTOM/openai 127/127 (games=127, wins=127, winRate=1.0) for models that never won, contradicting the real aggregated rates.

The fix (src/benchmark-report.ts): - Normalize before grouping — every model id is parsed into a canonical provider/name key (normalizeModel). openai/gpt-4o-mini and gpt-4o-mini (with provider: "openai") both collapse to key openai/gpt-4o-mini, so aggregation merges them into a single row. Record-level provider takes precedence over a prefix. - Delete the blanket win attribution — only the model whose id equals the game's normalized winner gets a wins++. Ties/no-winner games contribute games but zero wins. Wins are also never credited to a model that wasn't assigned.

// src/benchmark-report.ts (core of the fix)

/** "openai/gpt-4o-mini" -> { provider: "openai", name: "gpt-4o-mini", key: "openai/gpt-4o-mini" } */
export function normalizeModel(raw: string, recordProvider?: string | null): NormalizedModel {
  const trimmed = raw.trim();
  let provider = recordProvider?.trim() || null;
  let name = trimmed;
  const slash = trimmed.indexOf("/");
  if (slash > 0) {
    provider ??= trimmed.slice(0, slash);   // prefix is a provider only when unknown
    name = trimmed.slice(slash + 1);
  }
  return { provider, name, key: provider ? `${provider}/${name}` : name };
}

// inside buildReport(records):
for (const rec of records) {
  if (seenIds.has(rec.id)) continue;              // dedupe duplicate game rows
  seenIds.add(rec.id);
  const assigned = new Map<string, NormalizedModel>();
  for (const m of rec.models ?? []) {
    const norm = normalizeModel(m, rec.provider); // canonical key -> no split rows
    if (!assigned.has(norm.key)) assigned.set(norm.key, norm);
  }
  if (assigned.size === 0) continue;
  let winnerKey: string | null = null;
  if (rec.winner?.trim()) winnerKey = normalizeModel(rec.winner, rec.provider).key;
  for (const [key, norm] of assigned) {
    // games++ once per (game, model)
    // wins++ ONLY when winnerKey === key  <-- legacy fallback removed
  }
}

Side-by-side on the problem scenario (127 games, models ["openai/gpt-4o-mini","gpt-4o-mini","claude-3-5-sonnet"], winner always gpt-4o-mini):

BUGGY (before) FIXED (after)
openai/gpt-4o-mini 127/127 winRate 1.0 gpt-4o-mini (openai) 127/127 winRate 1.0 ← real winner
gpt-4o-mini 127/127 winRate 1.0 (split dup row) merged into the row above
claude-3-5-sonnet 127/127 winRate 1.0 (fabricated) claude-3-5-sonnet (openai) 127 games, 0 wins, winRate 0

Evidence & signatures

Verified by running the suite under the target runtime (`node v22.22.3` + `tsx v4.23.12`, matching the docker-compose node22 / typescript-node22-tsx environment):

```
# tests 10   # pass 10   # fail 0
```

Files: `src/benchmark-report.ts` (fix), `test/benchmark-report.test.ts` (10 tests via `node:test`).

Edge cases tested:
- **Provider-prefix merge** — prefixed and bare spellings of the same model collapse into exactly one row with merged games/wins.
- **No fabricated 127/127** — the exact problem scenario: the non-winning assigned model reports 127 games / 0 wins / winRate 0; exactly 2 rows, no split rows.
- **Consistency invariant** — one win per decisive game, never more; total row wins === number of decisive games.
- **Record-level provider precedence** — `provider:"CUSTOM"` + `model:"openai/gpt-4o-mini"` yields `CUSTOM/gpt-4o-mini`; the winner normalized the same way still matches.
- **Duplicate game ids** in the input counted once; **duplicate entries within one game's model list** counted once.
- **Winner not among assigned models** → no random model is credited (0 wins, defensive).
- **Tie / no winner** → zero wins, winRate 0; **empty input** → empty report with `totalGames: 0`.

The demo run above prints the buggy vs fixed rows for the 127-game scenario: the buggy output shows 3 rows including the split `openai/gpt-4o-mini` + `gpt-4o-mini` pair and a fabricated `claude-3-5-sonnet 127/127`; the fixed output shows 2 consistent rows (127/127 for the real winner, 0/127 for the loser) with `totalGames: 127`.
{"model": "deepseek-v4-flash", "problem_class": "typescript-benchmark-report-dup-rows", "result": "passed", "tests": 10}
Generated from the verified corpus · MIT licensedBack to the catalog