typescript-stats-aggregation-key-normalization
The bug had three independent causes, fixed in stats-aggregation.ts (full file at /tmp/mafia-stats/stats-aggregation.ts):
1. Key normalization — rows carried a provider prefix on the model column, so keying on raw provider/model produced openai/openai/gpt-4o-mini that never merged with openai/gpt-4o-mini:
// strip a "<provider>/" prefix before GROUP BY/merge
export function normalizeModelKey(
model: string,
provider?: string,
knownProviders: readonly string[] = DEFAULT_PROVIDER_PREFIXES,
): string {
const m = (model ?? "").trim();
if (!m) return m;
const p = (provider ?? "").trim();
// 1) exact provider prefix, repeated (collapses double-writes too)
if (p) {
const prefix = `${p.toLowerCase()}/`;
if (m.toLowerCase().startsWith(prefix)) {
let out = m;
while (out.length > p.length && out.toLowerCase().startsWith(prefix))
out = out.slice(p.length + 1);
return out;
}
}
// 2) any observed provider prefix (provider column missing / spelled differently)
for (const k of knownProviders) {
const prefix = `${k.toLowerCase()}/`;
if (m.toLowerCase().startsWith(prefix)) {
let out = m;
while (out.length > k.length && out.toLowerCase().startsWith(prefix))
out = out.slice(k.length + 1);
return out;
}
}
return m;
}
The key becomes provider.toLowerCase() + "/" + normalizeModelKey(model, provider, known) — so openai/gpt-4o-mini and gpt-4o-mini (both provider=openai) collapse onto openai/gpt-4o-mini, and missing-provider rows infer the provider from the model prefix:
export function inferProvider(row: PlayerRow, knownProviders: readonly string[]): string {
const p = (row.provider ?? "").trim();
if (p) return p;
const lower = (row.model ?? "").trim().toLowerCase();
for (const k of knownProviders)
if (lower.startsWith(`${k.toLowerCase()}/`)) return k;
return "unknown";
}
2. Wins only from real per-player rows — the game-level-winner fallback is deleted. buildReport never receives the games table (the SQL drops the winner join), and wins are counted strictly from players.won === 1:
// SQL feeding the report — the games table / winner column are NOT selected:
// SELECT game_id, player_id, provider, model, role, won,
// tokens_in, tokens_out, cost_usd, latency_ms
// FROM players ORDER BY id;
const won = row.won === 1 ? 1 : 0; // the ONLY source of wins
r.wins += won;
r.winRate = r.games > 0 ? r.wins / r.games : 0;
3. First/most-complete source on key collision — numeric metrics are always summed (nothing dropped); display metadata keeps the first source unless an incoming row is strictly more complete:
const sourceScoreBy = new WeakMap<ReportRow, number>();
function mergeRow(r: ReportRow, row: PlayerRow, knownProviders: readonly string[]): void {
// ... r.games++, r.wins += won, metrics += ... (never dropped)
const incoming = completenessScore(row); // count of populated fields
if (incoming > (sourceScoreBy.get(r) ?? 0)) {
r.model = normalizeModelKey(row.model, row.provider, knownProviders);
r.provider = inferProvider(row, knownProviders);
sourceScoreBy.set(r, incoming); // strictly more complete → replaces
}
}
knownProviders is data-driven (collectProviders: static safe list ∪ providers observed in the dataset), so a legitimately-slashed id like meta-llama/llama-3.1-70b is never corrupted unless meta-llama is actually a provider.
Ran with `node --experimental-strip-types evidence.test.ts` (Node 22) + `tsc --strict`: **25/25 tests pass, type-check clean** (`/tmp/mafia-stats/`). **Fixture reproduces the report exactly** — 5 real models; `gpt-4o-mini` written three ways (prefixed, bare, prefixed+missing-provider); `omni-everywhere` plays every role in 127 games, its side wins all 127, but it has only **73 real** `players.won=1` rows; `gpt-4o` includes 3 draw games (`sideWon=null`); `meta-llama/llama-3.1-70b` has a legit internal slash; one empty-model row. **Buggy (shipped) aggregator reproduces the reported failures:** ``` vllm vllm/vllm/omni-everywhere games=127 wins=127 winRate=1.0000 ← fabricated 127/127 openai openai/openai/gpt-4o-mini games= 10 wins= 0 ← duplicate of openai openai/gpt-4o-mini games= 18 wins= 0 ← gpt-4o-mini undefined undefined/openai/gpt-4o-mini games= 3 wins= 0 ← and a third ``` **Fixed aggregator (live):** ``` vllm omni-everywhere games=127 wins= 73 winRate=0.5748 tokIn=33401 tokOut=27051 cost=$1.054 lat=43191ms openai gpt-4o-mini games= 31 wins= 19 winRate=0.6129 tokIn= 3100 tokOut= 2480 cost=$0.031 lat= 7750ms anthropic claude-3-5-sonnet games= 25 wins= 15 winRate=0.6000 tokIn=12800 tokOut=10300 cost=$0.550 lat=22800ms openai gpt-4o games= 20 wins= 12 winRate=0.6000 tokIn= 2000 tokOut= 1600 cost=$0.020 lat= 5000ms together meta-llama/llama-3.1-70b games= 18 wins= 9 winRate=0.5000 tokIn= 1800 tokOut= 1440 cost=$0.018 lat= 4500ms ``` → 5 rows, 0 duplicate model strings, no winRate exactly 1.0, tokens/cost/latency sums identical to the raw rows (nothing dropped by merging). **Edge cases tested (25 assertions):** - Prefix strip: exact provider, missing provider column, mismatched provider spelling (`openai-compatible`), double-write `openai/openai/gpt-4o-mini`, case-insensitive (`OPENAI/`), whitespace trim, empty model skipped (no phantom row) - `meta-llama/llama-3.1-70b` preserved verbatim; stripped only when `meta-llama` is the actual provider - gpt-4o-mini's three write-styles merge into one row (31 plays / 19 wins / sources=31) - Draw games (`sideWon=null`) never credit wins (gpt-4o stays 12 wins, not 15) - Collision policy: `OpenAI` vs `openai` providers don't split a model; first most-complete source's metadata kept - Metric preservation: Σ report rows == Σ player rows for tokensIn/tokensOut/costUsd/latencyMs, games, wins
{"model": "deepseek-v4-flash", "problem_class": "typescript-stats-aggregation-key-normalization", "result": "passed", "tests": 25}