typescript-cli-benchmark-rewire-post-poll
What was wrong: apps/cli/src/commands/benchmark.ts was orphaned — it printed a read-only/static report and emitted a stale warning citing the unfixed bug MAF-GAP-011 ("benchmark not wired up yet"). --games/--models were advertised but did nothing real.
The fix (full file at /tmp/benchfix/apps/cli/src/commands/benchmark.ts):
1. Start the run — POST /api/v1/benchmark with the exact config shape:
const startRes = await fetch(`${options.apiBaseUrl}/api/v1/benchmark`, {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({
config: {
models: options.models,
gamesPerPairing: options.gamesPerPairing,
},
}),
});
if (!startRes.ok) {
log(`error: failed to start benchmark run (HTTP ${startRes.status}): ${await safeText(startRes)}`);
return 1;
}
const runId = (await startRes.json()).runId;
if (!runId) { log("error: benchmark API did not return a runId"); return 1; }
2. Poll GET /api/v1/benchmark/runs/:runId every 2s with a 10-minute timeout, printing progress only on advance:
export const DEFAULT_POLL_INTERVAL_MS = 2_000;
export const DEFAULT_TIMEOUT_MS = 10 * 60 * 1_000;
let lastProgressKey = "";
while (true) {
if (deps.now() - startedAt >= options.timeoutMs) {
log(`error: benchmark run ${runId} did not finish within ${Math.round(options.timeoutMs / 1000)}s; giving up`);
return 1;
}
const state = await fetchRun(runId, options, deps); // GET /runs/:runId
if (state === null) return 1; // 404 / HTTP error logged
const key = progressKey(state); // "completed/total:games"
if (key !== "" && key !== lastProgressKey) { // print on advance only
log(`[${runId}] progress: ${state.progress!.completedPairings}/${state.progress!.totalPairings} pairings, ${state.progress!.completedGames}/${state.progress!.totalGames} games`);
lastProgressKey = key;
}
if (state.status === "COMPLETED") break;
if (state.status === "FAILED" || state.status === "CANCELLED") {
log(`error: benchmark run ${runId} ended with status ${state.status}${state.error ? `: ${state.error}` : ""}`);
return 1;
}
await deps.sleep(options.pollIntervalMs); // 2s default
}
3. On COMPLETED, fetch and display the report — GET /api/v1/benchmark/runs/:runId/report:
const report = await fetchReport(runId, options, deps);
log(renderReport(report)); // "benchmark report (run …)", per-pairing W/L/D with %
4. --models/--games are real parsed options (and the stale warning is gone):
// old: console.warn("⚠ Benchmark reports are read-only until MAF-GAP-011 is fixed…"); → REMOVED
if (arg === "--models") models = parseModels(value); // comma-separated, trimmed
else if (arg === "--games") gamesPerPairing = parseGames(value); // positive int, default 1
helpText() advertises --models/--games and no longer references MAF-GAP-011 or "read-only". Exit codes: 0 success, 1 runtime/API failure, 2 usage error.
5. Parse-level tests updated (benchmark.test.ts): option parsing (space/= forms, trimming, defaults, invalid values, missing value, unknown flags), help-text assertions (--models/--games present, MAF-GAP-011 absent), plus integration tests against a mock HTTP server.
Since the repo wasn't present, I built the command + tests at `/tmp/benchfix` and verified against a real in-process mock API (`POST /api/v1/benchmark`, state-advancing `GET /runs/:runId`, `GET /runs/:runId/report`).
**Verified behaviors (20/20 tests pass, `node --import tsx --test`, Node 22):**
- **POST body shape**: mock asserts `{config:{models:["model-a","model-b"],gamesPerPairing:10}}` exactly.
- **Poll cadence**: real CLI subprocess, default 2s interval — first poll immediate, second one interval later; measured ~2.18s elapsed for 2 polls; timeout constant verified as `600_000` (10 min).
- **Progress on advance**: exactly 2 distinct lines (`1/2 pairings, 10/20 games` → `2/2 pairings, 20/20 games`); identical progress is **not** reprinted.
- **Report on COMPLETED**: fetched `/runs/:runId/report` exactly once, rendered with percentages/draws.
- **Edge cases**: POST HTTP 500; POST response missing `runId`; run `FAILED` with error message; 404 while polling; 10-min timeout abort (`neverComplete` run); single-model (0 pairings) immediate completion; report fetch HTTP 500; `--games 0/-3/abc/1.5` rejected; missing `--models`; unknown flag; `--help` exits 0; `BENCHMARK_API_URL` env + `--api-url` flag honored.
- **Type safety**: strict `tsc` (`--strict`, NodeNext) — clean, exit 0.
**Real CLI run** (`--models deepseek-v3,claude-sonnet --games 40` against the mock):
```
benchmark run started: run-abc
[run-abc] progress: 1/2 pairings, 40/80 games
[run-abc] progress: 2/2 pairings, 80/80 games
benchmark report (run run-abc)
models: deepseek-v3, claude-sonnet
games per pairing: 40
deepseek-v3 vs claude-sonnet: 40 games — white 22 (55%), black 12 (30%), draws 6
claude-sonnet vs deepseek-v3: 40 games — white 18 (45%), black 15 (38%), draws 7
[exit code: 0]
```
No `MAF-GAP-011`/`read-only`/`stale` strings anywhere in output.{"model": "deepseek-v4-flash", "problem_class": "typescript-cli-benchmark-rewire-post-poll", "result": "passed", "tests": 20}