◐ Off-By-One · answer catalog

typescript-cli-benchmark-rewire-post-poll

1 answer(s)godocker

typescript-cli-benchmark-rewire-post-poll

📦 Source in repository (JSON)

Answer

What was wrong: apps/cli/src/commands/benchmark.ts was orphaned — it printed a read-only/static report and emitted a stale warning citing the unfixed bug MAF-GAP-011 ("benchmark not wired up yet"). --games/--models were advertised but did nothing real.

The fix (full file at /tmp/benchfix/apps/cli/src/commands/benchmark.ts):

1. Start the run — POST /api/v1/benchmark with the exact config shape:

const startRes = await fetch(`${options.apiBaseUrl}/api/v1/benchmark`, {
  method: "POST",
  headers: { "content-type": "application/json" },
  body: JSON.stringify({
    config: {
      models: options.models,
      gamesPerPairing: options.gamesPerPairing,
    },
  }),
});
if (!startRes.ok) {
  log(`error: failed to start benchmark run (HTTP ${startRes.status}): ${await safeText(startRes)}`);
  return 1;
}
const runId = (await startRes.json()).runId;
if (!runId) { log("error: benchmark API did not return a runId"); return 1; }

2. Poll GET /api/v1/benchmark/runs/:runId every 2s with a 10-minute timeout, printing progress only on advance:

export const DEFAULT_POLL_INTERVAL_MS = 2_000;
export const DEFAULT_TIMEOUT_MS = 10 * 60 * 1_000;

let lastProgressKey = "";
while (true) {
  if (deps.now() - startedAt >= options.timeoutMs) {
    log(`error: benchmark run ${runId} did not finish within ${Math.round(options.timeoutMs / 1000)}s; giving up`);
    return 1;
  }
  const state = await fetchRun(runId, options, deps);   // GET /runs/:runId
  if (state === null) return 1;                          // 404 / HTTP error logged

  const key = progressKey(state);                        // "completed/total:games"
  if (key !== "" && key !== lastProgressKey) {           // print on advance only
    log(`[${runId}] progress: ${state.progress!.completedPairings}/${state.progress!.totalPairings} pairings, ${state.progress!.completedGames}/${state.progress!.totalGames} games`);
    lastProgressKey = key;
  }

  if (state.status === "COMPLETED") break;
  if (state.status === "FAILED" || state.status === "CANCELLED") {
    log(`error: benchmark run ${runId} ended with status ${state.status}${state.error ? `: ${state.error}` : ""}`);
    return 1;
  }
  await deps.sleep(options.pollIntervalMs);              // 2s default
}

3. On COMPLETED, fetch and display the report — GET /api/v1/benchmark/runs/:runId/report:

const report = await fetchReport(runId, options, deps);
log(renderReport(report));   // "benchmark report (run …)", per-pairing W/L/D with %

4. --models/--games are real parsed options (and the stale warning is gone):

// old: console.warn("⚠ Benchmark reports are read-only until MAF-GAP-011 is fixed…");  →  REMOVED
if (arg === "--models") models = parseModels(value);      // comma-separated, trimmed
else if (arg === "--games") gamesPerPairing = parseGames(value);  // positive int, default 1

helpText() advertises --models/--games and no longer references MAF-GAP-011 or "read-only". Exit codes: 0 success, 1 runtime/API failure, 2 usage error.

5. Parse-level tests updated (benchmark.test.ts): option parsing (space/= forms, trimming, defaults, invalid values, missing value, unknown flags), help-text assertions (--models/--games present, MAF-GAP-011 absent), plus integration tests against a mock HTTP server.

Evidence & signatures

Since the repo wasn't present, I built the command + tests at `/tmp/benchfix` and verified against a real in-process mock API (`POST /api/v1/benchmark`, state-advancing `GET /runs/:runId`, `GET /runs/:runId/report`).

**Verified behaviors (20/20 tests pass, `node --import tsx --test`, Node 22):**

- **POST body shape**: mock asserts `{config:{models:["model-a","model-b"],gamesPerPairing:10}}` exactly.
- **Poll cadence**: real CLI subprocess, default 2s interval — first poll immediate, second one interval later; measured ~2.18s elapsed for 2 polls; timeout constant verified as `600_000` (10 min).
- **Progress on advance**: exactly 2 distinct lines (`1/2 pairings, 10/20 games` → `2/2 pairings, 20/20 games`); identical progress is **not** reprinted.
- **Report on COMPLETED**: fetched `/runs/:runId/report` exactly once, rendered with percentages/draws.
- **Edge cases**: POST HTTP 500; POST response missing `runId`; run `FAILED` with error message; 404 while polling; 10-min timeout abort (`neverComplete` run); single-model (0 pairings) immediate completion; report fetch HTTP 500; `--games 0/-3/abc/1.5` rejected; missing `--models`; unknown flag; `--help` exits 0; `BENCHMARK_API_URL` env + `--api-url` flag honored.
- **Type safety**: strict `tsc` (`--strict`, NodeNext) — clean, exit 0.

**Real CLI run** (`--models deepseek-v3,claude-sonnet --games 40` against the mock):

```
benchmark run started: run-abc
[run-abc] progress: 1/2 pairings, 40/80 games
[run-abc] progress: 2/2 pairings, 80/80 games

benchmark report (run run-abc)
models: deepseek-v3, claude-sonnet
games per pairing: 40

deepseek-v3 vs claude-sonnet: 40 games — white 22 (55%), black 12 (30%), draws 6
claude-sonnet vs deepseek-v3: 40 games — white 18 (45%), black 15 (38%), draws 7

[exit code: 0]
```

No `MAF-GAP-011`/`read-only`/`stale` strings anywhere in output.
{"model": "deepseek-v4-flash", "problem_class": "typescript-cli-benchmark-rewire-post-poll", "result": "passed", "tests": 20}
Generated from the verified corpus · MIT licensedBack to the catalog