model-catalog-freshness-verification
Diagnosis complete. I built a runnable reference implementation and verified it offline against a synthetic registry that reproduces every trap in the problem class.
The freshness pass had five independent defects, all stemming from one structural fact: model identity is denormalized across seed files, and the job only understood the catalog.
/models is curated; zai-glm lists 10 ids but serves omitted ids at HTTP 200. Omission was treated as a verdict instead of a probe trigger.valid_to.benchmarks.jsonl (seed for model_perf/model_tier), PROFILE_MODELS in router_seed.py, and fallback_lanes.jsonl were not — so the renamed model silently lost tier categories and dropped out of P1_CODING.context_limit/vision/thinking were only filled on new rows.router_seed.py regenerates model_tier.jsonl/model_perf.jsonl, wiping manual fixes.solution/catalog_freshness.py implements:
- Fetch models.dev/api.json + {base_url}/models per enabled lane.
- Retire only on probe 404/410 with a real max_tokens=4 chat call; 401/403/429/5xx/transport all keep the model.
- Set valid_to (never delete), atomic writes, append-only catalog_probe_evidence.jsonl.
- Renames applied before probing (no false retirement), write model_aliases.jsonl, propagate to benchmarks.jsonl, fallback_lanes.jsonl, and the PROFILE_MODELS block.
- Capability refresh on existing rows: fill-when-null, overwrite-on-diff, retired rows frozen, cost_* never touched.
solution/test_catalog_freshness.py builds a synthetic registry with omitted-but-alive ids (probe 200), a real 404, a rename, a provider rename, stale/wrong capabilities, a frozen retired row, and a conflicting price. Result:
done: retired=2 aliases=1
ALL ASSERTIONS PASSED
It confirms 11 invariants including: omitted-but-alive kept, dead-model retired only via chat probe 404, no row deleted, vision False → True overwritten, cost_input stayed 1.5 despite models.dev 99.0, alias + benchmarks + fallback lanes + PROFILE_MODELS all propagated.
Artifacts
- ~/SOLUTION.md — full self-contained write-up (root cause, code, verification, production checklist, guardrails)
- ~/solution/catalog_freshness.py — runner
- ~/solution/test_catalog_freshness.py — offline end-to-end verifier
Run it yourself: cd ~/solution && python3 test_catalog_freshness.py
# Evidence - Problem class: model-catalog-freshness-verification - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-11T02:53:53.715Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "TR-035 weekly research lane 1/3 (model-catalog freshness) on the task-router registry. Approach that worked: (1) fetch https://models.dev/api.json AND GET {base_url}/models for every enabled lane in data/tables/probe_providers.jsonl; (2) NEVER retire a model on listing omission alone \u2014 the live listing is NOT the routable set (zai-glm's coding endpoint lists 10 ids but still serves glm-4.5-flash/4.5v/4.6v/4.7-flashx at HTTP 200; 24 of 48 omitted ids were alive). Double-check each candidate with a direct chat probe (max_tokens=4) and only retire on 404/410. (3) Set valid_to (never delete); add a model_aliases.jsonl row for renamed ids. (4) A provider rename must propagate to benchmarks.jsonl (the SEED INPUT for model_perf/model_tier), router_seed.py PROFILE_MODELS family maps and fallback_lanes.jsonl \u2014 otherwise the renamed model silently loses tier categories and drops out of profiles (P1_CODING test>=0). (5) models.dev sync only filled capability fields on NEW rows \u2014 make it refresh context_limit/vision/thinking on existing rows (fill-when-null, overwrite-on-diff, retired rows frozen, never touch price fields). Watch out: router_seed.py regenerates data/tables/*.jsonl, so edits to DERIVED tables (model_tier/model_perf) are wiped \u2014 fix the seed input instead.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "model-catalog-freshness-verification", "provider": "openrouter", "solved_at": "2026-09-11T02:53:53.715Z", "version": ""}Diagnosis complete. I built a runnable reference implementation and verified it offline against a synthetic registry that reproduces every trap in the problem class.
The freshness pass had five independent defects, all stemming from one structural fact: model identity is denormalized across seed files, and the job only understood the catalog.
/models is curated; zai-glm lists 10 ids but serves omitted ids at HTTP 200. Omission was treated as a verdict instead of a probe trigger.valid_to.benchmarks.jsonl (seed for model_perf/model_tier), PROFILE_MODELS in router_seed.py, and fallback_lanes.jsonl were not — so the renamed model silently lost tier categories and dropped out of P1_CODING.context_limit/vision/thinking were only filled on new rows.router_seed.py regenerates model_tier.jsonl/model_perf.jsonl, wiping manual fixes.solution/catalog_freshness.py implements:
- Fetch models.dev/api.json + {base_url}/models per enabled lane.
- Retire only on probe 404/410 with a real max_tokens=4 chat call; 401/403/429/5xx/transport all keep the model.
- Set valid_to (never delete), atomic writes, append-only catalog_probe_evidence.jsonl.
- Renames applied before probing (no false retirement), write model_aliases.jsonl, propagate to benchmarks.jsonl, fallback_lanes.jsonl, and the PROFILE_MODELS block.
- Capability refresh on existing rows: fill-when-null, overwrite-on-diff, retired rows frozen, cost_* never touched.
solution/test_catalog_freshness.py builds a synthetic registry with omitted-but-alive ids (probe 200), a real 404, a rename, a provider rename, stale/wrong capabilities, a frozen retired row, and a conflicting price. Result:
done: retired=2 aliases=1
ALL ASSERTIONS PASSED
It confirms 11 invariants including: omitted-but-alive kept, dead-model retired only via chat probe 404, no row deleted, vision False → True overwritten, cost_input stayed 1.5 despite models.dev 99.0, alias + benchmarks + fallback lanes + PROFILE_MODELS all propagated.
Artifacts
- ~/SOLUTION.md — full self-contained write-up (root cause, code, verification, production checklist, guardrails)
- ~/solution/catalog_freshness.py — runner
- ~/solution/test_catalog_freshness.py — offline end-to-end verifier
Run it yourself: cd ~/solution && python3 test_catalog_freshness.py
# Evidence - Problem class: model-catalog-freshness-verification - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-11T02:53:53.715Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "TR-035 weekly research lane 1/3 (model-catalog freshness) on the task-router registry. Approach that worked: (1) fetch https://models.dev/api.json AND GET {base_url}/models for every enabled lane in data/tables/probe_providers.jsonl; (2) NEVER retire a model on listing omission alone \u2014 the live listing is NOT the routable set (zai-glm's coding endpoint lists 10 ids but still serves glm-4.5-flash/4.5v/4.6v/4.7-flashx at HTTP 200; 24 of 48 omitted ids were alive). Double-check each candidate with a direct chat probe (max_tokens=4) and only retire on 404/410. (3) Set valid_to (never delete); add a model_aliases.jsonl row for renamed ids. (4) A provider rename must propagate to benchmarks.jsonl (the SEED INPUT for model_perf/model_tier), router_seed.py PROFILE_MODELS family maps and fallback_lanes.jsonl \u2014 otherwise the renamed model silently loses tier categories and drops out of profiles (P1_CODING test>=0). (5) models.dev sync only filled capability fields on NEW rows \u2014 make it refresh context_limit/vision/thinking on existing rows (fill-when-null, overwrite-on-diff, retired rows frozen, never touch price fields). Watch out: router_seed.py regenerates data/tables/*.jsonl, so edits to DERIVED tables (model_tier/model_perf) are wiped \u2014 fix the seed input instead.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "model-catalog-freshness-verification", "provider": "openrouter", "solved_at": "2026-09-11T02:53:53.715Z", "version": ""}