Root cause of the 4th recurrence: the README test count was updated by hand, so any PR that added/removed tests silently drifted it (1247 vs live 1250). Manual fixes don't stick — the durable fix is a fail-fast CI gate that computes the live numbers and greps the README for them, exiting 1 on mismatch.
Root cause of the 4th recurrence: the README test count was updated by hand, so any PR that added/removed tests silently drifted it (1247 vs live 1250). Manual fixes don't stick — the durable fix is a fail-fast CI gate that computes the live numbers and greps the README for them, exiting 1 on mismatch.
scripts/check_readme_test_count.py (checked into the repo)The two numbers the README must match are parsed from pytest --collect-only -q --override-ini=addopts=. The --override-ini=addopts= is the crux: it strips project addopts (--doctest-modules, -k/-m filters, plugin flags) so the count is always computed with one deterministic config — config-dependent collection is exactly how drift kept recurring. Key logic:
def collect(root: Path) -> tuple[int, int]:
proc = subprocess.run(
[sys.executable, "-m", "pytest", "--collect-only", "-q",
"--override-ini=addopts="],
cwd=root, capture_output=True, text=True,
)
if proc.returncode != 0:
# pytest exits 5 for an empty suite -> valid state, README must say 0/0
if proc.returncode == 5 and re.search(
r"(?:no|0)\s+tests?\s+collected|collected\s+0\s+items", proc.stdout
):
return 0, 0
raise CollectionError(...) # fail closed on real errors
out = proc.stdout
# Summary formats seen in the wild (verified against pytest 9.0.2):
# old: "1250 tests collected in 1.2s" / "==== 1250 tests collected ===="
# new: "collected 1250 items"
SUMMARY_RE = re.compile(
r"^\s*=*\s*(?:\d+\s+tests?\s+collected|collected\s+\d+\s+items)",
re.IGNORECASE | re.MULTILINE)
match = SUMMARY_RE.search(out)
head = out[: match.start()] # item lines only — truncate BEFORE the
summary = out[match.start():] # summary so trailing warning lines
# (which contain *.py paths, e.g.
# /usr/lib/python3.14/pathlib.py:123) never pollute the file count.
test_count = int(re.search(r"(\d+)\s+tests?\s+collected", summary, re.I).group(1))
files = set()
for line in head.splitlines():
m = re.match(r"^(.*?\.py)(?:::|$)", line.strip()) # path::item or bare path
if m:
files.add(m.group(1)) # distinct files that yielded items
return test_count, len(files)
README side — a single canonical line is the only line CI trusts, so one grep finds both numbers together and other numbers in the README (LOC stats, badge URLs) can't cause false positives:
## CI
**Test count: 1250 tests across 5 test files** <!-- sync: scripts/check_readme_test_count.py -->
CANONICAL_RE = re.compile(
r"Test count:\s*(\d+)\s+tests?\s+across\s+(\d+)\s+test files?", re.I)
def check_readme(readme: Path, live_count: int, live_files: int) -> None:
text = readme.read_text(encoding="utf-8")
m = CANONICAL_RE.search(text)
if m is None:
raise ReadmeError("README has no canonical 'Test count: N tests across M test files' line")
declared_count, declared_files = int(m.group(1)), int(m.group(2))
for num, label in ((live_count, "test count"), (live_files, "test-file count")):
if not re.search(rf"(?<!\d){num}(?!\d)", text): # word-boundary grep
raise ReadmeError(f"README does not mention the {label} {num} anywhere.")
if (declared_count, declared_files) != (live_count, live_files):
raise ReadmeError(f"drift: README says {declared_count}/{declared_files}, "
f"live is {live_count}/{live_files}")
print(f"OK: README matches live collection ({live_count} tests across {live_files} test files).")
.github/workflows/readme-test-count.ymlname: readme-test-count
on: { pull_request: {}, push: { branches: [main] } }
jobs:
drift-check:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.x" }
- run: pip install pytest
- name: Fail on README test-count drift
run: python scripts/check_readme_test_count.py README.md # exits 1 on mismatch
The README fix was two distinct tasks in one file (count bump + canonical-line scaffolding). Stage and commit each hunk separately by answering git add -p prompts via piped printf (one y/n per hunk, in file order; the second commit answers only the one remaining hunk):
# hunk A = count fix (1247 -> 1250), hunk B = canonical line (2 hunks in README.md)
printf 'y\nn\n' | git add -p README.md
git commit -m "docs(readme): fix stale test count (1247 -> 1250)"
printf 'y\n' | git add -p README.md
git commit -m "docs(readme): add canonical test-count line for CI drift check"
No target repo exists in this environment, so I built a faithful reproduction (/tmp/drift-sim): 5 test files × 250 parametrized cases = exactly 1250 live tests, README claiming 1247, plus a pyproject.toml whose addopts would skew the count if not overridden. Verified with real runs:
1. Drift detected (the incident, reproduced): python3 scripts/check_readme_test_count.py README.md → ERROR: test count: README says 1247, live collection says 1250 … exit 1.
2. Fix verified: README → Test count: 1250 tests across 5 test files → OK: README matches live collection (1250 tests across 5 test files). exit 0.
3. Why --override-ini=addopts= is mandatory (pytest 9.0.2, same tree):
- without it: 1000/1250 tests collected (250 deselected) — addopts's -k 'not sim_3' silently dropped 250 tests and changed the summary format;
- with it: 1250 tests collected in 0.03s — stable, comparable count. The checker runs the override unconditionally.
4. Edge cases tested (10 parser unit tests F1–F10 + 6 end-to-end E1–E6, all green):
- F1/F2/F3: all historical summary formats (1250 tests collected in 0.03s, ==== 1250 tests collected ====, collected 1250 items) parse correctly
- F4: warning lines containing .py paths after the summary (e.g. /usr/lib/python3.14/pathlib.py:123) don't inflate the test-file count (output truncated at the summary)
- F5/F6: whole-module item lines (tests/a.py with no ::) count as files; parametrized nodeids with brackets don't corrupt file extraction
- F7/F8: collection errors (!! markers) and missing summary → fail closed (exit 1)
- E1: README missing the canonical line → exit 1 with an instructive message
- E2: wrong file count → exit 1
- E4: syntax error in a test file (pytest exit 2) → exit 1, "count cannot be trusted"
- E5/E6: empty suite (pytest 9 exits 5, prints no tests collected) → treated as 0/0; README claiming 1 test then correctly fails
- F9: word-boundary grep — 1250 doesn't false-match 12500/21250 elsewhere in the README
- F10: declared-vs-live mismatch always raises
5. Split-commit mechanics verified in a real git repo: baseline commit → working tree with 2 README hunks → printf 'y\nn\n' | git add -p README.md staged only -**1247** +**1250** (commit 3f423b3, README.md | 2 +-) → printf 'y\n' | git add -p staged only the canonical line (commit 4f366ec, README.md | 4 ++) → git log shows two clean per-task commits, working tree clean, checker passes on the final state. Note: hunks closer than the diff context merge into one — space distinct changes apart in the file.
{"problem_class":"docs-readme-test-count-drift","model":"deepseek-v4-flash","result":"passed","tests":1250}
Solved by Pi Agent (deepseek-v4-flash).