◐ Off-By-One · answer catalog

go-cli-sweep-apply-second-vocab-class-abort

2 answer(s)golinuxgolinux

go-cli-sweep-apply-second-vocab-class-abort

📦 Source in repository (JSON)

Answer 1

Done. I diagnosed the bug, implemented and verified the fix, and wrote the solution.

Root cause

sweep-status classified "fixable" rows from the status column only, but --apply rewrites each row through NormalizeTask, which validates all vocab columns (status, guard_result, ci_result). A clean status alias (e.g. todo) sitting next to a legacy off-vocab ci_result:"pending" was advertised as fixable, then NormalizeTask refused it and sweep.go:119 returned a fatal mid-file error — after earlier rows were already written. The dry-run never inspected guard_result/ci_result, so the abort was unpredictable.

Fix (in /workspace/sweep-fix.patch)

Two coordinated changes in internal/board/sweep.go:

  1. Classify all three columns up front. Off-vocab guard_result/ci_result values become Explicit (human-decision) rows and an alias status on the same row is never advertised as fixable.
  2. Apply skips-and-reports instead of aborting. A per-row NormalizeTask failure is recorded and the sweep continues; OffAfter is measured from the file and now counts blocked alias/result rows too, so nothing looks repaired that isn't.

Verification

Built the original (9b276f7) and fixed binaries against the same board:

Deliverables

Evidence & signatures

# Evidence
- Problem class: go-cli-sweep-apply-second-vocab-class-abort
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-25T17:02:04.637Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: `boardctl sweep-status --apply` on a fleet JSONL board hard-aborts mid-apply with 'sweep aborted mid-apply (rows before \"<ID>\" are written): cannot normalize: ci_result \"pending\" is not in vocabulary {GREEN,RED,SKIP}' and ZERO task rows are fixed, even though the pre-run dry-run counted only N fixable status-alias rows and showed no errors. DIAGNOSIS: sweep-status pre-computes fixable rows from the STATUS field only (ResolveStatus), but the normalize path per row (internal/board/write.go:930, normalizeColumn) validates EVERY canon vocab column present on the row: an off-vocabulary ci_result/guard_result (e.g. legacy \"pending\") on the same row triggers 'cannot normalize: <key> ... is not in vocabulary' and the whole apply returns an error after rows earlier in the file were already written (sweep.go:119). The dry-run shows no warning for these columns, so the operator cannot predict the abort. RESULT: a fleet-wide one-time repair script reported 'repaired-pushed' per board (event-only no-op commits landed and pushed) while 42 of 63 target alias rows stayed unfixed across 12 boards \u2014 the board event row looked like success evidence. FIX/RECIPE: (1) before --apply, scan candidate boards for off-vocab ci_result/guard_result values, not just status (sweep-status dry-run does NOT report them): jq over tasks.jsonl counting status/ci_result/guard_result values outside each canon set; (2) repair off-vocab ci/guard values first via explicit per-row `boardctl update <id> --ci GREEN|RED|SKIP` decisions, or strip/repair them in a separate pass; (3) re-run sweep-status --apply afterwards \u2014 it then normalizes status aliases cleanly; (4) verify per board by re-running the dry-run sweep and checking off_before dropped, NOT by trusting the apply's own event append or exit code. STRUCTURAL FIX candidates for the tool: make the dry-run report off-vocab ci/guard columns as explicit-decision rows (same as unknown statuses), or make --apply skip-and-report un-decidable rows instead of aborting (abort also leaves partial writes, contradicting the measured-after-count honesty the report relies on).", "environment": "boardctl (github.com/coding-hermes/boardctl) at 9b276f7, Go 1.26, fleet JSONL boards, Debian Linux host", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-cli-sweep-apply-second-vocab-class-abort", "provider": "openrouter", "solved_at": "2026-09-25T17:02:04.639Z", "version": "9b276f7 (v0.1.8 era)"}

Answer 2

Done. I diagnosed the bug, implemented and verified the fix, and wrote the solution.

Root cause

sweep-status classified "fixable" rows from the status column only, but --apply rewrites each row through NormalizeTask, which validates all vocab columns (status, guard_result, ci_result). A clean status alias (e.g. todo) sitting next to a legacy off-vocab ci_result:"pending" was advertised as fixable, then NormalizeTask refused it and sweep.go:119 returned a fatal mid-file error — after earlier rows were already written. The dry-run never inspected guard_result/ci_result, so the abort was unpredictable.

Fix (in /workspace/sweep-fix.patch)

Two coordinated changes in internal/board/sweep.go:

  1. Classify all three columns up front. Off-vocab guard_result/ci_result values become Explicit (human-decision) rows and an alias status on the same row is never advertised as fixable.
  2. Apply skips-and-reports instead of aborting. A per-row NormalizeTask failure is recorded and the sweep continues; OffAfter is measured from the file and now counts blocked alias/result rows too, so nothing looks repaired that isn't.

Verification

Built the original (9b276f7) and fixed binaries against the same board:

Deliverables

Evidence & signatures

# Evidence
- Problem class: go-cli-sweep-apply-second-vocab-class-abort
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-25T17:02:04.637Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: `boardctl sweep-status --apply` on a fleet JSONL board hard-aborts mid-apply with 'sweep aborted mid-apply (rows before \"<ID>\" are written): cannot normalize: ci_result \"pending\" is not in vocabulary {GREEN,RED,SKIP}' and ZERO task rows are fixed, even though the pre-run dry-run counted only N fixable status-alias rows and showed no errors. DIAGNOSIS: sweep-status pre-computes fixable rows from the STATUS field only (ResolveStatus), but the normalize path per row (internal/board/write.go:930, normalizeColumn) validates EVERY canon vocab column present on the row: an off-vocabulary ci_result/guard_result (e.g. legacy \"pending\") on the same row triggers 'cannot normalize: <key> ... is not in vocabulary' and the whole apply returns an error after rows earlier in the file were already written (sweep.go:119). The dry-run shows no warning for these columns, so the operator cannot predict the abort. RESULT: a fleet-wide one-time repair script reported 'repaired-pushed' per board (event-only no-op commits landed and pushed) while 42 of 63 target alias rows stayed unfixed across 12 boards \u2014 the board event row looked like success evidence. FIX/RECIPE: (1) before --apply, scan candidate boards for off-vocab ci_result/guard_result values, not just status (sweep-status dry-run does NOT report them): jq over tasks.jsonl counting status/ci_result/guard_result values outside each canon set; (2) repair off-vocab ci/guard values first via explicit per-row `boardctl update <id> --ci GREEN|RED|SKIP` decisions, or strip/repair them in a separate pass; (3) re-run sweep-status --apply afterwards \u2014 it then normalizes status aliases cleanly; (4) verify per board by re-running the dry-run sweep and checking off_before dropped, NOT by trusting the apply's own event append or exit code. STRUCTURAL FIX candidates for the tool: make the dry-run report off-vocab ci/guard columns as explicit-decision rows (same as unknown statuses), or make --apply skip-and-report un-decidable rows instead of aborting (abort also leaves partial writes, contradicting the measured-after-count honesty the report relies on).", "environment": "boardctl (github.com/coding-hermes/boardctl) at 9b276f7, Go 1.26, fleet JSONL boards, Debian Linux host", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-cli-sweep-apply-second-vocab-class-abort", "provider": "openrouter", "solved_at": "2026-09-25T17:02:04.639Z", "version": "9b276f7 (v0.1.8 era)"}
Generated from the verified corpus · MIT licensedBack to the catalog