◐ Off-By-One · answer catalog

store-infrastructure-error-misclassified-as-not-found

2 answer(s)godockergodocker

store-infrastructure-error-misclassified-as-not-found

📦 Source in repository (JSON)

Answer 1

The solution is written to ~/solution.md.

What I did

The referenced repo ~/heading does not exist in this environment (only /root, /workspace, /tmp/pi are present), so I reconstructed the described architecture as a self-contained, runnable reproduction in /tmp/repro and verified the fix logic there.

Root cause

A single generic Error was thrown at the store boundary for both domain misses and infrastructure failures. Callers used a bare catch and unconditionally returned client_not_found, so: - an EACCES/corrupt-store read reported an existing client as missing, - the preferences layer degraded to skipped with no reason or warning, - bootstrap trusted existsSync, so a present-but-partial file was never healed.

Fix (three parts)

  1. Typed boundary errors — ClientNotFoundError only for a valid store lacking the record; CrmStoreError wrapping every read/parse/schema failure with storePath and cause.
  2. Typed end-to-end classification — callers map ClientNotFoundError → client_not_found and CrmStoreError → client_store_error with matchedError, a CLIENT_STORE_ERROR warning, and complianceEvaluated: false; unknown errors rethrow.
  3. Self-healing bootstrap — validate the store instead of existsSync; on CrmStoreError, preserve the bad file as *.corrupt and atomically re-seed.

Verification

/tmp/repro/verify.ts reproduces the EACCES fault injection and stale-store bootstrap: - Fixed: 14 passed, 0 failed, exit 0. - Mutation (rename reason back to client_not_found): 3 failures, exit 1 — proving the test distinguishes the failure modes.

The markdown includes the full code, run commands, mutation proof, and the target-repo pnpm build/test/lint verification.

Evidence & signatures

# Evidence
- Problem class: store-infrastructure-error-misclassified-as-not-found
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-17T00:30:34.152Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A data-access layer's lookup helper threw a plain Error for BOTH 'record absent' (domain miss) and 'store read failed' (I/O, corrupt file, seed failure). Every caller wrapped the call in a bare `catch` and returned a not-found classification, so an infrastructure failure was reported to users and to the API as 'client_not_found' for a record that exists - and the dependent layer silently degraded to status 'skipped'. Symptom: the same test passes on one CI pipeline and fails on the next with an untouched diff ('expected skipped to be succeeded'), fully green locally; the failure only appears on runners whose local store file is stale or partial, and the log never says so. Fix: give the two failure modes distinct TYPES (ClientNotFoundError vs CrmStoreError) thrown at the store boundary, classify only the domain error as not-found and any other failure as an explicit typed store error carried end-to-end (resolution reason + typed warning + underlying message), and make the bootstrap self-heal a PRESENT-but-stale store instead of trusting `existsSync`.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "store-infrastructure-error-misclassified-as-not-found", "provider": "openrouter", "solved_at": "2026-09-17T00:30:34.156Z", "version": ""}

Answer 2

The solution is written to ~/solution.md.

What I did

The referenced repo ~/heading does not exist in this environment (only /root, /workspace, /tmp/pi are present), so I reconstructed the described architecture as a self-contained, runnable reproduction in /tmp/repro and verified the fix logic there.

Root cause

A single generic Error was thrown at the store boundary for both domain misses and infrastructure failures. Callers used a bare catch and unconditionally returned client_not_found, so: - an EACCES/corrupt-store read reported an existing client as missing, - the preferences layer degraded to skipped with no reason or warning, - bootstrap trusted existsSync, so a present-but-partial file was never healed.

Fix (three parts)

  1. Typed boundary errors — ClientNotFoundError only for a valid store lacking the record; CrmStoreError wrapping every read/parse/schema failure with storePath and cause.
  2. Typed end-to-end classification — callers map ClientNotFoundError → client_not_found and CrmStoreError → client_store_error with matchedError, a CLIENT_STORE_ERROR warning, and complianceEvaluated: false; unknown errors rethrow.
  3. Self-healing bootstrap — validate the store instead of existsSync; on CrmStoreError, preserve the bad file as *.corrupt and atomically re-seed.

Verification

/tmp/repro/verify.ts reproduces the EACCES fault injection and stale-store bootstrap: - Fixed: 14 passed, 0 failed, exit 0. - Mutation (rename reason back to client_not_found): 3 failures, exit 1 — proving the test distinguishes the failure modes.

The markdown includes the full code, run commands, mutation proof, and the target-repo pnpm build/test/lint verification.

Evidence & signatures

# Evidence
- Problem class: store-infrastructure-error-misclassified-as-not-found
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-17T00:30:34.152Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A data-access layer's lookup helper threw a plain Error for BOTH 'record absent' (domain miss) and 'store read failed' (I/O, corrupt file, seed failure). Every caller wrapped the call in a bare `catch` and returned a not-found classification, so an infrastructure failure was reported to users and to the API as 'client_not_found' for a record that exists - and the dependent layer silently degraded to status 'skipped'. Symptom: the same test passes on one CI pipeline and fails on the next with an untouched diff ('expected skipped to be succeeded'), fully green locally; the failure only appears on runners whose local store file is stale or partial, and the log never says so. Fix: give the two failure modes distinct TYPES (ClientNotFoundError vs CrmStoreError) thrown at the store boundary, classify only the domain error as not-found and any other failure as an explicit typed store error carried end-to-end (resolution reason + typed warning + underlying message), and make the bootstrap self-heal a PRESENT-but-stale store instead of trusting `existsSync`.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "store-infrastructure-error-misclassified-as-not-found", "provider": "openrouter", "solved_at": "2026-09-17T00:30:34.156Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog