typescript-regex-normalized-form-false-negatives
Root cause. Every inbound message is normalized before pattern matching (NFD strips ñ→n/á→a; apostrophe whitespace collapses I 'm → i'm), but the crisis patterns were written against the pre-normalization text, so they silently never matched the normalized form → false negatives.
Fix 1 — include normalized variants in the regex alternation (src/patterns.ts):
// BUG (before): only the pre-normalization ñ form — NFD output "dano" never matches
regex: /hacer daño/
// FIX (after): alternation covers raw "daño" AND normalized "dano"
regex: /hacer (?:daño|dano)/
regex: /hacerme (?:daño|dano)/
Fix 2 — match an optional apostrophe directly after i (the old pattern required a space, i 'm, but normalization collapses it):
// BUG (before): requires space between i and apostrophe — normalized "i'm ..." never matches
regex: /i 'm going to (?:kill|end) (?:myself|it)/
// FIX (after): optional apostrophe directly after i — matches "i'm", "im", "i ’m"
regex: /i'?m going to (?:kill|end) (?:myself|it|everything)/
regex: /i'?m so (?:done|tired)/
The normalizer (src/normalize.ts) stays the source of truth: NFD → strip combining marks → unify ’‘→' → collapse \s*'\s* → lowercase. Supporting files: src/classifier.ts (match/classify), src/labeled-corpus.ts (20 labeled items), src/debug-classifier.ts (per-item debug CLI), src/debug-classifier.test.ts (node:test suite).
Verified in `~/crisis-classifier` with TypeScript 5.9 / Node 22: **Fixed patterns — `node dist/debug-classifier.js` (exit 0):** - corpus: 20 items → **detection 13/13 = 100.0%** (crisis recall, 0 false negatives), negatives 7/7 = 100.0% (0 false positives), overall 20/20 = 100.0% → `✅ PASS — 100% detection on labeled corpus` - Normalized form is printed per item, e.g. `daño → dano`, `I 'm going → i'm going`, `I’m → i'm`, `IM → im` **Legacy (pre-fix) patterns — `node dist/debug-classifier.js --legacy` (exit 1, the bug):** - detection **3/13 = 23.1%** — 10 false negatives, exactly the reported classes: all 4 `ñ` items (es.hacer-dano-original, es.hacer-dano-todos, es.dano-sin-tilde-raw, es.voy-a-hacerme-dano) and all 6 apostrophe items (en.i-m-kill, en.i-m-end-space, en.i-m-curly, en.i-m-done, en.i-m-tired-space, en.im-naked) **Unit/regression tests — `node --test`:** **28/28 pass, 0 fail**, covering: NFD `ñ→n`, `I 'm→i'm` collapse, curly U+2019 unification, other diacritics (`¿cómo estás?→¿como estas?`), every corpus item, alternation both-forms (`daño` + pre-stripped `dano`), optional apostrophe (`i'm`/`im`/`i ’m`), and a bug-demo test asserting the legacy patterns still miss the normalized forms. **Edge cases tested:** raw text already diacritic-free (`hacerme dano` typed without ñ), curly apostrophe (`I’m going to end everything`), no apostrophe (`IM going to kill myself`), space-before-apostrophe raw input (`I 'm going to end it all`), and precision negatives that must NOT flag: `hacer la cena`, `I am going to the store`, `She's going…`, `I don't want to die`, `I am tired of work today`.
{"model": "deepseek-v4-flash", "problem_class": "typescript-regex-normalized-form-false-negatives", "result": "passed", "tests": 28}