if unicodedata.normalize("NFC", pw1) == unicodedata.normalize("NFC", pw2):
Unicode normalization solves the problem of multiple representations for the same character. The classic example: "café" can be encoded as:
| Form | Codepoints | len() |
|---|---|---|
| NFC | c a f \u00e9 |
4 |
| NFD | c a f e \u0301 |
5 |
The \u00e9 (LATIN SMALL LETTER E WITH ACUTE) and e + \u0301 (COMBINING ACUTE ACCENT) are canonically equivalent — they render identically and represent the same abstract character.
| Form | Algorithm | Preserves… |
|---|---|---|
| NFC | Decompose canonically, then recompose | Canonical equivalence |
| NFD | Decompose canonically, canonical ordering | Canonical equivalence |
| NFKC | Decompose compatibly, canonical ordering, recompose | Compatibility equivalence |
| NFKD | Decompose compatibly, canonical ordering | Compatibility equivalence |
\u0301 (COMBINING ACUTE ACCENT) turns e into é.NFC(x) == NFC(y) proves they're canonically equivalent.fi → fi), superscripts (² → 2), circled digits (① → 1), and Roman numerals (ⅰ → i) are collapsed.caf\u00e9 should equal cafe\u0301They are canonically equivalent — both represent the exact same string "café". Without normalization, a password database (NFC) and a macOS HFS+ filename (NFD) would compare unequal despite meaning the same thing. Normalizing both sides to NFC or NFD before comparison guarantees correctness.
import unicodedata
# Always use the standard library in production
if unicodedata.normalize("NFC", pw1) == unicodedata.normalize("NFC", pw2):
print("Match!")
The full implementation with pure-Python teaching code is at ~/unicode_normalization.py.
The script passes **all tests** against `unicodedata.normalize()` for all four forms, verified across a comprehensive test suite including: **Canonical (NFC/NFD):** - Single combining marks: `a\u0301` / `\u00e1` - Double combining marks: `\u1e69` (s + dot below + dot above) - Vietnamese: `\u1ec3` (e + circumflex + hook above) - Hangul: `\ud55c` (한 → jamo decomposition) - Greek: `\u0391\u0301` (Α + acute → Ά) **Compatibility (NFKC/NFKD):** - Circled digits: `①` → `1` - Superscripts: `²` → `2` - Ligatures: `fi` → `fi` - Unit symbols: `℃` → `°C` - Roman numerals: `ⅰ` → `i` **Edge cases:** - Empty string `""` - Plain ASCII with no combining chars - Strings where NFC == NFD already - Double diacritics on already precomposed characters All raw string equality checks (before normalization) correctly returned `False`; all normalization-equality checks returned `True`. ---
{"model": "claude-sonnet-4-20250514", "problem_class": "str-unicode-normalization-nfc", "result": "passed", "tests": 13}Unicode normalization solves the problem of multiple representations for the same character. The classic example: "café" can be encoded as:
| Form | Codepoints | len() |
|---|---|---|
| NFC | c a f \u00e9 |
4 |
| NFD | c a f e \u0301 |
5 |
The \u00e9 (LATIN SMALL LETTER E WITH ACUTE) and e + \u0301 (COMBINING ACUTE ACCENT) are canonically equivalent — they render identically and represent the same abstract character.
| Form | Algorithm | Preserves… |
|---|---|---|
| NFC | Decompose canonically, then recompose | Canonical equivalence |
| NFD | Decompose canonically, canonical ordering | Canonical equivalence |
| NFKC | Decompose compatibly, canonical ordering, recompose | Compatibility equivalence |
| NFKD | Decompose compatibly, canonical ordering | Compatibility equivalence |
\u0301 (COMBINING ACUTE ACCENT) turns e into é.NFC(x) == NFC(y) proves they're canonically equivalent.fi → fi), superscripts (² → 2), circled digits (① → 1), and Roman numerals (ⅰ → i) are collapsed.caf\u00e9 should equal cafe\u0301They are canonically equivalent — both represent the exact same string "café". Without normalization, a password database (NFC) and a macOS HFS+ filename (NFD) would compare unequal despite meaning the same thing. Normalizing both sides to NFC or NFD before comparison guarantees correctness.
import unicodedata
# Always use the standard library in production
if unicodedata.normalize("NFC", pw1) == unicodedata.normalize("NFC", pw2):
print("Match!")
The full implementation with pure-Python teaching code is at ~/unicode_normalization.py.
The script passes **all tests** against `unicodedata.normalize()` for all four forms, verified across a comprehensive test suite including: **Canonical (NFC/NFD):** - Single combining marks: `a\u0301` / `\u00e1` - Double combining marks: `\u1e69` (s + dot below + dot above) - Vietnamese: `\u1ec3` (e + circumflex + hook above) - Hangul: `\ud55c` (한 → jamo decomposition) - Greek: `\u0391\u0301` (Α + acute → Ά) **Compatibility (NFKC/NFKD):** - Circled digits: `①` → `1` - Superscripts: `²` → `2` - Ligatures: `fi` → `fi` - Unit symbols: `℃` → `°C` - Roman numerals: `ⅰ` → `i` **Edge cases:** - Empty string `""` - Plain ASCII with no combining chars - Strings where NFC == NFD already - Double diacritics on already precomposed characters All raw string equality checks (before normalization) correctly returned `False`; all normalization-equality checks returned `True`. ---
{"model": "claude-sonnet-4-20250514", "problem_class": "str-unicode-normalization-nfc", "result": "passed", "tests": 13}