◐ Off-By-One · answer catalog

str-unicode-normalization-nfc

2 answer(s)pythonpython3pythonpython3

if unicodedata.normalize("NFC", pw1) == unicodedata.normalize("NFC", pw2):

📦 Source in repository (JSON)

Answer 1

Unicode normalization solves the problem of multiple representations for the same character. The classic example: "café" can be encoded as:

Form Codepoints len()
NFC c a f \u00e9 4
NFD c a f e \u0301 5

The \u00e9 (LATIN SMALL LETTER E WITH ACUTE) and e + \u0301 (COMBINING ACUTE ACCENT) are canonically equivalent — they render identically and represent the same abstract character.

The Four Normalization Forms

Form Algorithm Preserves…
NFC Decompose canonically, then recompose Canonical equivalence
NFD Decompose canonically, canonical ordering Canonical equivalence
NFKC Decompose compatibly, canonical ordering, recompose Compatibility equivalence
NFKD Decompose compatibly, canonical ordering Compatibility equivalence

Key Concepts

Why caf\u00e9 should equal cafe\u0301

They are canonically equivalent — both represent the exact same string "café". Without normalization, a password database (NFC) and a macOS HFS+ filename (NFD) would compare unequal despite meaning the same thing. Normalizing both sides to NFC or NFD before comparison guarantees correctness.

import unicodedata

# Always use the standard library in production
if unicodedata.normalize("NFC", pw1) == unicodedata.normalize("NFC", pw2):
    print("Match!")

The full implementation with pure-Python teaching code is at ~/unicode_normalization.py.


Evidence & signatures

The script passes **all tests** against `unicodedata.normalize()` for all four forms, verified across a comprehensive test suite including:

**Canonical (NFC/NFD):**
- Single combining marks: `a\u0301` / `\u00e1`
- Double combining marks: `\u1e69` (s + dot below + dot above)
- Vietnamese: `\u1ec3` (e + circumflex + hook above)
- Hangul: `\ud55c` (한 → jamo decomposition)
- Greek: `\u0391\u0301` (Α + acute → Ά)

**Compatibility (NFKC/NFKD):**
- Circled digits: `①` → `1`
- Superscripts: `²` → `2`
- Ligatures: `fi` → `fi`
- Unit symbols: `℃` → `°C`
- Roman numerals: `ⅰ` → `i`

**Edge cases:**
- Empty string `""`
- Plain ASCII with no combining chars
- Strings where NFC == NFD already
- Double diacritics on already precomposed characters

All raw string equality checks (before normalization) correctly returned `False`; all normalization-equality checks returned `True`.

---
{"model": "claude-sonnet-4-20250514", "problem_class": "str-unicode-normalization-nfc", "result": "passed", "tests": 13}

Answer 2

Unicode normalization solves the problem of multiple representations for the same character. The classic example: "café" can be encoded as:

Form Codepoints len()
NFC c a f \u00e9 4
NFD c a f e \u0301 5

The \u00e9 (LATIN SMALL LETTER E WITH ACUTE) and e + \u0301 (COMBINING ACUTE ACCENT) are canonically equivalent — they render identically and represent the same abstract character.

The Four Normalization Forms

Form Algorithm Preserves…
NFC Decompose canonically, then recompose Canonical equivalence
NFD Decompose canonically, canonical ordering Canonical equivalence
NFKC Decompose compatibly, canonical ordering, recompose Compatibility equivalence
NFKD Decompose compatibly, canonical ordering Compatibility equivalence

Key Concepts

Why caf\u00e9 should equal cafe\u0301

They are canonically equivalent — both represent the exact same string "café". Without normalization, a password database (NFC) and a macOS HFS+ filename (NFD) would compare unequal despite meaning the same thing. Normalizing both sides to NFC or NFD before comparison guarantees correctness.

import unicodedata

# Always use the standard library in production
if unicodedata.normalize("NFC", pw1) == unicodedata.normalize("NFC", pw2):
    print("Match!")

The full implementation with pure-Python teaching code is at ~/unicode_normalization.py.


Evidence & signatures

The script passes **all tests** against `unicodedata.normalize()` for all four forms, verified across a comprehensive test suite including:

**Canonical (NFC/NFD):**
- Single combining marks: `a\u0301` / `\u00e1`
- Double combining marks: `\u1e69` (s + dot below + dot above)
- Vietnamese: `\u1ec3` (e + circumflex + hook above)
- Hangul: `\ud55c` (한 → jamo decomposition)
- Greek: `\u0391\u0301` (Α + acute → Ά)

**Compatibility (NFKC/NFKD):**
- Circled digits: `①` → `1`
- Superscripts: `²` → `2`
- Ligatures: `fi` → `fi`
- Unit symbols: `℃` → `°C`
- Roman numerals: `ⅰ` → `i`

**Edge cases:**
- Empty string `""`
- Plain ASCII with no combining chars
- Strings where NFC == NFD already
- Double diacritics on already precomposed characters

All raw string equality checks (before normalization) correctly returned `False`; all normalization-equality checks returned `True`.

---
{"model": "claude-sonnet-4-20250514", "problem_class": "str-unicode-normalization-nfc", "result": "passed", "tests": 13}
Generated from the verified corpus · MIT licensedBack to the catalog