Precomposed vs Combining Accents in Unicode, Explained
Unicode can represent é in two ways: as a single precomposed character, U+00E9, or as the letter e followed by the combining acute accent U+0301. They look identical but are different sequences of code points.
NFC and NFD
Normalization Form C (NFC) composes sequences into single characters where possible. Normalization Form D (NFD) decomposes them. macOS file systems have historically stored names in a decomposed form, while most web text is NFC, which is why a file named café can fail to match a search for café.
Why it matters
- String comparison: "café" in NFC and NFD are not equal unless normalized.
- Length: the NFD form counts as five characters, the NFC form as four.
- Search and databases: index and query text in the same form.
How to normalize
JavaScript: s.normalize("NFC"). Python: unicodedata.normalize("NFC", s). PHP: Normalizer::normalize($s). Use NFC for storage and display, and NFD (then strip combining marks) when you need to remove accents for matching.
Letters with no precomposed form
Some combinations exist only as base plus combining mark, such as q with a dot above. Use the builder on any letter page to make them.