Grapheme Clusters and Multi-code-point Letters
Quick copy
Details for U+0065 U+0301
LATIN SMALL LETTER E + COMBINING ACUTE ACCENT
Exact sequence: é
U+0065 U+0301
Details for U+0915 U+094D U+0937
DEVANAGARI LETTER KA + DEVANAGARI SIGN VIRAMA + DEVANAGARI LETTER SSA
Exact sequence: क्ष
U+0915 U+094D U+0937
Click a character to copy it. Use Details for its technical record.
An extended grapheme cluster is a default text-segmentation unit. It can contain multiple Unicode scalars and multiple UTF-16 code units.
Default extended grapheme boundaries follow UAX #29. Combining sequences, emoji ZWJ sequences and Indic conjunct behavior require more than splitting a JavaScript string into individual code units.
The splitter uses Unicode 17 grapheme, Extended_Pictographic and Indic_Conjunct_Break properties. It preserves every scalar, including joiners. It does not certify that every cluster is an orthographically valid letter.
When copying, preserve the complete selected cluster. When inspecting, show constituent code points without inserting presentation characters into the copy value. A leading combining mark receives a display-only dotted circle.
Exact examples
Details for U+1F469 U+1F3FD U+200D U+1F4BB
Exact sequence
Exact sequence: 👩🏽💻
U+1F469 U+1F3FD U+200D U+1F4BB
Details for U+0301
Exact sequence
Exact sequence: ́
U+0301
| Original code points | NFC | NFD | NFKC | NFKD |
|---|---|---|---|---|
U+0065 U+0301 | éU+00E9 | éU+0065 U+0301 | éU+00E9 | éU+0065 U+0301 |
U+0915 U+094D U+0937 | क्षU+0915 U+094D U+0937 | क्षU+0915 U+094D U+0937 | क्षU+0915 U+094D U+0937 | क्षU+0915 U+094D U+0937 |
U+1F469 U+1F3FD U+200D U+1F4BB | 👩🏽💻U+1F469 U+1F3FD U+200D U+1F4BB | 👩🏽💻U+1F469 U+1F3FD U+200D U+1F4BB | 👩🏽💻U+1F469 U+1F3FD U+200D U+1F4BB | 👩🏽💻U+1F469 U+1F3FD U+200D U+1F4BB |
U+0301 | ́U+0301 | ́U+0301 | ́U+0301 | ́U+0301 |
Try the normalizer · Split clusters · Inspect encoded candidates · Precomposed and combining text
Sources, versions and limits
- Encoded properties
- Properties use Unicode 17.0.0. Normalization follows UAX #15; grapheme boundaries follow UAX #29. Script associations describe encoded properties, not language use.
- Language evidence
- CLDR 48.2and the IANA registry dated 2026-08-08. Main and auxiliary exemplars are distinct. CLDR exemplars are not complete orthographies.
- Review and limitations
- The dataset is derived from the sources above. Editorial guides cite additional scoped authorities. No native or specialist orthography approval is implied. Release scope date: 14 September 2026. Historic or specialist usage needs separate evidence. Registry-only records have unknown orthography in this release.
- Reproduction and corrections
- Downloads and licences · Coverage and held routes · Suggest a source-backed correction