Grapheme Cluster Splitter
Click a character to copy it. Use Details for its technical record.
Split text at Unicode 17 extended grapheme boundaries and inspect every cluster without changing its contents.
A grapheme is a default segmentation unit, not a guarantee of one visible glyph, one sound or one orthographically valid letter.
Processing is local. No input is stored or sent to a server. Inputs above 12,000 UTF-16 code units are rejected with an explanation rather than silently truncated.
Choose Process text to load the pinned Unicode tables.
Reference examples without processing
Details for U+00E9
Exact sequence
Exact sequence: é
U+00E9
Details for U+306F U+3099
Exact sequence
Exact sequence: ば
U+306F U+3099
Details for U+0628 U+0650
Exact sequence
Exact sequence: بِ
U+0628 U+0650
All tools · Character explorer · Grapheme boundaries · Normalization limits
Sources, versions and limits
- Encoded properties
- Properties use Unicode 17.0.0. Normalization follows UAX #15; grapheme boundaries follow UAX #29. Script associations describe encoded properties, not language use.
- Language evidence
- CLDR 48.2and the IANA registry dated 2026-08-08. Main and auxiliary exemplars are distinct. CLDR exemplars are not complete orthographies.
- Review and limitations
- The dataset is derived from the sources above. Editorial guides cite additional scoped authorities. No native or specialist orthography approval is implied. Release scope date: 14 September 2026. Historic or specialist usage needs separate evidence. Registry-only records have unknown orthography in this release.
- Reproduction and corrections
- Downloads and licences · Coverage and held routes · Suggest a source-backed correction