Grapheme Cluster Splitter

Click a character to copy it. Use Details for its technical record.

Split text at Unicode 17 extended grapheme boundaries and inspect every cluster without changing its contents.

A grapheme is a default segmentation unit, not a guarantee of one visible glyph, one sound or one orthographically valid letter.

Processing is local. No input is stored or sent to a server. Inputs above 12,000 UTF-16 code units are rejected with an explanation rather than silently truncated.

Choose Process text to load the pinned Unicode tables.

Reference examples without processing

Details for U+00E9

Exact sequence

Exact sequence: é

U+00E9

Details for U+306F U+3099

Exact sequence

Exact sequence: ば

U+306F U+3099

Details for U+0628 U+0650

Exact sequence

Exact sequence: بِ

U+0628 U+0650

All tools · Character explorer · Grapheme boundaries · Normalization limits

Sources, versions and limits

Encoded properties
Properties use Unicode 17.0.0. Normalization follows UAX #15; grapheme boundaries follow UAX #29. Script associations describe encoded properties, not language use.
Language evidence
CLDR 48.2and the IANA registry dated 2026-08-08. Main and auxiliary exemplars are distinct. CLDR exemplars are not complete orthographies.
Review and limitations
The dataset is derived from the sources above. Editorial guides cite additional scoped authorities. No native or specialist orthography approval is implied. Release scope date: 14 September 2026. Historic or specialist usage needs separate evidence. Registry-only records have unknown orthography in this release.
Reproduction and corrections
Downloads and licences · Coverage and held routes · Suggest a source-backed correction