Diacritic and Character Inspector
Click a character to copy it. Use Details for its technical record.
Inspect exact scalars, Unicode names, UTF encodings, escapes, grapheme boundaries and the four normalization forms.
Candidate records include Script and Script_Extensions. Values outside the candidate corpus are labelled separately. Unassigned or algorithmically named ranges are not given invented character names.
Processing is local. No input is stored or sent to a server. Inputs above 12,000 UTF-16 code units are rejected with an explanation rather than silently truncated.
Choose Process text to load the pinned Unicode tables.
Reference examples without processing
Details for U+00E9
Exact sequence
Exact sequence: é
U+00E9
Details for U+306F U+3099
Exact sequence
Exact sequence: ば
U+306F U+3099
Details for U+0628 U+0650
Exact sequence
Exact sequence: بِ
U+0628 U+0650
All tools · Character explorer · Grapheme boundaries · Normalization limits
Sources, versions and limits
- Encoded properties
- Properties use Unicode 17.0.0. Normalization follows UAX #15; grapheme boundaries follow UAX #29. Script associations describe encoded properties, not language use.
- Language evidence
- CLDR 48.2and the IANA registry dated 2026-08-08. Main and auxiliary exemplars are distinct. CLDR exemplars are not complete orthographies.
- Review and limitations
- The dataset is derived from the sources above. Editorial guides cite additional scoped authorities. No native or specialist orthography approval is implied. Release scope date: 14 September 2026. Historic or specialist usage needs separate evidence. Registry-only records have unknown orthography in this release.
- Reproduction and corrections
- Downloads and licences · Coverage and held routes · Suggest a source-backed correction