Diacritics by Language
Worldwide coverage begins with an honest record of evidence. Search all 8,276 language subtags in the language coverage ledger, including languages whose orthography is unknown in this release.
How to read a coverage state
- Verified
- A scoped authority supports a specific usage claim. The published guides above cite their own sources. The frozen dataset ledger retains its original evidence states and does not imply native or specialist approval.
- Dataset-derived
- Direct CLDR exemplar evidence exists. Its locale, exemplar type and draft status are preserved.
- Registry-only/unknown
- An IANA record exists, but direct exemplar evidence has not established the orthography.
- Historic/specialist
- A usage classification requiring its own source and scope. It cannot be inferred merely from an encoded character.
Language guide review queue
The remaining language priorities below are held until their source and review gates pass. These are English editorial subjects. They are not translated versions of this site.
- Catalan Diacritics and Middle Dot
- Galician Accents
- Bosnian Diacritics
- Serbian Cyrillic and Latin Marks
- Estonian Diacritics
- Faroese Accents and Special Letters
- Irish Síneadh Fada Accents
- Scottish Gaelic Accents
- Welsh Diacritics
- Maltese Diacritics and Special Letters
- Albanian Special Letters
- Belarusian Diacritics and Ў
- Bulgarian Stress Marks
- Persian Diacritics
- Urdu Diacritics
- Pashto Diacritics
- Kurdish Diacritics Across Scripts
- Yiddish Pointing Marks
- Hindi Matras and Diacritics
- Sanskrit Marks in Devanagari
- Marathi Marks and Matras
- Nepali Marks and Matras
- Bengali Marks and Vowel Signs
- Assamese Marks and Vowel Signs
- Punjabi Gurmukhi Marks
- Gujarati Marks and Vowel Signs
- Odia Marks and Vowel Signs
- Tamil Vowel Signs and Marks
- Telugu Vowel Signs and Marks
- Kannada Vowel Signs and Marks
- Malayalam Vowel Signs and Marks
- Sinhala Vowel Signs and Marks
- Thai Tone Marks and Vowel Signs
- Lao Tone Marks and Vowel Signs
- Khmer Vowel Signs and Marks
- Burmese Marks and Vowel Signs
- Yoruba Tone Marks and Letters
- Igbo Diacritics
- Hausa Tone and Special-letter Marks
- Amharic and Ethiopic Combining Marks
- Navajo Diacritics
- Māori Macrons
- Hawaiian ʻOkina and Kahakō
- Quechua Diacritics
- Guaraní Diacritics
- Northern Sámi Letters and Diacritics
- Samoan Glottal and Long-vowel Marks
- Tongan Fakauʻa and Long-vowel Marks
Browse encoded script coverage · Find a character · Compose exact text
Sources, versions and limits
- Encoded properties
- Properties use Unicode 17.0.0. Normalization follows UAX #15; grapheme boundaries follow UAX #29. Script associations describe encoded properties, not language use.
- Language evidence
- CLDR 48.2and the IANA registry dated 2026-08-08. Main and auxiliary exemplars are distinct. CLDR exemplars are not complete orthographies.
- Review and limitations
- The dataset is derived from the sources above. Editorial guides cite additional scoped authorities. No native or specialist orthography approval is implied. Release scope date: 14 September 2026. Historic or specialist usage needs separate evidence. Registry-only records have unknown orthography in this release.
- Reproduction and corrections
- Downloads and licences · Coverage and held routes · Suggest a source-backed correction