Unicode Diacritics and Mark Corpus

This corpus includes exactly Diacritic=Yes OR General_Category Mn/Mc/Me from Unicode 17.0.0. There are 1,247 Diacritic characters and 2,543 Mark characters. Their intersection is 811, giving a union of 2,979.

Diacritic only
436 characters
Mark only
1,732 characters
Nonspacing marks (Mn)
2,059
Spacing combining marks (Mc)
471
Enclosing marks (Me)
13

What the set does and does not mean

Every member is an assigned scalar under the stated rule. The data retains names, category, combining class, decomposition, Script and Script_Extensions. It does not classify every member as an orthographic diacritic: variation selectors and other technical marks are included.

Precomposed letters whose decompositions contain marks are not automatically members of this union. Existing accent-letter references remain a separate exact-copy surface. Language associations require scoped evidence.

Versioned downloads

Exact sequence and grapheme records · Workbook contract export (compressed JSON)· Download checksums · Latest release manifest

Search the candidate corpus · Read the coverage ledger · Script counts

Sources, versions and limits

Encoded properties
Properties use Unicode 17.0.0. Normalization follows UAX #15; grapheme boundaries follow UAX #29. Script associations describe encoded properties, not language use.
Language evidence
CLDR 48.2and the IANA registry dated 2026-08-08. Main and auxiliary exemplars are distinct. CLDR exemplars are not complete orthographies.
Review and limitations
The dataset is derived from the sources above. Editorial guides cite additional scoped authorities. No native or specialist orthography approval is implied. Release scope date: 14 September 2026. Historic or specialist usage needs separate evidence. Registry-only records have unknown orthography in this release.
Reproduction and corrections
Downloads and licences · Coverage and held routes · Suggest a source-backed correction