Diacritical Marks Data Downloads

Download reproducible, versioned data for exact character lookup and language coverage. JSON is the exact scalar interchange format; CSV adds a protective apostrophe to cells that could be interpreted as formulas by spreadsheet software.

Reproduction and evidence

The generator reconciles the Unicode union against the implementation inventory, verifies source SHA-256 hashes, parses the pinned IANA registry, and joins direct CLDR main-exemplar evidence. Auxiliary and other exemplar types remain separate in the CLDR evidence download. Root is excluded from language-locale counts.

Registry evidence does not establish complete orthographies. The source register identifies the exact retrieval locations, versions and checksums; the included README explains the schema and CSV safeguards.

Versioned downloads

Exact sequence and grapheme records · Workbook contract export (compressed JSON)· Download checksums · Latest release manifest

Search the candidate corpus · Read the coverage ledger · Script counts

Sources, versions and limits

Encoded properties
Properties use Unicode 17.0.0. Normalization follows UAX #15; grapheme boundaries follow UAX #29. Script associations describe encoded properties, not language use.
Language evidence
CLDR 48.2and the IANA registry dated 2026-08-08. Main and auxiliary exemplars are distinct. CLDR exemplars are not complete orthographies.
Review and limitations
The dataset is derived from the sources above. Editorial guides cite additional scoped authorities. No native or specialist orthography approval is implied. Release scope date: 14 September 2026. Historic or specialist usage needs separate evidence. Registry-only records have unknown orthography in this release.
Reproduction and corrections
Downloads and licences · Coverage and held routes · Suggest a source-backed correction