Editorial method & sources

Character facts come from official Unicode data. Explanations are written for this reference; the data is generated from a pinned source rather than copied from another character website.

Inclusion and classification

We parse code points, official names, general categories, combining classes, and decompositions from UnicodeData.txt. We recursively resolve canonical decompositions, retaining uppercase and lowercase Latin letters only when the result is an ASCII A–Z or a–z letter followed entirely by marks. Compatibility decompositions are excluded.

Ligatures, stroked letters, and other Latin forms that fail this rule remain in a separate related-forms dataset. The Fahrenheit sign is a symbol, not an accented Latin letter, and is excluded. Names describe Unicode identities; we do not infer language usage from names.

Authoritative sources

Validation and limitations

The generator records source URLs, retrieval dates, and SHA-256 checksums. Automated checks compare regenerated data, reject duplicates, and test character formatting and site routes. CP1252 shortcuts are mapping-verified, not tested across every Windows setup. Browser fonts and keyboard support vary.

The text tool uses the browser’s normalization implementation. Removing marks is not full transliteration. See the tool’s limitations and the typing guide.

Corrections

Compare a questionable character with the pinned source and include its code point when documenting the discrepancy. Our contact page explains current feedback availability.