Unicode Data and Testing Methodology

How we source Unicode 17.0 data, build 89 collections, measure font coverage across 3 bundled fonts, review claims and record what we did not test.

At a glance

The short version

01

Use trusted source files

We use Unicode 17 files for names, codes, and emoji lists.

02

Count from stored data

We count each saved mark and check which fonts can draw it.

03

Say what we did not test

We tell you when a claim still needs a real user or expert check.

By the numbers

Figures anyone can check against the pages themselves.

89
symbol collections, each counted from stored data
9,651
distinct characters and sequences on the site
17.0
Unicode version the character data comes from
20
guides, every one citing its sources
41
FAQ answers reviewed against the standard
3
bundled fallback fonts measured by cmap coverage
1,293
code points only the maths font can draw
0
claims of testing we did not carry out

01

Unicode sources and versioning

The site bundles Unicode 17 character and emoji data. That data generates every formal name, general category, emoji flag and qualified sequence.

Category counts come from the stored values after processing, not from a rounded marketing figure. A Unicode name proves what a character is.

It does not prove a translation, a pronunciation, a religious meaning or a right to use a trademark.

02

How collections are assembled

Original page addresses are preserved so old links keep working. Inventories are then expanded from official ranges and reference sets.

Each page groups characters by what people are searching for. It also shows formal names and code points so you can tell lookalikes apart.

Multi-character emoji keep their stored order, including modifiers and joiners. Corrupt data is dropped rather than published as a real symbol.

03

Font coverage measurement

Compatibility notes come from reading each bundled font's character map. We compare it against the code points the collections actually use.

A mapped code point means the font declares a glyph for it. It does not promise good design, correct spacing, colour or support in another app.

Three fallback fonts are still needed, because the current inventory includes 1,293 code points only the math font covers.

04

Destination and platform checks

A rendering claim should record five things: the exact sequence, the operating system, the app or browser, the font if known, and the date. We separate a rule we can reproduce from an observation that a platform update could change.

Important text is checked after saving or exporting, not only in the editor. Normalisation, filtering and font fallback can all change text at that step.

05

Accessibility, language and cultural boundaries

Our guidance follows published Unicode and accessibility documentation. Reading documentation is not the same as running a study with assistive-technology users.

It is also not the same as review by a fluent speaker. Pages say so where it matters.

We do not claim testing that did not happen. Teams should check critical workflows with their own users.

Permanent language, identity, religious, medical or legal uses need a qualified reviewer.

06

Dates, corrections and reproducibility

A page's modified date changes only after a real review of its data or guidance. To report a problem, send the page URL, the exact copied text, the code points if you have them, and where it went wrong.

Screenshots help show appearance. Pasted text is what we need, because invisible characters and lookalikes do not show up in an image.

Copied