Unicode fonts and fancy text explained

Sites that offer “fonts” to copy and paste aren't handing out fonts at all. They swap each letter for a different character that happens to look bold, curly or boxed. This guide explains where those characters come from, why Unicode has them, why a styled name can count as twice its length, and what assistive software and search engines make of the result.

A real font changes how a letter is drawn. Fancy text changes which letter it is, and everything that reads text afterwards notices the difference.

Characters, not fonts

Unicode gives every character a number and a name. The capital A you type is U+0041, LATIN CAPITAL LETTER A. Make it bold in a word processor and it is still that character; only the drawing changes, and the bold vanishes when you paste into a box that ignores formatting. The bold-looking 𝐀 produced by a fancy text tool is a different character entirely, U+1D400, MATHEMATICAL BOLD CAPITAL A. It survives a paste because the boldness lives in the character, not in the formatting. That is the whole trick, and the source of every snag that follows.

Where the bold and script letters come from

Most styled alphabets sit in one block, Mathematical Alphanumeric Symbols, U+1D400 to U+1D7FF, added in Unicode 3.1. Mathematicians use letter style to carry meaning: a bold v and an italic v can name different things in one formula, so the styles needed to survive in plain text. Unicode's own names list says the block is for mathematical variables where style matters, and that general text should use standard letters with markup (NamesList.txt, Unicode Character Database 18.0.0 (Unicode Consortium)).

The alphabets have holes. A few styled letters, such as the double-struck ℂℍℚ, were already encoded in Unicode 1.1, in the Letterlike Symbols block, so the newer block leaves reserved slots where they would otherwise go. A careful tool fills each slot with the older character; a careless one leaves a gap or an empty box. The script small l is a later oddity, added in Unicode 4.0.

BlockRangeUsed for
Mathematical Alphanumeric SymbolsU+1D400 to U+1D7FFbold, italic, script, gothic, double-struck, sans-serif, monospace
Letterlike SymbolsU+2100 to U+214Fthe letters missing from those alphabets
Enclosed AlphanumericsU+2460 to U+24FFcircled letters and digits, parenthesised small letters
Enclosed Alphanumeric SupplementU+1F100 to U+1F1FFsquared, filled and parenthesised capitals
Halfwidth and Fullwidth FormsU+FF00 to U+FFEFwide letters
Phonetic ExtensionsU+1D00 to U+1D7Fmost small capitals and some superscripts
Combining Diacritical MarksU+0300 to U+036Fstrikethrough, underline and slash
RunicU+16A0 to U+16FFthe rune-like style

Ranges come from Blocks.txt in the Unicode Character Database, version 18.0.0 (Blocks.txt, Unicode Character Database 18.0.0 (Unicode Consortium)).

Unicode remembers the plain letter

Each styled letter in the database carries a note tying it back to the plain one: U+1D400 is marked as a “font” variant of U+0041 (UnicodeData.txt, Unicode Character Database 18.0.0 (Unicode Consortium)). Software that applies compatibility normalisation, known as NFKC, folds the styled text back into ordinary letters, so a search or a filter built that way sees the plain name. Software that skips normalisation treats the two as unrelated, so a styled name won't match its plain spelling. You can't tell from outside which approach a given app takes.

Why styled names count double

Unicode is divided into planes of 65,536 code points. The first, the Basic Multilingual Plane, holds nearly every everyday letter. The mathematical letters live in the second. UTF-16, a common way of storing text inside programs, fits a first-plane character into one unit but needs two for anything beyond. So “Tui” is three units plain, while 𝐓𝐮𝐢 is 3 characters and 6 units. A length check written in JavaScript, and in many other languages, counts those units, which is how a name that looks short gets refused.

Marks, circles and lookalikes

Strikethrough and underline styles use combining marks, characters that draw over or under the one before. T̶u̶i̶ looks like three letters but is 6 characters, and marks are among the first things a name box strips or refuses. Circled, squared and wide letters are their own characters in the enclosed and full-width blocks. Upside-down, mirrored and rune-like styles borrow from other alphabets by shape alone, so a reader or translator meets Cyrillic, Greek or runic letters standing in for Latin ones. A name that mixes alphabets like this can look like an attempt to imitate someone else's, and may be refused for that reason.

What screen readers and search make of it

A screen reader may spell styled words letter by letter, read each character's full name, or skip them. Search may or may not normalise. Neither outcome can be fixed by the person writing the text, so styled characters belong in decoration, not in names people search for, links or anything that must be understood.

What this guide could not check

Character names, ranges and versions here were read from the Unicode Character Database files for version 18.0.0; the code chart for the block (Code chart: Mathematical Alphanumeric Symbols, U+1D400 to U+1D7FF (Unicode Consortium)) and the full text of The Unicode Standard, Version 18.0 (Unicode Consortium) are the places to read further. How each character looks depends on the fonts on a device, and that couldn't be checked on every phone, console and computer. To try the styles yourself, open the fancy text and name styler.

Last reviewed: