Text, Unicode and UTF-8
"How long is this string?" has four correct answers: the number of bytes, the number of code units, the number of code points, and the number of characters a reader sees. Lantern's catalogue holds titles in dozens of scripts, and Noor has met each answer the hard way: a title that fails to match its own search, a length check that cuts an emoji in half, a column that takes three times the bytes it was sized for.
All three bugs come from treating one layer as another. This topic separates the four layers of a string, shows how UTF-8 turns numbers into bytes, and prices each choice a language or a database makes about which layer to count.
Four Layers of One String
A grapheme is what a reader sees as one character. A code point is Unicode's number for one abstract character. A code unit is the fixed-size chunk an encoding works in: 8 bits in UTF-8, 16 bits in UTF-16. And bytes are what is actually stored. For plain English text all four counts agree, and so much code gets away with confusing them.
They come apart quickly. The letter é can be one code point, or two: a plain e followed by a combining accent. The family emoji of a man, a woman and a girl is one grapheme built from five code points: three people joined by two invisible zero-width joiners. In UTF-16 those five code points take 8 code units, and in UTF-8 they take 18 bytes. Every language's length function counts one of these layers, and they do not all count the same one.
Unicode, the Numbering
Unicode assigns a number to every character in every script, from U+0000 to U+10FFFF, a space of 1,114,112 code points. Unicode 17.0, the version Python 3.15 ships, assigns 159,801 characters, most of them Chinese, Japanese and Korean ideographs. The numbers are the agreement: U+00F1 is ñ on every machine in the world.
A number is not yet a byte. An encoding is the rule for writing code points as bytes, and there are several, each with a different price. Mixing up the numbering and the encoding is the root of most text bugs, because two programs can agree completely on Unicode and still disagree on the bytes.
How UTF-8 Writes a Code Point
UTF-8 uses 1 byte for code points up to U+007F, identical to ASCII, 2 bytes up to U+07FF, 3 bytes up to U+FFFF, and 4 bytes for the rest. The first byte's leading bits say how many bytes the character takes: a leading 0 means one, 110 means two, 1110 means three, 11110 means four. Every continuation byte starts with the bits 10. The code point's own bits are poured into the free slots.
So é, U+00E9, has 11 significant bits and becomes the two bytes C3 and A9 in hexadecimal. The kanji for forest, U+68EE, needs 16 bits and becomes three bytes, E6, A3 and AE. Because continuation bytes are marked, a reader dropped anywhere in a UTF-8 stream can step back to the start of the current character, and a search for an ASCII byte such as a slash or a quote can never match the inside of another character: every byte of a multi-byte character has its top bit set, and no ASCII byte does.
Lantern's Titles
Measured in UTF-8 on Python 3.15, "The Dispossessed" is 16 bytes for 16 characters. "Cien años de soledad" is 21 bytes for 20 characters, because ñ takes two. The Russian title "Война и мир" is 20 bytes for 11 characters. The Japanese title ノルウェイの森 is 21 bytes for 7 characters, three bytes each. And the Hindi title गोदान shows the fourth layer: it is 5 code points and 15 bytes, but a reader sees 3 characters, because two of the code points are vowel signs that attach to the consonant before them.
| Title | Graphemes | Code points | UTF-8 bytes |
|---|---|---|---|
| The Dispossessed | 16 | 16 | 16 |
| Cien años de soledad | 20 | 20 | 21 |
| Война и мир | 11 | 11 | 20 |
| ノルウェイの森 | 7 | 7 | 21 |
| गोदान | 3 | 5 | 15 |
A title field sized as "100 characters, so 100 bytes" holds 100 English characters and only 33 Japanese ones. A cut at an arbitrary byte can land inside a character: cutting the Japanese title after 10 bytes leaves three whole characters and one stray byte, which decodes as the replacement character, a question mark in a diamond, or raises an error, depending on how the reader was told to handle bad input. Python's standard library counts code points and has no function that counts graphemes. The grapheme counts here follow the Unicode segmentation rules and were checked with the segmenter built into JavaScript.
Normalization and Matching
The letter ñ can be stored as one code point, U+00F1, or as n followed by a combining tilde, U+0303. The two look identical on every screen and compare unequal, because they are different code points: 2 bytes in one form and 3 in the other. A record catalogued in the decomposed form never matches a query typed on a kiosk in the composed form. Lantern has to normalize both sides to one form, usually the composed form called NFC, before indexing and again before searching.
Case needs the same care. Lowercasing is not enough for case-insensitive matching, because some characters do not map one to one: the German ß uppercases to SS, so "Straße" and "STRASSE" lowercase to different strings. Case folding, which Python offers as a separate string method, maps both to "strasse" and makes them match.
What Each Choice Costs
UTF-8 costs 1 byte per character for English, 2 for Cyrillic, Greek and accented Latin, and 3 for most Chinese and Japanese, where UTF-16 costs 2 for all of them. Finding the thousandth code point in UTF-8 means scanning from the start, because characters differ in width, so indexing by code point is O(n) and languages choose which layer to index. Go and Rust index bytes. JavaScript and Java index UTF-16 units. CPython stores each string at 1, 2 or 4 bytes per code point, depending on its widest character, and keeps indexing O(1).
CPython's choice has a price of its own. On 3.15 a string of 1,000 ASCII letters takes 1,041 bytes. Change one letter to 森 and the whole string moves to 2 bytes per character, 2,058 bytes. Change it to an emoji instead and it moves to 4 bytes per character, 4,060 bytes: one emoji roughly quadruples the memory of the whole string. A service holding millions of short strings pays for its widest character in each one.
UTF-8 uses 1 to 4 bytes per code point, is compatible with ASCII and has no byte-order question. It is the encoding of the web, of files and of network protocols.
UTF-16 uses 2 or 4 bytes, with pairs of units called surrogate pairs above U+FFFF. JavaScript, Java and Windows use it inside the program, and its length counts code units, so an emoji counts as 2.
UTF-32 uses 4 bytes for every code point, which makes indexing trivial and text four times the size of ASCII. It is rare outside program internals. Store and send UTF-8, and know which one your language uses inside.
- "UTF-8 is one byte per character." Only for ASCII. Accented Latin letters take 2 bytes, most Chinese, Japanese and Korean characters 3, and emoji 4, so byte limits sized for English truncate every other script.
- "The length function counts the characters the user sees." Python counts code points, JavaScript counts UTF-16 units, Go counts bytes. The family emoji is 5 in Python, 8 in JavaScript, 18 in Go and 1 on the screen, so a 280-character limit enforced in one language disagrees with the same limit enforced in another.
- "Two strings that look the same are equal." The composed and decomposed forms of é or ñ render identically and compare unequal. Searching, deduplication and unique constraints miss matches until both sides are normalized.
- "Lowercasing makes comparison case-insensitive." Some characters change length or have no one-to-one mapping: ß uppercases to SS, and Turkish has a dotted and a dotless i. Case folding exists because lowercasing is not enough.
- "A column named utf8 stores all of Unicode." MySQL's historical utf8 character set stores at most 3 bytes per character and rejects emoji. The full encoding there is called utf8mb4, and the name has caused real data loss.
- Store and transmit text as UTF-8, and name the encoding at every boundary. Defaults vary by platform and by year, and Python itself changed its default in 3.15.
- Normalize text to one form before indexing, comparing or enforcing uniqueness. Visually identical strings must be identical in code points before equality means anything.
- Size and truncate by bytes where bytes are the limit, and by graphemes where a person will see the result. Cutting by code points can still split a flag or a family emoji.
- Use case folding, not lowercasing, for case-insensitive matching. It handles the characters whose upper and lower case are not one to one.
Knowledge Check
A title column holds at most 60 bytes of UTF-8. Roughly how many characters of a Japanese title fit?
- 60, because UTF-8 stores one character per byte
- 30, because each character takes two bytes in UTF-8
- About 20, because most Japanese characters take three
- 15, because every non-ASCII character takes four bytes
The same family emoji is checked against a length limit in Python, in JavaScript and in Go. What lengths do they report?
- 5 in Python, 8 in JavaScript and 18 in Go
- 1 in all three, since it is one visible character
- 18 in Python, 5 in JavaScript and 8 in Go
- 3 in all three, one for each person in the emoji
A record titled "Cien años de soledad" never appears when a patron types the same title on a kiosk. The text looks identical. What is the most likely cause?
- The kiosk sends the title in the opposite byte order
- The index compares case, and the patron used capitals
- The index cannot store characters outside of ASCII
- The ñ is stored in a different Unicode form on each side
A service keeps millions of short ASCII labels in CPython. One label in each gains an emoji. What happens to memory?
- Each label grows by the 4 bytes of one emoji
- Each label grows by the emoji's 4 UTF-8 bytes
- Every character in a changed label widens to 4 bytes
- Nothing, because CPython stores emoji elsewhere
Why can a byte-level search for the slash character never match inside a multi-byte UTF-8 character?
- Because the search decodes the text into code points first
- Because every byte of such a character has its top bit set
- Because UTF-8 skips the byte value of a slash altogether
- Because multi-byte characters are wrapped in a marker byte
You got correct