Tokens: How the Model Reads
Tessa pastes a paragraph of tour description into the chat box and idly wonders whether the model reads it the way she does — word by word, left to right, maybe skimming. Close, but not quite, and the difference turns out to run through everything: what the model costs, how much it can hold, even why it occasionally stumbles over an unusual name. Before any text reaches the model, it is cut into tokens — small chunks, usually a word or a piece of a word — and tokens are the only thing the model ever sees.
Think of LEGO bricks. When Tessa looks at her paragraph she sees the castle — sentences, meaning, tone. The model sees the bricks it was built from. A common short word like the or travel is one brick. A longer or rarer word gets snapped together from several: unbelievable might arrive as un, believ, able. The bricks are what get counted, and two different counters in this book — one for size, one for money — both count in bricks.
Text Becomes Chunks
The cutting happens automatically, before the model does anything at all. Your sentence goes in; a stream of tokens comes out; the model reads that stream. There is no page in there, no layout, no fonts — a bulleted list survives only as the characters that make it up. The model then produces its answer the same way, token by token, like beads coming off a string one at a time. You have already seen the evidence without knowing it: that typewriter effect, where the answer appears in little spurts rather than all at once, is the token stream arriving as it is made.
How big is a token? For English text, a dependable rule of thumb: a token is about three-quarters of a word. So 100 tokens is roughly 75 words, and a full page of text — say 400 words — is around 530 tokens. You will never need to be more precise than that, but you will use this rule surprisingly often, the way you use rough centimetres-to-inches without ever getting out a ruler.
Why the Chunks Are Uneven
The cutting is not arbitrary — common stretches of text become single tokens, rare ones get split into pieces. That is why an everyday word is one token while a rare surname, a chemical name, or a made-up tour code like WYM-2288 gets chopped into fragments. This also explains a small mystery you may have noticed: the model handles common words flawlessly and occasionally fumbles rare ones — miscounting the letters in a strange word, or mangling an unusual name. It never saw the word as a whole. It saw three bricks, and worked with the bricks.
Different providers cut slightly differently, and other languages cut less efficiently than English — the same paragraph in German or Japanese usually costs more tokens. None of that detail needs memorizing. The durable fact is that the cutting exists, it is uneven, and everything downstream is measured in the pieces.
Why You Will Care
Two hard limits govern every conversation you will ever have with a model, and both are measured in tokens. The first is the context window — how much the model can consider at once. It is the very next page, and it is the reason Tessa's 120-page contract will cause trouble in Chapter 4. The second is the bill. When Waymark starts paying for model access in Chapter 7, the price will be set per token — so many tokens in, so many tokens out, each direction metered. A question that costs half a cent and a question that costs fifteen cents differ in exactly one way: how many bricks went through.
So the humble ¾-of-a-word rule is quietly a business skill. It lets Tessa look at a document and estimate, on a napkin, what it weighs in the model's terms: this email is trivial, this report is a few thousand tokens, this contract is well over a hundred thousand and will not be going through in one piece. That estimating habit — casual, approximate, constant — is the first fully practical thing this book teaches, and it never stops being useful.
- "A token is just a word." Often, not reliably — long or rare words split into several tokens. Keep the working rule instead: about three-quarters of a word, so 100 tokens is roughly 75 English words.
- "The model reads the page like I do." It reads a stream of tokens. There is no layout, no glance, no skimming — formatting survives only as characters, and the answer comes out the same way, one token at a time.
- "Token counts are trivia for engineers." They are the unit of both limits that will shape your daily use: the context window (next page) and the bill (Chapter 7). This is one of the most practical numbers in the book.
- "The model sees whole words, so it can always spell and count them." A rare word may arrive as fragments, which is exactly why models sometimes fumble unusual names or miscount letters — the whole word was never there.
- Tokens are the unit of the model's two hard limits — the window and the bill. Every "why did it cut off?" and every "why did that cost so much?" in this book traces back to this page.
- The ¾-word rule turns you into someone who can estimate any document's size in the model's terms on a napkin — a small skill you will use weekly for as long as you work with these tools.
Knowledge Check
Roughly how many English words is 100 tokens?
- About 100 words, because each token corresponds to exactly one English word
- About 75 words, since a token is roughly three-quarters of a word
- About 400 words, because one token usually covers several words
- It cannot be estimated, because token sizes vary far too much
What does the model actually receive when Tessa pastes a formatted, bulleted paragraph?
- A stream of tokens, where the formatting survives only as characters
- An image of the page, so it can see the layout and bullets
- A list of sentences, which it reads and weighs one at a time in order
- A compressed summary of the text, so that a long document still fits in the window
Why does the model sometimes fumble a rare name or miscount the letters in an unusual word?
- Rare words make the model less careful, so it pays less attention to them
- Rare words are missing from the dictionary the model looks words up in
- A rare word arrives cut into fragments, so the model never sees it as a whole
- The model is built to skip unusual words to keep its answers fast
Which two things in your daily use of a model are measured in tokens?
- The speed of the answer and the quality of the writing
- The context window and the bill for using the model
- The safety limits and the topics the model will discuss
- The accuracy of answers and how often they need checking
You got correct