The Context Window
It is late afternoon, and Tessa has been going back and forth with the chat box for an hour — refining a brochure, trying angles, pasting drafts. Then something odd happens: the model's latest suggestion cheerfully contradicts an instruction she gave near the start. "Keep every mention of pricing out of the text," she had written, and here is pricing, back in the text. Nothing broke. She has just met the model's most important limit, and it deserves the plainest possible explanation.
The model has exactly one working memory, called the context window: the amount of text it can consider at once, measured — as everything is — in tokens. Everything that should influence the next answer has to fit inside it: your messages, the model's own replies, anything you pasted. It is big — modern windows hold a book's worth of text — but it is never infinite. Picture the whiteboard in a meeting room: everything the meeting has decided must be written on the board, and when the board fills up, someone erases the oldest corner to make room. The model answers from the board. Never from the meeting's memory — there is no meeting's memory.
What Happens When It Fills
An hour of back-and-forth adds up. Every message and reply lands in the window, and once the conversation outgrows it, the oldest turns stop fitting — Tessa's pricing instruction among them. The model then answers using only what remains in view. From the outside this reads as forgetting; mechanically, it is simpler and stranger: the instruction was not forgotten, it was no longer shown. The model followed everything it could see, perfectly. What it could see no longer included her rule.
This explains a family of frustrations that every regular user has met. Long chats drift off their brief. Constraints quietly lapse. A carefully established persona goes flat. The model contradicts "what you told it" — except you told it something that has since slid off the whiteboard. Nothing is malfunctioning, and — this is the useful part — nothing is random about it. Overflow has symptoms you can now recognize on sight.
Every Chat Starts Empty
There is a second half to the picture. Open a new conversation and the window starts blank: the model carries nothing over from yesterday's chat, or from the chat sitting in the next browser tab. Whatever it seemed to know about Tessa yesterday, it does not know now. (Why the chat box nonetheless shows her old conversations, and how some products give a real impression of remembering you, is a wrapper's trick — Chapter 7 reveals it, and it is a good one.)
For now, take the practical fact: a conversation is a self-contained world. Everything the model should know in this world, you put in this world.
Living with the Window
Three habits turn this limit from a trap into a tool, and they cost nothing. Start fresh for a new task. A long, wandering chat drags its whole history through the window on every turn; a clean chat starts sharp. Restate what matters. If an instruction must hold, repeat it in your latest message rather than trusting the one from an hour ago — professionals do this constantly, without embarrassment. Put the important thing near the end. Recent text is always safely in view; ancient text is what falls off.
And when Tessa meets a document too big for the window — the 120-page supplier contract is waiting in Chapter 4 — the answer will not be to push harder. It will be craft: choosing what goes onto the whiteboard, because choosing is the skill the window forces, and it turns out to be the skill that separates good model use from bad far more than clever phrasing ever does.
- "The model remembers our whole conversation." It sees what currently fits in the window. A long chat quietly loses its own beginning — which is why the pricing rule from an hour ago stopped working.
- "It forgot my instruction — it's unreliable." It followed everything it could see. The instruction had slid out of view; restating it in a recent message fixes the problem instantly, every time.
- "If I said it once in this chat, it's in there for good." Nothing in the window is permanent. Repeating what matters is normal practice, not a workaround for a defect.
- "The model remembers me from yesterday's chat." Every conversation starts with an empty window. When a product seems to remember you across chats, the product is re-supplying your details — a mechanism Chapter 7 shows in full.
- Half of everyday model frustration — drift, dropped constraints, contradictions in long chats — is window overflow, and you can now recognize it and fix it with a restatement instead of blaming the tool.
- The window is the reason big documents must be cut to fit (Chapter 4) and the reason long conversations cost more than short ones (Chapter 7) — this one limit stands behind both.
Knowledge Check
Why did the model bring pricing back after Tessa told it not to, an hour earlier?
- It decided that pricing information would make the brochure more persuasive and put it back
- Her instruction had slid out of the context window and was no longer shown
- A random glitch made it drop one instruction while keeping the others
- Instructions automatically expire after about an hour of chatting
What does a brand-new conversation start with?
- A short summary of your previous chats, so the model starts with some background
- Whatever the model learned about you and your work in earlier conversations
- The contents of your other open chat tabs, merged into one shared memory
- An empty window, since the model carries nothing over from any other chat
An important instruction must hold through a long working session. What is the professional habit?
- Restate the instruction in a recent message instead of trusting the early one
- Write it in stronger words at the very start so the model takes it seriously
- Ask the model to promise it will remember the instruction throughout
- Repeat the instruction three times in the same message for emphasis
What is the context window, in one sentence?
- The full record of every message you have ever sent, kept in your chat history
- The time the model is allowed to spend considering your question
- The finite amount of text, in tokens, the model can consider at once
- The maximum length of the answer the model is allowed to write in one reply
You got correct