Why the Same Question Gets Different Answers
Tessa asks the chat box for three subject lines for the Lakeland tour email, likes them, and — wanting one more option — asks the identical question again. She gets three completely different subject lines. Also good. Which set did the model "really mean"? Her instinct says one of them must be the true answer and the other a variation. The honest answer is stranger: neither. There is no stored answer inside the model at all, and understanding why will explain half of the model's personality.
Here is what actually happens, and it is worth slowing down for. The model builds every answer one token at a time. At each step it asks, in effect: given everything so far, what token would fit next? Usually several tokens would fit. The model picks among them — and here is the part nobody expects — with a deliberate touch of chance. Then it asks the question again for the next token, and the next, thousands of times, until the answer is done. An answer is not retrieved. It is assembled, choice by weighted choice, fresh each time.
A Weighted Roll, Not a Lookup
Picture asking a good improviser the same question twice. Both answers are in character, both are correct in spirit, and neither is a recording — the second performance simply took different turns. The model works the same way, mechanically. At some early step, two or three tokens were all plausible; chance picked differently on the second run; and from that fork the two answers kept diverging, each one internally consistent, each one a legitimate assembly of the patterns.
This is why the word "guessing" — which people reach for when they learn this — is the wrong picture. The choices are not wild; they are drawn from what fits, weighted by how well it fits. Call it controlled chance: variation among plausible continuations, not chaos. The distinction matters, because chaos would make the tool useless, and controlled chance is precisely what makes it useful.
Why Chance Is There on Purpose
The natural follow-up: why would anyone build it this way? Because the alternative was tried, and it reads badly. A model forced to always take the single most likely token produces prose that is stiff, repetitive, and strangely dead — the most probable next word, it turns out, is often the most boring one. A measured dose of randomness is what makes the writing feel alive, lets a second attempt actually differ from the first, and gives brainstorming its variety.
So the randomness is a setting, chosen by whoever built the product — and it is adjustable. There is a dial behind every chat box that controls how adventurous the token-picking is, turned low for tasks that need consistency and higher for tasks that need spark. Tessa will meet that dial properly in Chapter 8, when products come apart into their pieces. For now it is enough to know the variation is by design, and someone chose the amount you are getting.
What This Means at the Keyboard
Three practical consequences, all of them liberating. First: re-asking is a legitimate technique, not cheating. The second run is a genuinely new assembly; when an answer feels almost-right, another roll of the same question is often faster than surgery on the first attempt. The regenerate button exists precisely because of this page. Second: your colleague getting a different answer to the same question means nothing is broken — two runs differ by design, and neither of you got the "wrong" one.
Third, and this one carries a warning: since every answer is assembled fresh, consistency is something you build, not something you get. When Tessa needs the same shape of output every time — and by Chapter 5 she will, badly — the sameness will have to come from her side of the window: fixed instructions, shown formats, a turned-down dial. And one more thing follows from assembly-not-retrieval, something big enough to get its own chapter: if answers are assembled to fit patterns rather than looked up from a store of facts, what exactly makes them true? Hold that thought. It is Chapter 3.
- "Different answers mean it's guessing wildly." The variation is among plausible continuations, weighted by fit — controlled chance, not chaos. Both of Tessa's subject-line sets were legitimate assemblies.
- "One of the two answers is the real one." There is no stored answer to be "the real one." Every answer is built fresh, token by token; truth is a separate question, checked against the world (Chapter 3).
- "Computers are exact — something must be broken." The randomness is deliberate, added because always-most-likely text reads stiff and dead. It is a designed feature with an adjustable dial, not a defect.
- "If I ask better, I'll get the one true answer every time." Phrasing shapes the answers but cannot remove the variation. When you need consistency, you build it — fixed instructions, shown formats, a lower dial (Chapters 5 and 8).
- Variability explains the regenerate button, the changed answer, and the colleague who got something different — a whole family of "is it broken?" moments that now read as design.
- Re-asking becomes a deliberate technique instead of a guilty habit — often the fastest way from almost-right to right.
- It sets up the book's central question about trust: a system that assembles answers fresh each time must be verified differently from one that looks answers up. That is Chapter 3's whole subject.
Knowledge Check
How does the model produce an answer?
- It finds the best matching answer in its stored collection and returns it
- It builds the answer one token at a time, choosing among tokens that fit
- It writes several complete drafts internally and shows you the best one
- It copies the closest passage from the documents it was trained on
Why do two runs of the identical question produce different answers?
- The model remembers the first answer and deliberately avoids repeating it
- Tiny rounding errors build up differently on each run and shift the wording
- Chance picked a different plausible token at some step, and the answers diverged from there
- The provider's servers were busier the second time, changing the result
Why is a deliberate dose of randomness built into the token choices?
- Because always picking the most likely token makes prose stiff and dead
- Because picking at random is faster to compute than weighing every candidate
- To stop anyone proving what the model said to them in a dispute
- So that two users asking the same thing get different answers, and neither is favoured
Tessa's draft is almost right. What does this page suggest is often the fastest move?
- Open a fresh chat and rebuild the whole request from the beginning, in more detail
- Ask the model to slow down and go through the draft more carefully this time
- Retype the question word for word, in case a typo in the first one spoiled it
- Simply run the same question again, since a fresh assembly may land closer to it
You got correct