Settings That Change the Answer
Two products at Waymark draft emails. The CRM's suggest-a-reply is reliably crisp: ask it the same thing twice and you get almost the same sentence twice. The marketing tool Tessa uses for campaign lines is the opposite — run it three times and you get three genuinely different ideas, one of which is usually worth keeping.
For a year she assumed the marketing tool had the better model. It does not. It may well be the same model. What differs is a pair of numbers that ride along with every request, chosen by whoever built each feature.
Tessa met them in Chapter 7, sitting at the bottom of the payload under the message list, and that page deliberately did not explain them. This page does. There are two that matter, they explain nearly everything a user experiences, and both are ordinary enough to describe in a paragraph each.
The Randomness Dial
Chapter 1 established how the model produces text: one piece at a time, each time choosing among several plausible continuations rather than the single most likely one, which is why the same question gives different answers on different days. The chapter said someone chooses how much variation there is. This is that choice.
The randomness dial sets how adventurous the picking is. Turn it low and the model takes the safest continuation almost every time: focused, repeatable, dull in the way a form is dull. Turn it high and it reaches further down the list of plausible next pieces: varied, surprising, occasionally strange.
Its usual name in the documentation is temperature, and it is worth knowing that word once so that a vendor's page or a settings screen does not read as a foreign language. It comes from a piece of physics with no relevance here, it means nothing warmer or cooler, and after this paragraph this book calls it the randomness dial.
The choice follows the task, and it is not subtle.
Low is for work with a right answer and a fixed shape. Sorting a review into one of four buckets. Pulling five fields out of a document. Anything where two runs over the same input ought to agree, and where a creative flourish is a defect. Tessa's review template wants the dial as low as it goes.
High is for work where the point is options. Six taglines for the Lakeland tour, and five of them can be bad if the sixth is good. Names, angles, opening lines — anything where you are going to choose, and sameness is the failure.
Products choose per feature, which is exactly why the CRM and the marketing tool behave like different species. The raw chat box picks something in the middle, because it has no idea which of those two things you are about to ask for.
The Length Cap
The second setting is blunter and it caused a mystery Tessa has been carrying since Chapter 7.
The length cap is a hard ceiling on how much the model may write in one reply, counted in tokens. It has two effects and they are worth separating, because most people only know about the good one.
The good one is money. Output is the expensive half of the bill, and a cap is the only mechanism that puts an absolute limit on it. No matter what the model was minded to produce, no matter how a user phrased the request, the reply cannot exceed the ceiling. For a vendor running a feature across ten thousand customers, that is not a preference, it is a budget.
The other effect is the mystery. When the cap bites, the model does not wrap up gracefully. It stops. Mid-sentence, mid-word, wherever it happened to be when the ceiling arrived — which is precisely the stop reason Tessa read in the response body two pages ago, the one that says the answer ran out of room rather than finished.
So the truncated reply she once got from a product support bot, the one that ended halfway through a sentence about baggage, was never a crash and never a network problem. It was a cap set lower than the answer needed, and now she can say so.
Which leads to the thing the cap is not. It is not a way to ask for a short answer. Asking for a short answer is Chapter 2's job, done in words: answer in three sentences, one line per review. That produces a complete short answer. A cap produces an answer that ends, and whether it ends anywhere sensible is luck. Sensible builders do both — ask for brevity in the prompt, and set the cap above what a good answer needs, as a stop against runaway output rather than as an editor.
Waymark's office oven has two knobs, and between them they explain nearly every outcome that ever comes out of it: how hot, and for how long. There are other controls on the panel — a fan setting, a grill, something with a picture of a chicken — and in four years nobody has touched them. The two knobs are not a simplification of baking. They are the part of baking you actually operate.
Where the Settings Get Set
Same two fields, three different places, and only one of them is visible to a user.
In a product, the vendor chooses per feature and you never see it. The suggest-a-reply button was built by someone who set the dial low on purpose, because a reply that reinvents itself every time you press the button is an annoying reply. That decision was made once, at the vendor, and every customer lives inside it.
In a direct request — the payload from Chapter 7 — whoever sends it chooses per call. When Milo writes Tessa's review script on the next page, he sets the dial as low as it goes and the cap tight, and both choices come straight off this page: 2,300 identical extractions want repeatability, and a one-line answer does not need room for a paragraph.
And a few chat products expose a version of the dial to users, usually as a creativity slider or a choice between a precise mode and an exploratory one. That is this setting wearing a friendlier name. If Tessa ever meets one, she now knows what it does and which way to turn it for which task.
The Dials That Stay in the Box
Providers offer a handful of further settings beyond these two. They control finer details of how the next piece of text gets chosen, and one or two of them overlap heavily with the randomness dial.
This book leaves them alone, deliberately, and the reason is honest rather than protective. The defaults are sensible. Products almost never expose them. And the two settings on this page account for nearly everything anyone actually experiences using these tools — the difference between crisp and surprising, and the answer that stops mid-sentence.
They are not a secret and they are not advanced magic; they are a page in the provider's documentation, waiting for the day somebody has a specific reason to read it. That day belongs to whoever builds these systems for a living, which is the next course on this shelf rather than this one.
What Tessa leaves with is the last piece of a four-part explanation she started building two pages ago. A product's behaviour is the model, plus the standing instructions, plus the gathered context, plus these settings. Nothing else is in there. Every AI feature she will ever be shown is some arrangement of those four things — and on the next page, so is hers.
- "A low randomness dial makes the answers more accurate." It makes them more consistent, which is a different property. A low dial repeats a wrong answer just as faithfully as a right one, and accuracy remains verification's job.
- "The cap makes the model concise." It makes the model stop, mid-sentence if that is where the ceiling lands. Concision is something you ask for in words; the cap is a ceiling, not an editor.
- "I need to master all the settings before using these tools well." Two of them explain nearly everything a user meets, and the rest are defaults that products do not even expose. Prompting well pays far more than knob-turning.
- "A product that answers differently every time is broken." It is a feature with the dial set high, on purpose, probably for a task where options beat repetition. Broken looks like an error, not like variety.
- These two dials complete the account of why any AI feature behaves as it does — model, instructions, context, settings — which is the whole of product literacy and the end of guessing about it.
- They are also the configuration of Tessa's own script. When the next page sets the dial low and the cap tight for 2,300 classifications, both choices are hers to understand rather than to accept.
Knowledge Check
What does turning the randomness dial down actually do?
- It makes the model check its answers, so fewer are wrong
- It shortens the reply by cutting the weaker sentences
- It makes the model pick safe continuations, so runs look alike
- It switches the request to a smaller and cheaper model to answer
A product's answer ends in the middle of a sentence. What is the most likely cause?
- The reply hit the length cap set in the request
- The account went over its allowed rate of requests
- Something failed inside the provider while writing
- The conversation grew too long for the context window
Tessa wants six different taglines for a tour, and she wants them to be genuinely different from each other. What should the dial be?
- Low, because a low dial is the safe default for creative work
- Left alone, because the model adjusts it to suit the task it is given
- Irrelevant, because the length cap is what controls variety
- High, because the whole value of the batch is that they differ
Where are these two settings chosen for a feature inside a product?
- By the provider, which sets them per model for all of its customers
- By the vendor, once, when the feature was built
- By the user, in the product's own settings screen
- By the model, which adapts them to each question
You got correct