Topic 15

Bias and the Average Answer

Concept

Tessa is sketching a marketing campaign and needs a customer to write for, so she asks the model to describe a typical Waymark customer. What comes back is a retired couple in their sixties, comfortable, from the suburbs, looking for a gentle week away. Pleasant. Plausible. Useless — and quietly expensive.

Waymark's actual bookings last season ran from a university walking club to solo hikers in their thirties to three-generation family groups. The retired couple is in there. So is everyone the answer left out, and the campaign Tessa writes from that sketch will speak to one of them.

The model did not choose the retired couple. It learned from human writing, inherited the regularities of human writing, and handed back the most common pattern when asked for a typical one. This page is about noticing when the average is speaking.

Asking for a typical customer, and asking for the range
One typical customer
A retired couple in their sixties, comfortable, from the suburbs, after a gentle week away. The densest point in the pattern, arriving without a label.
The actual range
A university walking club, solo hikers in their thirties, three-generation family groups, the retired couple, a pair of first-time travellers — all of them booked last season.

Where the Lean Comes From

Chapter 1 said the model learned patterns from an enormous amount of text. That was the whole mechanism, and it does not get a carve-out for the patterns we would rather it had not picked up. Which jobs go with which names. Which customers count as typical. Which neighbourhoods get described as up-and-coming and which get described as rough. Which pronoun follows which profession. All of that is in the text, so all of that is in the patterns.

Call that lean bias: a systematic tilt in what the model produces, inherited from the material it learned on rather than invented by it. Note how ordinary the mechanism is. There is no separate prejudice module bolted on, and nothing malfunctioned when Tessa got her retired couple. Pattern reproduction is the machine's only trick, and this is that trick working exactly as described.

An autocomplete makes it concrete. Feed one every sentence ever typed and start it with "the nurse said" — it offers "she", not out of any opinion, but because that is what followed most often in what it read. No malice, just frequency. Now imagine that same frequency machine writing the customer profiles for your next campaign, and you have the shape of the problem.

The Average as Erasure

Certain words are trapdoors, and Tessa fell through one. Typical. Normal. Standard. The average. Each of those asks the model to collapse a range into one representative, and the representative it returns is the densest point in the pattern — with everyone else silently deleted on the way.

Silently is the operative word. The answer does not arrive labelled "one of several possible customers, chosen by frequency". It arrives as a description, with the confidence of the previous two pages behind it, and it reads like a finding about Waymark rather than a fact about text.

For marketing, that lands twice. There is the ethical cost — stereotype reproduced in public-facing copy, with real people reading it and recognizing who was assumed and who was forgotten. And there is the commercial one, which is the same mistake wearing a suit: the solo hikers and the student groups also had money, and the campaign was not written for them.

Working Against the Lean

The counter-moves are small and they work. Ask for a range instead of an example: five different customer profiles, as distinct from each other as you can make them beats a typical customer every time, because the request no longer rewards the densest point.

Name the dimensions you actually care about, rather than leaving them to be filled in by frequency — age range, group size, budget, travel experience, reason for the trip. Whatever you leave unspecified, the average will supply. And where you have real information, supply it: Waymark's own booking records make a far better basis for a customer sketch than the model's inherited notion of a traveller, and putting them in front of it is exactly the Chapter 2 habit, applied here.

Then read generated depictions of people with the same eye you would bring to picking stock photos for the same campaign — a check you already know how to run, and one you would not skip just because the images looked professional.

Calibration, Not Panic

Two things are true at once, and holding both is the point of this page. The lean is real and it shows up in mundane work, not just in the sensitive topics people expect. And it is a lean to correct for, like sycophancy one page ago — not a reason to put the tool down, and not something a clever instruction fixes permanently.

That last part catches people. Adding "avoid stereotypes" helps that answer, and only that answer. Open a fresh chat tomorrow, ask for a typical customer again, and the average is waiting where you left it. This is a property to keep correcting for, not a setting to switch off — and why the training process cannot simply subtract it is a genuinely deep question that belongs to a different book: Machine Learning from Zero, Chapter 10.

That completes the chapter's toolkit, and it is worth seeing the four together. Hallucination: it invents specifics. Fluency: good writing is not evidence. Sycophancy: it leans toward you. Bias: it leans toward the average. Four leans, one posture — stand upright, ask against the tilt, and get anything that matters from outside the chat box. With that in hand, Chapter 4 finally does the thing this whole book has been circling: puts Waymark's real documents in front of the model.

Common Confusions
  • "The model is neutral — it is just maths." The maths faithfully reproduces the leanings of the text it learned from, and faithful reproduction of a bias is a bias. Neutral machinery does not make neutral output.
  • "Bias only matters on sensitive topics." It shapes ordinary output constantly — who gets pictured as the customer, the boss, the tourist, the complainer. Marketing copy lives precisely in that territory.
  • "I will tell it to be unbiased and move on." An instruction helps that one answer. The lean is back with the next fresh chat, which is why this is a habit of noticing rather than a fix you apply once.
  • "So the answer about the retired couple was wrong." It was not wrong; it was narrow. Retired couples do book Waymark tours. The failure is what the single answer left out, and single answers always leave out.
Why It Matters
  • Tessa's output is public-facing. Inherited stereotypes in Waymark's copy carry a real human cost and a real reputational one, and she is the last checkpoint before they reach a guest.
  • Noticing when the average is speaking completes the trust toolkit. Hallucination, fluency, sycophancy, bias — four leans in one machine, and one habit of standing upright against all of them.

Knowledge Check

Where does the model's bias come from, mechanically?

  • From the patterns in the human text it learned on
  • From a separate filtering layer added by the provider
  • From the earlier questions you asked it in that chat
  • From an error that crept into the training process

Why is asking for "a typical customer" a trapdoor?

  • The model cannot describe people it has no real data about
  • It collapses a range to its densest point, silently dropping the rest
  • The word is vague, so the answer comes back confused and mixed
  • It pushes the model into inventing a customer who does not exist

Which request works best against the lean, for Tessa's campaign sketch?

  • "Describe our typical customer, and please avoid any stereotypes"
  • "Describe a fair and balanced version of our average customer"
  • "Give five distinct customer profiles, varying age, group size and budget"
  • "Describe our typical customer in much more detail than usual"

You add "avoid stereotypes" to a prompt and the answer improves. What does that buy you?

  • A setting that stays applied to every future chat you open
  • A permanent correction to how the model was trained
  • Evidence the lean has been removed from this topic
  • A better answer this time, and nothing beyond that one

You got correct