Check the Output, Don't Trust It
Two hundred rows into the finished file, Tessa stops. The row looks like all the others: a review, a tour, a rating, a bucket, a one-line note. The note says the guest was disappointed by the rooms at the Harbourview Annex.
She goes back to the review it came from. The guest wrote four vague sentences about a trip that did not go well and never named a hotel at all. The Annex is Tessa's old acquaintance from Chapter 3, back in a new costume — no longer a suggestion in a chat, now a value in a cell, sitting in a column, in a file of 2,300 rows that was about to go to the operations meeting.
That is this page. Hallucination does not stop when output gets tidy. It puts on columns.
Structure Launders Errors
A table carries an air of authority that prose does not, and the air is entirely unearned. Rows and columns are the visual language of records — of things counted, filed, checked by somebody. Reading a neat table, your eye reports "this was extracted", and extraction sounds like copying.
Nothing was copied. Every value in every cell was generated, one piece at a time, exactly as Chapter 3 described. A cell is not a smaller, safer kind of text; it is the same text in a box. If anything the box makes it worse, because a five-word cell gives you far less to be suspicious of than five sentences would.
This is the fluent costume from Chapter 3 with a new outfit. There, invented content was protected by good writing. Here it is protected by good formatting, and formatting is even harder to see through — nobody reads a spreadsheet cell critically. They read the column heading, once, and trust every value under it.
Check a Sample, Not the File
The honest question is what checking 2,300 rows can possibly mean, given that reading them all would cost more than the model saved. The answer is the one every factory settled on long ago: you do not check everything, you check a sample properly.
Tessa's version takes twenty minutes. Pick a random dozen rows — genuinely random, not the first twelve, which are the ones the model produced when the pattern was freshest. Open the review each one came from. Ask a plain question of every field: is that what this review actually says?
What the dozen tells you depends on what it finds. One error, and you have not found one bad row — you have found a class of bad row, because the same template produced all 2,300. Tessa's Annex row is not a one-off; it is what her template does with a vague review, and there will be dozens more like it in the file. That is the real value of sampling: a single failure points at the rule that caused it.
Twelve clean rows, on the other hand, are confidence and not proof. They say the common cases work. They say nothing about the rare shapes that were not in your dozen. This is Chapter 3's calibrated trust at the scale of a spreadsheet: you have not verified the file, you have earned a reasonable belief about it, and knowing the difference is the whole discipline.
Think of goods arriving at a warehouse. Nobody opens every box from a supplier they have used for years. They pull a few from each pallet, measure what a gauge can measure, and if a sample is bad the pallet is held rather than shelved. Sampling plus a gauge is not a compromise on quality — it is how quality is actually managed at volume.
The Gauge: Checks That Need No Judgement
Some errors need a person who has read the source. Others need nothing but a rule, and those are worth separating out, because they are cheap and they catch a surprising amount.
The rules for Tessa's file write themselves from the shape she chose on the earlier page. Every rating must be a whole number from one to five. Every category must be one of exactly four words — not a fifth invented on the fly, not "praise (mostly)". Every review identifier must be filled in, and no two rows may carry the same one. Every item must be valid in the labelled format, which the last page's strictness makes an instant, judgement-free test.
None of that asks whether a value is true. A rule-check cannot know that the Annex does not exist. What it can know is that a row is malformed, a category is not on the list, or a rating says seven — and errors of that kind cluster around exactly the inputs that confuse a model, so a rule-check often flags the same rows a human sample would have caught.
Today Tessa runs these by sorting and filtering in a spreadsheet, which works fine and takes ten minutes. Note what kind of task it is, though: fixed rules, applied to every row, with no judgement anywhere. Chapter 8 is going to point that out again.
Close the Loop
Finding the Annex row is not the end of the job, and treating it as one wastes what it cost to find. An error is information about the template.
So Tessa edits the document. The new line tells the model what to do with a review that names nothing: keep whichever of the four buckets fits, leave the specifics empty, and never supply a name the guest did not write. The note at the top of the template gains a version: v4: vague reviews keep their bucket, specifics stay empty, nothing invented. Then she re-runs the batches that the old version produced — not all of them, only the ones the fix affects.
That loop is the craft, and it is worth stating plainly because the instinct it replaces is so common. One bad row does not mean the batch is garbage and the whole idea was a mistake. It means the template had a gap, the gap is now closed, and the next 2,300 will be better than the first — which is more than can be said for most manual processes.
What Tessa ends the chapter with is not a perfect file. It is something better, and rarer: a file with known error behaviour. She can say which rules were enforced, how large a sample was checked, what was found and what was fixed. That sentence is the difference between "the AI made this, I think it's fine" and a piece of work an operations meeting can act on.
Which closes the chapter. Ask for the artifact, not the answer. Ask for a shape a machine can unpack. Save the prompt that produces it. Check before you trust. The next chapter changes the subject entirely, to the one question nobody at Waymark has asked yet: 2,300 guest reviews went into a chat box this afternoon, and nobody checked what they were allowed to send.
- "It is in a table, so it was extracted rather than invented." Generation filled those cells, the same way it filled the paragraphs. Tables hallucinate too; they just do it in smaller boxes, where it is harder to notice.
- "Checking properly would mean reading all 2,300 rows." A random sample plus a set of rule-checks is the honest and affordable standard, and it is the one factories have run on for a century.
- "One bad row means the whole batch is garbage." One bad row is a lead. It names a class of error, the template gets a fix and a version, and only the affected batches are re-run.
- "A clean sample proves the file is correct." It says the common cases work. Twelve clean rows out of 2,300 are grounds for reasonable confidence, and confidence is not the same thing as proof.
- This habit is what makes a model-made file something you can hand to other people — the difference between hoping it is right and being able to say what was checked and what was found.
- Rule-checks a machine could run are the second half of the Chapter 8 blueprint: template in, checked file out, and a person sampling the middle where judgement is actually required.
Knowledge Check
Why does the page say a tidy table is not evidence that its contents are correct?
- Tables usually contain figures, and the model's arithmetic is unreliable in every column
- Table formatting is fragile, so rows are often scrambled between the model and the file
- Cells are too short to hold the full source text, so detail is lost when it is cut down
- Rows and columns look like records, while every cell is still generated text in a box
Tessa's random dozen turns up one bad row. What does that tell her?
- A class of error is in the file, because one template made every row
- That one row is faulty and can simply be corrected on its own
- The batch is unusable, so the whole file should be discarded and rebuilt
- The sample was too small, so a much larger one should be drawn instead
Which of these is a rule-check rather than a judgement-check?
- Whether the one-line note is a fair summary of what the guest actually wrote
- Whether every rating is a whole number between one and five
- Whether a hotel named in a row is a real place that a guest could have stayed in
- Whether a review that grumbles but ends warmly belongs in the right bucket
Tessa has found the invented hotel name. What does the page say she does next?
- Delete the offending row and carry on, since the other rows checked out fine
- Stop classifying reviews this way and go back to reading them by hand instead
- Add a rule to the template, note the new version, and re-run the affected batches
- Ask the model to go back over the file and mark any values it thinks it invented
You got correct