Habits of Verification
Two pages of bad news deserve a page of method. Tessa cannot fact-check every sentence the model writes — she has forty itineraries a month and a job that is not research. If verification means re-doing the work by hand, nobody will do it, and the tool goes back to being a trap.
It does not mean that. In practice, catching most of the damage takes a handful of cheap habits: find the fact the decision rests on, ask for the form of a claim you can look up, and let the stakes decide how hard you look. Verification is a reflex to install, not a research project — and installing it takes one page.
Find the Load-Bearing Fact
Most answers are mostly filler, in the good sense. A hotel recommendation is a paragraph of atmosphere resting on one claim that carries weight: this hotel exists and takes bookings. A summary of a refund policy is several sentences resting on one number: fourteen days. A draft email is entirely harmless except for the date it promises. Find that piece — the load-bearing fact, the one that makes the rest either useful or embarrassing — and check it.
Two minutes on the hotel's own website was the entire Harbourview defence. Not a review of the model's reasoning, not a second opinion, not a policy. One search, one absence, one hotel dropped from the itinerary. The skill being learned here is not checking; it is spotting — reading an answer and seeing which two sentences are holding the weight.
Ask for the Checkable Form
The second habit takes one extra message: where could I verify this? or what is this based on? Be clear about what that does and does not do. It does not make the model honest, and it does not produce a guarantee — a source can be invented as fluently as a hotel, complete with an author, a publisher and a year that all look right.
What it buys you is a shape you can act on. A vague paragraph turns into named claims with named places to look, and looking becomes a thirty-second job instead of an open-ended one. There is a bonus effect worth knowing: asked to point at where a claim comes from, an invention sometimes visibly deflates — the answer softens, generalizes, or quietly drops the specific. That is not proof of anything, but it is a useful flinch to watch for.
Route by Stakes
The reason blanket rules fail is that they ignore what an answer is for. Checking everything is impossible; checking nothing is how the Annex reaches a client. So route by consequence, in three lanes.
Thinking for yourself — brainstorming tour names, drafting your own notes, getting unstuck: read it and go. If it is wrong you will find out cheaply, and nobody else is affected. Anything client-facing or money-adjacent — an itinerary, a quote, a policy summary sent to a guest: check the load-bearing facts before it leaves your hands. Anything contractual, legal, medical or regulatory: the model may draft, and a qualified human owns every fact in it, with no exceptions for a busy afternoon.
That is proofreading logic, and you already have it. When a colleague's report crosses your desk before it goes to a client, you do not re-research it. You check the two numbers the decision hangs on and read the rest with one eyebrow raised. Same reflex, same effort, new colleague.
The Second Run as a Smoke Test
One more trick, cheap enough to use constantly. Ask the same specific question again in a fresh chat — not the same conversation, where everything already said is in view and shapes what comes next (Chapter 1), but a clean start.
Chapter 1 also explained why the answers will not be word-for-word identical: there is genuine variability in how the text gets built. So read the result carefully. A stable answer across fresh runs is not proof — a firmly-learned pattern is repeatable whether or not it is true. But wildly different answers are a loud signal that there is no firm pattern underneath the question at all, and that is exactly the territory where inventions live. Cheap test, one-directional result: it can raise the alarm, it cannot clear you.
Four habits, then, and none of them takes more than a couple of minutes: spot the load-bearing fact, ask where it could be confirmed, route by stakes, and re-ask when something feels assembled. Between "never trust it" and "trust it blindly" there is a calibrated middle, and that middle is where all the real productivity turns out to live. Stakes-routing in particular comes back as the decision frame of Chapter 10; you are learning its everyday version now.
- "Verifying means fact-checking every sentence." It means finding the claims that carry weight and checking those. Total checking is neither possible on a working day nor necessary — most of an answer carries nothing.
- "I asked whether it was sure, and it said yes." Its self-assessment is more generated text, produced by the same process as the claim it is assessing. Verification happens outside the chat box or it has not happened.
- "Asking for sources proves the answer is right." Citations can be invented too. The habit's value is that it converts prose into something you can look up — the looking up is still yours to do.
- "If two runs agree, the answer is confirmed." Agreement means the pattern is stable, not that it is true. The second run is a smoke alarm: it can warn you, it cannot clear you.
- These habits are the practical bridge between refusing to use the model and trusting it blindly. The calibrated middle is where the productivity is, and it is reachable in about two minutes per answer.
- Routing by stakes scales to everything that comes later — documents in Chapter 4, generated tables in Chapter 5, the whole decision frame in Chapter 10. Learn it here on a hotel, use it there on a contract.
Knowledge Check
What is the "load-bearing fact" in an answer?
- The claim the model states first, since answers are ordered by importance
- The longest passage in the answer, because detail is where errors collect
- The one claim that makes the rest either useful or embarrassing
- The claim the model hedged, since hedges mark the shaky parts
You ask "where could I verify this?" and get a named source. What have you gained?
- Confirmation, since the model would not name a source it had made up
- Something specific to look up, which turns checking into a short job
- Evidence that the model searched for the claim before answering you
- A shift of responsibility onto whoever published the source named
Which task belongs in the strictest lane of stakes-routing?
- Brainstorming twenty possible names for a new walking tour
- Turning your own messy meeting notes into a tidy summary
- Drafting a friendly reply to a guest asking about the weather
- Summarizing the cancellation clause of a supplier contract
You re-ask the same question in a fresh chat and get a very different answer. What does that mean?
- The model corrected itself, so the newer answer is the more reliable one
- A warning sign that the answer is being assembled rather than recalled
- Nothing at all, because two runs never produce identical wording anyway
- Run it a third time and take whichever of the two answers appears again
You got correct