The Bill: Paying per Token
Waymark's first month on the provider's account closes, and the invoice comes to $19. Tessa's manager asks the question managers ask about every technology bill: what is this?
And for once there is a complete answer. Not an estimate, not "usage", not a shrug — a full reconciliation, down to which afternoon spent what. The meter has been visible since two pages ago: the usage block in every response, counting tokens in and tokens out. Add them all up and you get the invoice, because that is literally what the invoice is.
This page reconciles the $19 line by line, then turns to the three things that decide what next month costs.
The Pricing Model
The whole scheme fits in three sentences, and everything else is multiplication.
You are charged per token, in bulk — the published figures are per million tokens, which is why single requests cost fractions of a cent and nobody quotes them individually. What you send and what comes back are priced separately, and output costs several times more per token than input. And bigger models cost more per token than smaller ones, often by a wide margin, for identical text.
That is it. Tokens in, tokens out, at rates that depend on which model you chose. As of 2026 the actual figures keep moving and differ between providers, which is exactly why this book prints none of them: the numbers would be wrong within the year, and the structure will not be. Learn the structure, look the numbers up on the day you need them.
One implication is worth stating on its own, because it surprises people. The question you ask is usually cheap. The answer is where the money goes — priced higher per token and often far longer than what you sent. Asking a long question and requesting a short answer is a genuinely different bill from the reverse.
Reconciling the $19
Waymark's month had three kinds of work in it, and each has its own character on the invoice.
The classification trial: $4. Fifty reviews, one request each, run again through four successive versions of the template on the largest model, to see whether the whole 2,300 would be worth doing by program. Every request carried the same instruction template plus one short review, and got back a single line of labelled data. Input-heavy, output-tiny — the cheapest possible shape of work, which is why four rounds of tuning cost the price of a coffee and the full run, on a sensibly chosen model, will cost less than the trial did.
Drafting: $7. A month of tour descriptions, guest emails and social posts, produced in short conversations. Short asks, long answers — the opposite shape, and the reason this line is nearly twice the trial despite being far less text overall. Output is where the money is.
The contract week: $8. One week in which Tessa worked through the 120-page supplier contract, feeding the model long extracts and asking question after question in the same conversation. This is the expensive line, and its expense is not an accident — it is Chapter 1 arriving on an invoice.
Follow that line all the way back, because it closes a loop the book opened in its first chapter. The model holds nothing between requests, so the chat box re-sends the whole conversation every turn. The conversation contained a forty-thousand-token contract extract. So the tenth question in that thread did not send one question — it sent the extract, plus nine questions, plus nine answers, plus the new question, and paid for every token of it. The first question in a long thread is cheap. The fortieth is not, and it is not because the model is working harder. It is because the same text keeps being sent.
Four, seven and eight. Nineteen dollars, fully explained, for a month in which a forty-person company did real work.
The Three Dials
Everything anyone can do about cost turns one of three dials, and none of them requires programming.
Send less. This is Chapter 4's relevance-cutting, now with a price attached. The three clauses that bear on the question rather than the whole section. A tighter instruction template rather than a discursive one. A fresh conversation for a new topic, instead of dragging an hour of unrelated history through every request. That last one is the single cheapest habit in this book, and most people already know they should do it for quality reasons.
Receive less. Cap the length. Ask for one line per review rather than a paragraph of reasoning about each. "Answer in three sentences" is not only a style preference; it is a lever on the more expensive half of the bill.
Right-size the model. The largest model is not required for most work. Sorting a review into one of four buckets is not a task that needs the flagship, and a smaller, cheaper model does it just as accurately — a claim Tessa should check on a sample rather than take on trust, which is precisely the checking habit Chapter 5 built. Drafting that a human will edit anyway is another candidate. Reserve the expensive model for the work that visibly benefits.
Turned deliberately, those three dials take this exact workload — the same reviews, the same drafts, the same contract questions — from $19 to $6. Not by doing less work. By sending less text, asking for shorter answers, and choosing the right model for each job. Chapter 8 shows the turning, one dial at a time, on Milo's script.
Chat Plans Versus the Meter
Tessa has been paying for a chat subscription all along: a flat monthly fee, unlimited within some fair-use bounds, no meter anywhere in sight. The API account bills by what it measures. It is worth being clear about what actually changed between them, because it is less than it appears.
Nothing got more expensive. The chat subscription was never free of these costs — it averaged them across everyone and hid the mechanics behind a flat number. The meter does not add cost; it makes cost visible, one request at a time.
Waymark had a flat-rate arrangement for water at the old office, and the new one is metered. Nobody's shower got dearer. What changed is that a long shower now has a number attached to it — and a number attached is a thing you can decide about, which is exactly what a flat rate never lets you do.
Both arrangements are perfectly reasonable, and at Waymark's volumes both are affordable. Flat is simpler when a person is doing the asking. Metered is what programs use, and it rewards every single habit this book has taught: shorter context, tighter instructions, the right tool for the task. Knowing why is the point. The chapter began with Milo turning a laptop around, and it ends with a bill that Tessa can read line by line — which means there is nothing left under the chat box that is dark to her.
- "Working with models is expensive." Waymark's entire first month, classification trial included, came to less than a team lunch. The usual problem is not that this work is costly but that it is unmeasured, which is a different complaint with a different fix.
- "Only my questions cost money; the answers are free." Output is billed too, and at a higher rate per token than input. Asking for concise answers is one of the few cost levers available to someone who writes no code at all.
- "Every message in a chat costs about the same." Each turn re-sends the whole conversation, so a late message in a long thread bills a multiple of an early one. It is the quietest line item on any invoice, and the easiest to avoid.
- "The biggest model is the safe default." It is the expensive default, and for sorting, tagging and first drafts a smaller one is usually indistinguishable. Check on a sample rather than assuming either way.
- Cost literacy completes the wire. Tessa can now read the request, the response, the errors and the invoice — and an invoice she can explain is an invoice she can argue for at a budget meeting.
- The three dials are the practical takeaway that managers actually need: cost here is a design choice rather than a fixed price, and the next chapter proves it by turning $19 into $6 without doing less work.
Knowledge Check
How are input and output priced?
- At the same rate, since a token is a token whichever way it travels
- Only input is charged, because the answer is what you are paying to receive
- Separately, with output costing several times more per token than input
- Separately, with the rate falling once a request passes a certain length
Why was the contract week the most expensive line on Waymark's invoice?
- Because contract language is harder for the model to read than ordinary prose
- Because legal work requires the largest and most expensive model available
- Because the answers that week were far longer than anything else that month
- Because each question re-sent the long extract and the entire thread with it
Which of these is one of the three cost dials?
- Sending requests more slowly across the working day
- Choosing a smaller model where a big one is not needed
- Retrying failed requests until one of them succeeds cheaply
- Splitting the work across several keys to spread the cost
What actually changes when you move from a flat chat plan to a metered account?
- The cost becomes visible rather than becoming higher
- The model becomes more capable than the chat version
- The limits on how fast you may ask disappear entirely
- The same work starts costing substantially more money
You got correct