Topic 77

Build or Buy

Procurement

There is a hosted agent product for support queues, several of them, and for a company Sundry's size any of them is a serious option. Vera evaluated three before writing a line of the loop, which is the right order and not the order most teams use. The question is not whether a platform can resolve a ticket about a cracked shelving unit. It can.

The decision is about which parts of this book you want to own and which parts you are content to accept somebody else's answers on. That reframing does more work than any feature comparison, because it turns a procurement exercise into a list of things your business actually depends on — and because the parts that took Sundry the longest are the parts that stay yours in every scenario.

The decision is which parts of this book you must own — four conditions decide it, one crossover prices it, and four things never transfer either way
No engineering team, or one small enough that a support agent would consume itBuy
A standard queue — order status, returns, shipping — and rules the platform already modelsBuy
Below the crossover: somewhere between 800 and 2,600 tickets a week on these quotes, because the fee scales with volume and the engineering does notBuy, and faster
The domain rules are the business — a body of policy nobody else has, and a configuration surface would be the ceiling on your productOwn the loop
Consequences the vendor's model of authority cannot express: a ceiling that applies per customer per ticket, an approval that expires declined, an intent that settles exactly once across a four-hour pauseOwn the loop
At 4,200 tickets a week with engineering time in the sheet on both sides: $242,000 against $107,000 a yearOwn the loop
Tracing, eval tooling, prompt and version management, run persistence — undifferentiated, and rebuilt out of pride for a fortnightBuy either way
The tool surface, the policy library, the eval set and the approval design — most of the elapsed time on the buildYours either way
A year has passed, volumes have moved and platforms have added policy expressivenessRe-run the table

What a Platform Actually Sells

Five things, and you have built four of them. A loop, hardened and hosted. A tool catalogue with connectors into the systems support teams usually run, which is where the integration depth lives. An eval and testing surface. Tracing, dashboards and a review queue. And a user interface for the people who work the queue — the one component this book never touched, and the one a support team notices on day one.

The differentiators are narrower than the marketing suggests. Integration depth is the first: a platform already wired into your helpdesk, your order system and your payment provider is selling you weeks of connector work that has nothing to do with agents. The second is how much of the policy layer it lets you express. Sundry needs a refund ceiling that is an integer comparison in code, an authorization check against the ticket's verified customer rather than the order id the model supplied, and an approval item that shows the four facts from Chapter 12 in a form a lead can read in fifteen seconds. A platform that offers a refund limit but no way to say "and only for the customer who opened this ticket" cannot express Sundry's policy, and no amount of prompt configuration closes that gap.

What Never Transfers

Four things stay yours whichever way the decision goes, and between them they were most of the elapsed time on Sundry's build. The tool surface is first: nine tools with argued boundaries, error messages written for a reader who must act, and shaped results that do not blow the context budget. That design is Chapter 3, it is specific to Sundry's systems, and a platform's connector gives you a call to your order API without giving you the judgement about what get_order should return.

Then the policy library and the retrieval over it, which is where Sundry's actual business rules live: a 30-day own-stock window, a statutory 14-day right, a damage procedure, a marketplace refund matrix and around 2,000 per-seller supplements, indexed with scope and effective dates so a supplement can outrank a general rule. Then the eval set — 120 graded tickets, a rubric rewritten three times, and a judge whose agreement with human labels is measured rather than assumed. Then the approval design: which four categories require a person, what the reviewer is shown, and what happens when a decision never arrives.

The eval set is the one teams assume transfers and the one that transfers worst, because a case is graded against a trajectory — which tools were called, in what order, with what arguments — and those names belong to the surface it was built on. Chapter 9 measured outcome and trajectory separately for exactly this reason, and outcome-level grading is the half that survives a platform change. Write it that way from the start if you think a migration is ever likely: grade "was the buyer refunded the right amount from the right source, with the governing clause cited" rather than "did it call search_policy before issue_refund".

The Real Cost Comparison

Price both sides fully or the exercise is theatre. Hosted support agents are sold per resolution, and the quotes Sundry collected in 2026 ran from $0.60 to $1.80 depending on volume and term. At 4,200 tickets a week and 88% resolution that is 192,000 resolutions a year, so the fee alone lands between $115,000 and $346,000. Against that, the owned agent costs $0.06 a ticket in model spend — $13,000 a year — plus the engineering nobody puts in the spreadsheet. Sundry's build was two engineers for eleven weeks; steady state is about a day and a half a week of one engineer, and that is the line that makes the comparison honest.

Annual line item, at 4,200 tickets a weekBuyBuild
Platform fee, 192,000 resolutions at $1.10$211,000
Model spend at $0.06 a ticketIncluded in the fee$13,000
First-year engineering: tools, policy, evals, approval45 engineer-days110 engineer-days
Ongoing engineering~0.5 day a week~1.5 days a week
Steady-state total, at $1,200 an engineer-day$242,000$107,000

Read the last row against the first column of the table rather than against zero. Buying does not remove engineering; it removes about two thirds of it, because the tool integrations, the policy library and the eval set still have to be built by somebody who understands Sundry. What the platform removes is the loop, the tracing, the persistence and the queue interface, and those are worth real money — just not the whole difference.

Because the fee scales with volume and the engineering does not, there is a crossover, and it is worth computing for your own queue rather than borrowing Sundry's. On these numbers it sits somewhere between 800 and 2,600 tickets a week depending on where the per-resolution quote lands. Below that, buying is cheaper and probably also faster. Sundry runs at 4,200, so the arithmetic pointed one way — and the arithmetic was not what decided it. Two other costs belong in the same sheet: the switching cost of an eval set coupled to somebody else's internals, and the definition of a resolution in the contract, because a vendor that bills a deflection as a resolution is charging you for a number Chapter 9 would not have counted.

The Middle Path

Buy the surrounding infrastructure, own the loop and the domain logic. That is the same conclusion the previous topic reached from the framework direction, arrived at here from procurement, and the agreement is not a coincidence: both questions are asking which code you can afford not to understand. Tracing platforms, eval tooling, prompt and version management, and run persistence are bought without regret. The loop, the dispatcher, the tool surface, the policy retrieval and the approval design are where your product's behaviour is decided, and every incident in this book landed in one of them.

One practical note on doing this well. When you buy the surrounding tools, buy them on the same criteria the previous topic listed — can you export your traces, does the eval set live in your repository, can you leave with your own data in a format you can read. A hosted eval tool holding the only copy of your 120 graded cases is the same lock-in as a platform, for a fraction of the value.

When Buying Is Clearly Right

Three conditions, and when all three hold the answer is not close. No engineering team, or one small enough that a support agent would consume it. A standard queue: order status, returns, shipping, the same shapes every retailer has. And no unusual policy, meaning your rules are the ones the platform already models rather than 2,000 per-seller supplements with conflicting effective dates. Under those conditions a platform gets you to a working, monitored, staffed agent in weeks, and building the same thing is an expensive way to learn what you could have read in this book.

It is clearly wrong in two situations. When the domain rules are the business — when the thing that makes your support good is a body of policy nobody else has — you cannot outsource the layer that expresses it, and a platform's configuration surface will be the ceiling on your product. And when actions carry consequences the vendor's model of authority cannot express: a $150 ceiling that must apply per customer per ticket, an approval that must expire into declined rather than approved, an intent record that settles exactly once across a four-hour pause. Chapter 12 built all three, and each is a sentence in a contract negotiation you will lose.

Revisit the decision annually rather than treating it as identity. Volumes move, platforms add policy expressiveness, and an owned agent that has stopped needing a day and a half a week is a different line in the sheet than one that needs three. Sundry's decision is re-run each January with the same table and current quotes, which takes an afternoon and has flipped the recommendation for exactly one of Sundry's queues — the seller-onboarding queue, at 210 tickets a week, which now runs on a bought product.

Common Mistakes
  • Comparing a platform fee against zero — an owned agent costs a day and a half a week of engineering forever, and a comparison that omits it makes the wrong answer look obvious.
  • Assuming the eval set transfers — trajectory-graded cases name the tools they were built against, and a platform migration invalidates the instrument you would measure the migration with.
  • Buying to avoid the domain work — the policy library, the tool boundaries and the approval design are the work, they took most of Sundry's eleven weeks, and no platform does them for you.
  • Building because it is interesting — a good reason to prototype for a fortnight and a bad reason to own a production system that pages somebody at 2am for the next three years.
Best Practices
  • Write down which parts of this book you must own for business reasons before looking at any vendor, and buy everything that is not on that list.
  • Price both options with engineering time, first-year build and switching cost included, and compute the crossover volume for your own queue.
  • Grade the eval set on outcomes rather than trajectories wherever you can, so the instrument survives a change of tool surface or platform.
  • Re-run the comparison every year with current quotes and current volumes, and treat a flip as an ordinary result rather than a reversal.
Comparable toolsHelpdesk agent products the loop inside the ticketing system you already runProvider agent platforms the loop hosted, with the model included in the feeEval and tracing services the middle path, bought rather than builtAutomation platforms fixed workflows, for the steps that are knowableYour own loop Chapters 1 to 13, around 300 lines

Knowledge Check

Which part of Sundry's system stays the team's own work whether they build or buy?

  • The persistence layer that survives a restart, since every platform stores state differently
  • The policy library and the retrieval over it, including the per-seller supplements and their dates
  • The tracing and dashboards, because a support queue needs its own view of every run
  • The interface the support team works the queue in, since the workflow is specific to Sundry

A team compares a platform quote of $190,000 a year against $14,000 of model spend and concludes that building saves $176,000. What is wrong with that?

  • The platform quote is a list price, and volume commitments would bring the fee down substantially
  • The model spend will rise as the queue grows, so the saving shrinks every year the product runs
  • Building costs ongoing engineering, and buying still leaves the tools, the policy and the evals
  • The platform would resolve a different share of the queue, so the two figures are not comparable

Under which conditions does this book say buying is clearly the right call?

  • High ticket volume, because a platform's per-resolution fee falls fastest at large scale
  • No engineering team to speak of, a standard queue shape, and no unusual policy to express
  • Actions with money attached, since a vendor carries the liability for an incorrect refund
  • A large and unusual policy library, which a platform's configuration surface handles for you

Why does an eval set built against your own agent transfer badly to a bought platform?

  • The set is too small to be meaningful next to the benchmark suite that a platform already provides
  • The judge model is tuned to your agent's writing style and scores a platform's replies lower
  • The tickets contain internal identifiers that a hosted platform would not be permitted to process
  • Trajectory grading names your tools and their ordering, which a different surface does not have

You got correct