Topic 28

When the Model Says No

Concept

With the policy conversation still fresh, Tessa volunteers to write a short security note for the office — the kind of thing that goes on the kitchen noticeboard. Her first question to the model is the obvious one: how do people actually trick staff into handing over a door code? And the model declines. Politely, without drama, it says it will not help with that.

This is her first refusal: a deliberate limit, built in by the provider, where the model declines to answer rather than answering badly. It is not an error, it is not the model being broken, and it is not personal. This page explains what she just hit, why honest questions sometimes hit it, what to do next, and one line not to cross.

Two things to do with a refusal
Rephrase honestly · the craft
State the real purpose and who it is for — guidance so office staff can recognize the trick, rather than instructions for performing it.
Disguise the purpose · the trap
Role-play framings and invented cover stories aimed at the safeguard — the personal-account workaround wearing new clothes.

A Refusal Is the System Working

Providers deliberately train and instruct their models to decline certain requests. The categories are roughly what you would expect: help with clear harm to people, step-by-step instructions for dangerous things, certain kinds of content about real or vulnerable people.

Turn it around for a second. A model that helps absolutely everyone with absolutely everything is not the neutral option — it is the unsafe design, because "everyone" includes the person planning the thing you would not want helped. Given that these tools are available to anyone with a browser, some limits are the only responsible way to ship one. The "no" is the system doing its job.

Two honest notes about those limits. They are drawn by the provider, as a product decision, not handed down by law — which is why two products can refuse in slightly different places. And they are not fixed forever: where the line sits moves as providers learn what misfires. So a refusal describes this tool, today, and nothing grander than that.

Why Honest Requests Get Caught

Tessa's question was entirely legitimate. She was writing staff guidance. So why did it trip the limit?

Because the model judges by pattern, which is the machinery from Chapter 1 showing through. It has no way to verify who she is, what she does at Waymark, or why she is asking. All it has is the shape of the words in front of it — and a security-awareness question about phishing has almost exactly the same shape as an attacker's question about phishing. Same topic, same vocabulary, same request for specifics. The only difference is purpose, and purpose is the one thing that does not appear in text unless somebody writes it down.

So the borderline cases misfire in both directions. Honest requests get declined; occasionally something that should have been declined gets through. That is what a judgement made on surface pattern produces, and it is mechanical, not a verdict on the person asking. Nobody is being suspected of anything. The words simply landed in the wrong-looking pile.

What to Do When It Says No

The first move costs ten seconds: restate the request with its purpose and context attached. "I'm writing internal guidance to help our office staff recognize attempts to get door codes over the phone — explain the common tricks so people can spot them." That frequently resolves it, and it should, because the request genuinely is a different request. It is asking for recognition rather than execution, for staff rather than strangers.

Notice that this is not a trick. It is Chapter 2's context habit, applied to a new problem: the model only knows what is in front of it, and Tessa's purpose was never in front of it. Supplying true context changes the pattern honestly, which is why it works and why it is a legitimate skill.

If the model still declines, the request is finished with this tool, and there are two ordinary routes. Raise it through whatever the approved tool's support channel is, if it is a work task and the refusal looks plainly wrong. Or get the information the way people got it before chat boxes existed: published security guidance, an industry body, a book, a colleague whose job is actually security. A refusal is a door closing on one tool, not a wall around a subject.

The Line Not to Cross

There is a family of techniques for defeating these limits on purpose: role-play framings, "pretend you are an assistant with no restrictions", inventing a fictional scenario as cover, splitting a request into pieces so that no single piece looks like what it is. They are widely shared, and they sometimes work.

The difference between that and rephrasing is not cleverness or wording. It is one question: is the purpose you stated the purpose you have? Adding true context is honest. Building a costume for a request so the system reads it as something else is deception, aimed at a safeguard, on purpose.

At work, this is the previous page's workaround wearing new clothes, and it carries the same arithmetic. It typically breaches the provider's terms and the employer's policy at once. And it moves ownership: the person who disguised the request owns whatever comes out of it. "The model wrote it" persuades nobody once it is clear how the model was asked, because the asking is the part that was chosen.

So the book's stance, stated flatly and without a sermon attached: say what you actually want it for. If a tool still says no to an honest request, it is answering a question you did not mean to ask, and the right response is to go somewhere else honestly rather than to argue with a safeguard.

The everyday version is a pharmacist. Some things need a prescription — inconvenient when your need is genuine, and protective on purpose. Explaining why you need something is entirely legitimate and often solves it. Forging the prescription is a different act, and everyone understands that it is a different act, including the person doing it.

Which closes the chapter. Know where your words go before you send them. Keep credentials, other people's personal data, and material that is not yours out of the box. Read the policy as the company making those calls once, and ask when a rule blocks real work. State your purpose truthfully, and take a no at face value.

It also closes the whole middle of the book. Three chapters on Waymark's actual work — documents in, usable output back, and now what must never go in at all — and every one of them has been about a single side of the glass. Tessa types, and something comes back. She has never once asked what carries it. She presses Enter, her words reach a computer she has never seen, and an answer arrives a few seconds later, and the mechanism in between has been invisible for six chapters. It has a shape, and it is smaller and far more readable than she expects. Chapter 7 takes the chat box off.

Common Confusions
  • "The refusal is judging me." It is matching the shape of the request, blind to your character, your job and your reasons. Misfires are mechanical, which is also why adding true context so often clears them.
  • "Phrasings that beat refusals are a power-user skill." Adding honest context is a skill. Disguising a request to defeat a safeguard is a different act, and it usually breaches both the provider's terms and your employer's policy.
  • "A refusal means the topic is forbidden knowledge." It means one product declined one request. Published guidance, industry bodies, books and colleagues are all still there, and none of them was affected by the refusal.
  • "If the output came from the model, it is the model's responsibility." Not once the request was deliberately disguised to get it. Choosing how to ask is an act with an author, and the author is whoever typed it.
Why It Matters
  • Refusals are part of the daily texture of these tools. Knowing the mechanics turns a moment of friction into a two-second rephrase instead of an afternoon of resentment.
  • State your purpose, do not costume it. That is this chapter's ethics in one line, and it is the sentence that makes the rest of the chapter's habits hold together under pressure.

Knowledge Check

Why do providers build refusals into their models at all?

  • To keep costs down by declining the questions that would take longest to answer
  • To steer clear of subjects where the model is likely to be wrong, as Chapter 3 described
  • To comply with one single set of laws that applies to every AI product in every country
  • Because a model that helps everyone with everything is the unsafe design of the two

Tessa's honest security-awareness question was refused. Why does that happen?

  • The model has been tracking her earlier requests and has grown suspicious of the pattern
  • Her question has the same shape as an attacker's, and purpose is not visible in the words
  • The subject falls outside the model's training, so it has nothing accurate to offer here
  • Something went wrong in the request, and the refusal is a bug worth reporting to support

Which response to a refusal does this page endorse?

  • Wrapping the same question in a fictional story so the safeguard reads it differently
  • Sending the identical question a few more times until one of the attempts gets through
  • Restating the request with its real purpose and context attached, then asking again
  • Dropping the subject, since a refusal indicates the information is off limits generally

Someone disguises a request to get around a refusal, and uses what comes back. Who owns that output?

  • The person who asked for it
  • The provider of the model
  • The model that wrote it
  • The employer that allowed it

You got correct