AI guardrails written as instructions to the model fail, and you can’t predict when. In my own tests, guardrails written as a script that checks the agent’s output before it acts stopped the failures, as long as the script caught every way the mistake could be written.
Key Takeaways
- AI guardrails written into an agent's instructions are requests. The model treats them as one input among many and can skip them.
- In my tests, filler words got through 6 of 6 times with instructions alone and 0 times once a script removed them.
- Reality check: a script enforces only what you can describe precisely enough for a machine to check. Judgment calls still need a person.
- Before you buy an agent, ask the vendor which of its rules the software enforces and which are only instructions to the model.
What this series is
I run my consulting business with help from AI agents I built myself in Claude Code. They research companies, keep track of my projects, and work with me on my writing. I set the angle, the sources, and the rules before anything gets drafted, we write together, and I do the final read.
What I’m Learning from My AI Experiments is where I share what broke. Each post covers one experiment: the problem, what I wanted, what I did, what happened, how I fixed it, whether that worked, and what a marketing team should take from it. The numbers come from my own test logs. They’re small samples from one person’s setup, so treat them as something to test in your own environment.
This first post covers the lesson the rest of the series keeps coming back to: when I told the agent to follow a rule, it broke the rule now and then, and I couldn’t predict when. When a script checked the agent’s work before I saw it, most of those mistakes stopped.
The problem: my agent kept breaking rules I’d written down
My first AI guardrails were written instructions to the agent: no filler words in drafts, never invent a quote, avoid one sentence pattern I’d banned. The agent read those rules every session and broke them anyway, some more often than others.
The problem was that I couldn’t predict which session would produce the next mistake. So I checked every draft for the same mistakes by hand, time I needed for the argument itself.
What I wanted was easy to state: rules the agent follows every time, so I could read each draft to keep the writing moving instead of combing it for the same mistakes.
What I did: moved the rules out of the instructions
In mid-August I took 3 rules that kept failing and rebuilt each one as a script. A script here is a short program that reads the agent’s output and acts on it before I see it. One deletes banned filler words. One compares every quoted phrase against the source text and rejects the draft if a quote doesn’t match. A third flags a sentence pattern I’d banned.
The difference is where the rule is checked. An instruction is part of what the model reads, so the model decides whether to follow it. A script runs outside the model and applies the same way every time. Anthropic describes the hooks in its Claude Code tool in those terms: they give “deterministic control: certain actions always happen rather than relying on the LLM to choose to run them” (2. Anthropic, 2026).
What happened
Filler words were the clearest case, because a draft full of them reads like a machine wrote it. With the rule in the instructions, filler got through in 6 of 6 tests. With the script, it got through in none, and I stopped looking for it.
Invented quotes mattered more, because a made-up quote attached to a recommendation makes a wrong answer look trustworthy. With the instruction alone, 3 of 5 quoted passages were still wrong. With the script, none were. The banned sentence pattern showed up less often but kept appearing, because my script matched only some of the ways the pattern can be written. A later version that matched more of those ways cut how often it appeared. Testing each version, then fixing what the test found, is how I brought it down.
A different rule taught me the same lesson. Through August, my agent kept recommending things backed by quotes that appeared nowhere in what it had read. I logged that failure 5 times in August, and each time my fix was more instruction text. On September 15 I replaced the instruction with a script that reads every reply before I do. If a reply recommends something, it has to quote a file or page the agent opened in that session, word for word, or the reply is blocked.
That script still catches mistakes. While I was planning this series, it blocked a reply that quoted a project file the agent had seen but hadn’t opened in that session.
Others have found the same thing. Elizabeth Fuentes L, writing on the AWS developer community on DEV, built a travel-booking agent whose instructions said payment had to be verified before a booking was confirmed. The agent confirmed a booking anyway. A check added outside the model blocked the invalid requests in the test without changing a word of the prompt (3. Fuentes L, 2026).
Good or bad: mostly good, with one real limit
The good part is plain. For filler words and invented quotes, the failures stopped, and so did my hand-checking for them.
The limit is that a script enforces only what you can describe precisely enough for a machine to check. Filler words are a list. A quote either matches the source or it doesn’t. Whether a paragraph sounds like me, or whether a recommendation is any good, can’t be reduced to a check like that. Those calls take a writer’s judgment, and no script replaces that, however hard you try.
Scripts also cost upkeep. Each one is code someone has to maintain, and a script with a mistake in it blocks good work as reliably as bad work. My most expensive mistake, and the subject of the next post, was a checking script that rejected writing I already knew was good.
What I learned, and the question to ask your vendors
If a rule matters, enforce it in code that runs outside the model. Keep instructions for preferences, where an occasional miss costs little.
Marketing teams face the same problem across their whole stack. Martech researcher Frans Riemersma describes the hard part of putting AI agents to work in martech as “aligning probabilistic outputs with deterministic systems without breaking control, compliance or consistency” (4. Riemersma, 2026). In plain terms: AI gives a likely answer, your systems of record follow fixed rules, and something has to check the AI’s answer against those rules.
So when a vendor demos an agent and talks about its AI guardrails, ask one question: which of these rules does your software enforce, and which are instructions to the model? Then ask to see what happens when the agent tries to break one. If the answer for anything touching customer data, spending, or messages to customers is “the model is instructed not to,” you’ve found a rule you’ll need to enforce yourself .
Frequently Asked Questions
What are AI guardrails?
Why do AI agents ignore instructions?
What's the difference between a prompt instruction and a code-enforced guardrail?
How do I check a vendor's AI agent guardrails?
What is the What I'm Learning from My AI Experiments series?
References
- Jaroslawicz, D., Whiting, B., Shah, P., & Maamari, K. (2025). How many instructions can LLMs follow at once? arXiv. https://arxiv.org/abs/2507.11538
- Anthropic. (2026). Automate actions with hooks. Claude Code Docs. https://code.claude.com/docs/en/hooks-guide
- Fuentes L, E. (2026). AI agent guardrails: Rules that LLMs cannot bypass. DEV Community. https://dev.to/aws/ai-agent-guardrails-rules-that-llms-cannot-bypass-596d
- Riemersma, F. (2026). Why AI adoption is high but integration is failing in martech. MarTech. https://martech.org/why-ai-adoption-is-high-but-integration-is-failing-in-martech
