AI Experiments

My AI Guardrails Held Only When a Script Checked the Agent's Work

Published: • 6 min read

Bronze statue in the pose of Rodin's Thinker, wearing glasses, beside the headline What I'm Learning from My AI Experiments on a dark background

ChatGPT

AI guardrails written as instructions to the model fail, and you can’t predict when. In my own tests, guardrails written as a script that checks the agent’s output before it acts stopped the failures, as long as the script caught every way the mistake could be written.

Key Takeaways

  • AI guardrails written into an agent's instructions are requests. The model treats them as one input among many and can skip them.
  • In my tests, filler words got through 6 of 6 times with instructions alone and 0 times once a script removed them.
  • Reality check: a script enforces only what you can describe precisely enough for a machine to check. Judgment calls still need a person.
  • Before you buy an agent, ask the vendor which of its rules the software enforces and which are only instructions to the model.

What this series is

I run my consulting business with help from AI agents I built myself in Claude Code. They research companies, keep track of my projects, and work with me on my writing. I set the angle, the sources, and the rules before anything gets drafted, we write together, and I do the final read.

What I’m Learning from My AI Experiments is where I share what broke. Each post covers one experiment: the problem, what I wanted, what I did, what happened, how I fixed it, whether that worked, and what a marketing team should take from it. The numbers come from my own test logs. They’re small samples from one person’s setup, so treat them as something to test in your own environment.

This first post covers the lesson the rest of the series keeps coming back to: when I told the agent to follow a rule, it broke the rule now and then, and I couldn’t predict when. When a script checked the agent’s work before I saw it, most of those mistakes stopped.

The problem: my agent kept breaking rules I’d written down

My first AI guardrails were written instructions to the agent: no filler words in drafts, never invent a quote, avoid one sentence pattern I’d banned. The agent read those rules every session and broke them anyway, some more often than others.

The problem was that I couldn’t predict which session would produce the next mistake. So I checked every draft for the same mistakes by hand, time I needed for the argument itself.

What I wanted was easy to state: rules the agent follows every time, so I could read each draft to keep the writing moving instead of combing it for the same mistakes.

What I did: moved the rules out of the instructions

In mid-August I took 3 rules that kept failing and rebuilt each one as a script. A script here is a short program that reads the agent’s output and acts on it before I see it. One deletes banned filler words. One compares every quoted phrase against the source text and rejects the draft if a quote doesn’t match. A third flags a sentence pattern I’d banned.

The difference is where the rule is checked. An instruction is part of what the model reads, so the model decides whether to follow it. A script runs outside the model and applies the same way every time. Anthropic describes the hooks in its Claude Code tool in those terms: they give “deterministic control: certain actions always happen rather than relying on the LLM to choose to run them” (2. Anthropic, 2026).

What happened

Filler words were the clearest case, because a draft full of them reads like a machine wrote it. With the rule in the instructions, filler got through in 6 of 6 tests. With the script, it got through in none, and I stopped looking for it.

Invented quotes mattered more, because a made-up quote attached to a recommendation makes a wrong answer look trustworthy. With the instruction alone, 3 of 5 quoted passages were still wrong. With the script, none were. The banned sentence pattern showed up less often but kept appearing, because my script matched only some of the ways the pattern can be written. A later version that matched more of those ways cut how often it appeared. Testing each version, then fixing what the test found, is how I brought it down.

A different rule taught me the same lesson. Through August, my agent kept recommending things backed by quotes that appeared nowhere in what it had read. I logged that failure 5 times in August, and each time my fix was more instruction text. On September 15 I replaced the instruction with a script that reads every reply before I do. If a reply recommends something, it has to quote a file or page the agent opened in that session, word for word, or the reply is blocked.

That script still catches mistakes. While I was planning this series, it blocked a reply that quoted a project file the agent had seen but hadn’t opened in that session.

Others have found the same thing. Elizabeth Fuentes L, writing on the AWS developer community on DEV, built a travel-booking agent whose instructions said payment had to be verified before a booking was confirmed. The agent confirmed a booking anyway. A check added outside the model blocked the invalid requests in the test without changing a word of the prompt (3. Fuentes L, 2026).

Good or bad: mostly good, with one real limit

The good part is plain. For filler words and invented quotes, the failures stopped, and so did my hand-checking for them.

The limit is that a script enforces only what you can describe precisely enough for a machine to check. Filler words are a list. A quote either matches the source or it doesn’t. Whether a paragraph sounds like me, or whether a recommendation is any good, can’t be reduced to a check like that. Those calls take a writer’s judgment, and no script replaces that, however hard you try.

Scripts also cost upkeep. Each one is code someone has to maintain, and a script with a mistake in it blocks good work as reliably as bad work. My most expensive mistake, and the subject of the next post, was a checking script that rejected writing I already knew was good.

What I learned, and the question to ask your vendors

If a rule matters, enforce it in code that runs outside the model. Keep instructions for preferences, where an occasional miss costs little.

Marketing teams face the same problem across their whole stack. Martech researcher Frans Riemersma describes the hard part of putting AI agents to work in martech as “aligning probabilistic outputs with deterministic systems without breaking control, compliance or consistency” (4. Riemersma, 2026). In plain terms: AI gives a likely answer, your systems of record follow fixed rules, and something has to check the AI’s answer against those rules.

So when a vendor demos an agent and talks about its AI guardrails, ask one question: which of these rules does your software enforce, and which are instructions to the model? Then ask to see what happens when the agent tries to break one. If the answer for anything touching customer data, spending, or messages to customers is “the model is instructed not to,” you’ve found a rule you’ll need to enforce yourself .

About the Author

Gene De Libero, Independent Martech Advisor, Digital Mindshare LLC

Gene De Libero has spent more than thirty years in marketing technology — as buyer, seller, builder, and advisor. He is the creator of the Marketing Technology Transformation® framework, sponsor of How Marketing Technology Works®, and Independent Martech Advisor at Digital Mindshare LLC, a New York consultancy serving CMOs whose stacks have stopped paying for themselves. He believes most martech investments fail not because the technology is wrong, but because the organization was never built to use it. He fixes that.

Frequently Asked Questions

What are AI guardrails?

AI guardrails are the rules that limit what an AI system says or does. They come in two kinds. Some are instructions written into the model’s prompt, which the model usually follows. Others are code that runs outside the model and checks or blocks its output, which applies every time. The second kind is the one to rely on.

Why do AI agents ignore instructions?

A model weighs each instruction against everything else in its prompt, and it follows them less reliably as the list grows. When researchers gave 20 leading models up to 500 instructions at once, the best followed only 68% of them at 500 (1. Jaroslawicz et al., 2025). An agent’s instructions are guidance it usually follows, with no guarantee.

What's the difference between a prompt instruction and a code-enforced guardrail?

A prompt instruction asks the model to behave a certain way, and the model decides whether to comply. A code-enforced guardrail is a separate program that inspects what the agent produced or is about to do, then blocks it or fixes it. The model has no way to override a check that runs outside it.

How do I check a vendor's AI agent guardrails?

Ask the vendor to list the rules its agent follows and say, for each one, whether the software enforces it or the model is instructed to follow it. Then ask for a demonstration of the agent trying to break a rule. Any rule about customer data, spending, or customer messages should be enforced in code.

What is the What I'm Learning from My AI Experiments series?

It’s a series on How Marketing Technology Works about the AI agents I work with in my consulting business, and what broke along the way. Each post covers one experiment: the problem, what I wanted, what I did, what happened, how I fixed it, whether it worked, and what a marketing team should take from it.
References
  1. Jaroslawicz, D., Whiting, B., Shah, P., & Maamari, K. (2025). How many instructions can LLMs follow at once? arXiv. https://arxiv.org/abs/2507.11538
  2. Anthropic. (2026). Automate actions with hooks. Claude Code Docs. https://code.claude.com/docs/en/hooks-guide
  3. Fuentes L, E. (2026). AI agent guardrails: Rules that LLMs cannot bypass. DEV Community. https://dev.to/aws/ai-agent-guardrails-rules-that-llms-cannot-bypass-596d
  4. Riemersma, F. (2026). Why AI adoption is high but integration is failing in martech. MarTech. https://martech.org/why-ai-adoption-is-high-but-integration-is-failing-in-martech