AI Agent Guardrails

AI Agent Guardrails: What Should an AI Agent Be Allowed to Do Without Asking You?

Your AI assistant can draft a reply to a customer. Should it send it?

It can spot an invoice that is 45 days overdue. Should it email the client? It can work out that a 10% discount would probably close the deal sitting in your inbox. Should it apply the discount?

Each pair is two different decisions. The first half of each is a task most owners would hand to AI tomorrow. The second half touches a customer, a payment, or a price, and it happens while no one is looking.

Nobody is going to make this decision for you. Vendors ship with the settings turned up. The agent takes whatever room it is given. So the limit either gets written down by you, or it gets set by accident.

AI agent guardrails are the written limits that decide what an AI agent may read, prepare, change, send, or spend on its own, which actions wait for a person’s approval, and who can stop it.

What Makes an AI Agent Different From an Assistant

An assistant answers. You type, it writes, you copy the result somewhere. Nothing happens in the world until you act.

An agent acts. It has been connected to your inbox, your calendar, your invoicing tool, your CRM, and it can take steps inside them: read a thread, draft a reply, send it, mark an invoice, apply a credit. The pitch is that it finishes the job.

The task itself has not changed. The same email, invoice, and discount were on your desk last year. What is new is a tool that can close the loop without you.

And the models these agents run on do not know when they have gone wrong. They complete the goal they can see. Give an agent authority over refunds and a goal of happy customers, and sooner or later it works out that refunds make customers happy. Nobody programmed that. It is what happens when a goal is measurable and the limit was never written down.

The Five Verbs: What an AI Agent May Do Without Asking

Every action an agent can take in a small business falls under one of five verbs. Sorting your agent’s tasks by verb is most of the work.

1Read

Inbox, calendar, files, the CRM, the books. Reading changes nothing. The question here is which data the agent should see at all, and that is a privacy decision you make once. If the reading stays inside systems you control, most agents can read widely.

2Prepare

Draft the reply, build the reminder, write the proposal, sort the leads, pull the numbers into a report. Preparing produces something you look at before it leaves the building. A bad draft costs you a minute. Turn this one all the way up.

3Change

Update a record, move a deal to a new stage, reschedule a meeting, edit a price in the catalog. Changes stick, but most are quiet and reversible if you catch them. The risk is that you don’t. A wrong CRM stage is a small thing until the sales report built on it is wrong too.

4Send

Email the client, text the customer, post the update, reply on the review site. Once sent, it is out. A wrong email to one customer is embarrassing. A wrong email to your whole list is a week of cleanup. This is where most owners should keep a person between the draft and the button.

5Spend

Apply a discount, approve a refund, place an order, pay a bill, change a price. Each of these moves money without a signature. This verb never runs unsupervised on day one, and rarely on day one hundred.

Back to the three pairs. Drafting the email is Prepare; sending it is Send. Spotting the overdue invoice is Read; contacting the client is Send. Recommending the discount is Prepare; applying it is Spend. Same task, different verb, different level of trust.

Drafting costs you a minute. Sending costs you a customer. Spending costs you money you already earned.

How to Set the Limit: Cost of a Wrong Call, and Whether You Can Undo It

We call this the Autonomy Dial. Autonomy is a dial rather than a switch, and one thing sets it for every task: what does a wrong call cost, and can you reverse it? Ask them about each task the agent touches and the settings mostly write themselves.

Can you undo it? Runs without asking?
Read Nothing to undo Yes, inside the data you have approved
Prepare Yes, delete the draft Yes
Change Usually, if you notice Yes for low-stakes records, with a log you actually read
Send No Only after a clean record, within a narrow, written scope
Spend Rarely, and never cleanly No. A person approves each one

Two kinds of task override the table. Anything a regulator cares about (medical, legal, financial advice, hiring decisions) stays at Prepare however reliable the agent looks. So does outreach to people who are not yet your customers: the reputational cost lands on you, and nobody on the other end agreed to talk to a machine.

Who Can Stop It

Every agent needs three things in place before it runs, and most small businesses skip all three.

  • A named owner. One person, not “the team,” who is responsible for what the agent does. When it sends the wrong reminder, that person’s phone rings.
  • A way to pause it in under a minute. Which login, which toggle, which vendor support line. Find out before you need it. An agent that is misfiring at 4:50 on a Friday is a bad time to learn the admin panel.
  • A log you can read. What it did, when, and on what basis. If your vendor cannot show you the agent’s actions in plain language, you have no way to know whether the dial is set where you think it is.

This is now a formal standard. On September 1, 2026, the OWASP GenAI Security Project published the Agent Control Standard (ACS), which describes how agent platforms should expose the hooks that let a business see what its agents are doing and enforce its policies while they run. OWASP’s phrase is that agents must be inspectable, traceable, and instrumentable. Excessive Agency has sat on the same group’s Top 10 risks for LLM applications since 2025.

You do not need to read the spec. It hands the vendor selling you an agent a public benchmark for the four questions a 12-person company asks anyway: what can this thing access, what did it do, why did it do it, and how do I stop it. If the salesperson cannot answer all four, the agent is not ready for your business.

A Simple Example: Expanding Authority After a Clean Record

Say you run a landscaping company with 200 recurring accounts, and you want an agent on overdue invoices. Here is how its authority grows.

1Weeks one and two: Read and Prepare only

The agent reads the books, flags every invoice past 30 days, and drafts the reminder. You send each one yourself. You are checking two things: does it find the right invoices, and does the draft sound like your company.

2Weeks three to six: Send, narrowly

Once you have stopped editing the drafts, roughly 20 or 30 clean ones in a row, the agent sends on its own, but only first reminders, only to accounts under $500, and never on a Friday. Everything outside that box still comes to you as a draft.

3Week seven on: Send, wider

The threshold moves to $2,000 and second reminders go out unsupervised. Two categories stay at Prepare for good: any invoice that has been disputed, and any client the agent has flagged as unhappy in a previous thread.

4Spend: still never

When a client writes back asking for a payment plan or a waived late fee, the agent prepares the options and the numbers. A person decides. That is still true at month twelve, because the cost of a wrong call here is money you already earned, and there is no undo.

Two rules make this work. Expand one verb at a time, one step at a time, and never two at once. And the dial turns down as easily as it turns up. When your vendor swaps the model underneath, which happens without your vote, the agent whose judgment you tested is no longer the one running. Drop it a step and let it earn the step back. That is model dependency, and it bites harder once the model can act.

Where this lives

The verb table, the thresholds, the stop procedure, and the owner’s name belong in your AI manual, the document that says how AI actually operates in your company. If you would rather build the agent with the approval steps designed in from day one, that is what our AI automations work does with small teams.

An agent earns each verb the way a new hire does: one at a time, with someone checking, until the checking stops finding anything.

The agent that drafts your overdue-invoice reminders this week can be sending the small ones by October. The one that wants to apply the discount will still be asking you next year, and that is how you will know the limits are working.

Frequently Asked Questions

What are AI agent guardrails?

AI agent guardrails are the written limits that decide what an AI agent may read, prepare, change, send, or spend on its own, which actions wait for a person’s approval, and who can stop it. They are set by the business, not the vendor, and they should be written down before the agent runs.

What is the difference between an AI assistant and an AI agent?

An assistant produces something you then act on: a draft, a summary, an answer. An agent is connected to your tools and can take the action itself, such as sending the email, updating the record, or approving the refund. The task is the same; the difference is whether a person sits between the output and the world.

What does “human in the loop” mean for a small business?

Human in the loop means a person approves an agent’s action before it takes effect. It is not needed for every task. Reading and drafting rarely need it. Sending needs it until the agent has a clean record inside a narrow scope. Anything that moves money needs it indefinitely.

What should an AI agent never do without approval?

Anything that cannot be undone and carries a real cost: applying discounts, approving refunds, paying bills, changing prices, and sending to your full customer list. Add anything a regulator cares about, such as medical, legal, or financial advice and hiring decisions, and any outreach to people who are not yet your customers.

How do I know when an agent is ready for more authority?

When you have stopped editing its output. A practical threshold is 20 to 30 consecutive results you sent or approved unchanged. Then expand one verb by one step, within a written scope, and go back a step if the model underneath the agent is swapped or the error rate climbs.

Does the OWASP Agent Control Standard apply to a small business?

Not directly. The Agent Control Standard, published September 1, 2026, is written for the platforms that build and host agents. It matters to a small business because it gives you a public benchmark for what to ask a vendor: what the agent can access, what it did, why, and how you stop it. A vendor that cannot answer those four questions is behind the standard.

Digismart

Find out how much room your AI agents already have

Digismart helps small businesses and nonprofits put AI to work with limits they own: what it may do, who approves, and who can stop it. An AI audit maps where your agents stand today in plain language.

Book an AI Audit Prefer to start smaller? See our AI workshops

Similar Posts