Guardrails for AI in Operations: Approvals, Audit Trails and Keeping a Human in the Loop
How to control an AI that acts: least privilege, approval matrices by reversibility and amount, maker-checker, idempotency, kill switches, audit records, prompt-injection defences, and what the EU AI Act, ISO 42001 and NIST ask in 2026.
Author
Anichur Rahaman
1 month ago13 min read1 views
Thirty-seven refunds worth $8,400 were approved overnight, and every one of them was inside the limit. The support agent a 12-outlet retailer switched on last month may refund up to $300 at a time, and nothing in its setup capped the day.
A wave of messages claiming damaged parcels arrived after midnight, each one polite and plausible, and the agent approved them one by one. The finance lead finds the total on Tuesday morning. This is an illustrative scenario, though each step in it is an ordinary failure of a limit that checks one action at a time.
The agent did exactly what its prompt allowed, and the prompt was the only fence. A prompt is a request, and requests can be ignored, misread or overridden by a clever message from a stranger. Real safety comes from controls outside the model: permissions, thresholds, approvals, limits and records that work the same way whether the model behaves or not.
This article is a design guide for those controls. It covers who may approve what, how to make actions safe to repeat, what an audit record must contain, how to defend against hostile text, and what the rules in the EU, ISO and NIST ask of a normal business as of early September 2026.
A language model decides what to do by predicting text. It is usually right and sometimes confidently wrong, and nobody can promise which one you will get on a given day. If the only thing between the model and your payment system is a sentence saying "never refund more than $50", you have a hope, not a control.
A control is something the model cannot talk its way past. The refund tool itself refuses amounts above the limit. The approval screen cannot be skipped by the agent. The agent's account has no right to touch payroll at all. These rules are boring, deterministic code, and that is exactly why they work.
Think of it as the same discipline you already use for new staff. A new hire gets limited access, a spending limit and a manager who signs off big decisions. You do not rely on their good intentions alone, and you should not rely on the agent's.
Principle 1: least privilege, one identity per agent
Give each agent its own account, named for its job, such as "purchasing-agent". Never let it borrow a person's login or a shared admin key. Then grant only what the task needs.
Read scope: the modules and fields it needs. A purchasing agent does not need customer phone numbers.
Write scope: create drafts, not post to the ledger. Change tags, not prices.
Time and place: run only when triggered by a person or a schedule, from known systems.
Budget: a cap on spend, on model usage and on the number of actions per hour.
A separate identity also pays off later. When something looks odd, the log says "purchasing-agent did this on behalf of Maria", and you can switch that one identity off without disturbing anyone else.
Principle 2: approval depends on reversibility and amount
Part 1 sorted agent actions into four levels, from answering to irreversible money moves. In practice you need a second axis as well: how big is the action? A refund of $8 and a refund of $8,000 are the same kind of action but not the same kind of risk.
Combine the two into a matrix and write it down. The amounts below are an illustrative example; set yours from what a mistake would cost you, not from what looks tidy.
Illustrative approval matrix: the harder an action is to undo and the larger it is, the more humans must be involved.
Three habits make the matrix work. Decide the cell from the action and the amount in code, not by asking the model how risky it feels. Send the action to the approver as a proposal that nothing has yet happened. And re-check the amount at the moment of approval, because a cart or a draft can change between request and click.
Put the matrix inside the full path of one action and it becomes a flow with four gates. Every exit, including the refusals, leaves a record.
Where one proposed action runs, waits or stops. Refusals and rejections are recorded as carefully as successes.
Principle 3: maker-checker, previews and dry runs
Accountants have used maker-checker for generations: the person who prepares a payment is not the person who releases it. Apply the same split to agents. The agent is always the maker. A human, or for the highest-risk cells two humans, is the checker. An agent should never be the checker of another agent's work on money.
A checker can only check what they can see, so every proposal needs a clear preview:
What exactly will change, shown as before and after.
Why the agent proposes it, in two or three plain sentences, with links to the records it used.
What it will cost, and what cannot be undone.
A "dry run" result where possible: the same action run against a copy or in simulation, showing the outcome without saving it.
Beware of approval fatigue. If a person is asked to approve two hundred items a day, they will click through. Keep approvals for the cells that matter, batch the low-risk ones into a daily sampled review, and track how often approvers edit or reject. A rubber stamp that never rejects is a warning sign, not a success.
Make every action safe to repeat
Agents retry. Networks time out, jobs restart, and a model sometimes calls the same tool twice. If "create purchase order" runs twice, you buy twice.
The fix is an idempotency key: a unique reference attached to each proposed action. The system remembers which keys it has already completed and quietly ignores a repeat. Duplicate supplier invoices, double refunds and repeated stock transfers are all prevented by this one habit.
Add three more brakes beside it:
Rate limits: for example, no more than 20 order edits per hour per agent. A runaway loop then stops on its own, long before it becomes a disaster.
Circuit breakers: if error rates or rejections climb, the agent pauses itself and alerts a person.
A kill switch: one clearly labelled control that suspends an agent, or all agents, immediately. Test it on a calm day. The worst moment to discover it does not work is the moment you need it.
The audit record: what to log for every action
When something goes wrong, you need to answer five questions quickly: who asked, what did the agent see, what did it do, which model made the call, and who approved. If the log cannot answer all five, you cannot investigate, and you cannot prove to a customer, auditor or regulator what happened.
One audit record per action, written before the action runs and completed after.
Three details separate a useful log from a decorative one. Write the record when the action is proposed, not only when it finishes, so failed and refused actions are visible too. Store the model name and version together with the prompt or policy version, because behaviour changes when either changes. And keep records tamper-evident: append-only storage, or at least a rule that agents and their administrators cannot edit history.
Mind privacy as well. A log that stores full customer messages is itself sensitive data. Record references and short excerpts where you can, set a retention period, and restrict who may read it.
Defending against hostile text
Part 2 explained prompt injection: text from an email, a review or a product description that tells the model to do something its owner never asked. The OWASP Top 10 for LLM applications lists it as the first risk. No filter removes it completely, so design as if some injections will succeed.
Treat everything the agent reads from outside, such as emails, web pages, reviews and uploaded files, as untrusted data, never as instructions.
Remove or confirm dangerous tools whenever untrusted content is in the conversation. An agent reading a customer email should not be able to issue refunds in the same step.
Keep the sensitive data and the outbound channel apart. An agent that can read all customers and also send email to anyone can be tricked into mailing your customer list out.
Show the approver the original source text next to the proposal, so a human can spot an instruction that does not belong.
Log and review blocked attempts. A rising count tells you someone is probing.
Test before you change anything
Models, prompts, tools and your own data all change. Any of them can quietly alter how the agent behaves. So build a small evaluation set before you go live: fifty to a few hundred real or realistic cases with the answer you expect, including awkward ones such as duplicate orders, missing data, angry customers and injection attempts.
Run the set whenever you change the model, the prompt or the permissions. Compare results with the previous run, and only ship the change if the numbers hold or improve. This is the same idea as regression testing in software, applied to behaviour.
After launch, watch a handful of numbers weekly: acceptance rate of proposals, rejection reasons, actions blocked by policy, cost per action and errors that reached a customer or supplier. And plan rollback in advance. For each action type, know how to undo it, who does it and how long you have. If the answer is "we cannot undo this", the action belongs in the human-only cell.
What the rules ask of you, as of September 2026
Most operations uses of AI in a growing business, such as stock questions, purchase drafts and invoice matching, are not classed as "high-risk" under the EU AI Act. That does not mean there is nothing to do. Here is the position as of early September 2026.
Framework
Status
What it means for you
EU AI Act, transparency (Article 50)
Applies from 2 August 2026. A short extra period to 2 December 2026 applies only to the marking of AI-generated content by systems already on the market.
Tell people when they are talking to an AI, and label synthetic content where required.
EU AI Act, high-risk systems (Annex III)
The Digital Omnibus on AI, agreed in May 2026 and adopted by Parliament and Council in June 2026, moves the date from 2 August 2026 to 2 December 2027. Products already covered by EU product-safety law move to 2 August 2028.
Applies if AI decides on hiring, staff evaluation, credit or access to essential services. Requires logging, human oversight and records.
ISO/IEC 42001:2023
Published December 2023. A certifiable management-system standard for AI.
Optional. A useful structure for policy, roles, risk review and improvement, and a signal to larger customers.
NIST AI Risk Management Framework 1.0
Published January 2023, with a Generative AI Profile in July 2024. Voluntary.
Four functions to borrow: govern, map, measure, manage.
Two cautions. Dates in a regulation like this have already moved once, and summaries of the final text still differ in wording, so check the official text or take advice before relying on a date. And the postponement is a delay, not a cancellation: the obligations remain, and the habits in this article, which are logging, oversight and documented limits, are exactly what high-risk rules ask for.
Even a low-risk use needs two things: tell people when they are dealing with an AI, and keep the records. For the standards themselves, see the ISO page for ISO/IEC 42001 and the NIST AI Risk Management Framework. Outside the EU, national rules differ and keep changing, so check where you sell.
An AI action policy you can copy
Put your rules on one page that staff, vendors and auditors can read. Here is a template with illustrative values. Adjust every number to your own risk.
Action
Agent may
Limit
Approval
Undo
Answer stock and order questions
Read only
Own role's data
None
Not needed
Draft purchase order
Create draft
20 drafts per day
Buyer approves before sending
Delete draft
Tag or reserve orders
Change
Orders under $200, 50 per hour
Daily sampled review
One-click revert
Draft customer reply
Create draft
No refund or compensation promises
Agent sends nothing alone
Not sent
Refund
Propose only
Up to $1,000 one approver; over $1,000 two
Always human
Reverse in payment provider
Pay supplier, payroll, ledger posting, deletion
Nothing
Not allowed
Human only
Not applicable
Review the policy every quarter, and after every incident. When the evidence supports it, move one row up a level. Never move a row up because a vendor's slide says so.
A rollout checklist
Create a separate identity for each agent with the narrowest rights that work.
Write the approval matrix and the action policy, and get the owner and finance lead to sign them.
Enforce limits and approvals in the system, not in the prompt.
Add idempotency keys, rate limits and a tested kill switch.
Log request, agent, model version, inputs, action, decision, approval, result and undo reference.
Build an evaluation set, including injection attempts, and rerun it on every change.
Pilot with one team, review the log weekly, and widen access only on evidence.
Back to that Tuesday. With these controls, each refund would have been a proposal waiting for a person, a daily cap on refund value would have stopped the run after the first few, and the circuit breaker would have paused the agent when the proposals began to look alike. The finance lead would have found a pause alert and a short queue to review, not an $8,400 problem. The agent is no less capable than before. Its worst night is simply a bounded one.
Key takeaways
Guardrails belong outside the model: permissions, limits and approvals enforced by the system, not requested in a prompt.
Give every agent its own narrow identity, and decide approval by reversibility and amount together.
Use maker-checker with clear previews, keep approvals for risky cells, and watch for approval fatigue.
Make actions idempotent and add rate limits, circuit breakers and a kill switch you have tested.
Log who asked, what the agent saw, what it did, which model version acted and who approved.
Most operations uses are not high-risk under the EU AI Act, whose high-risk date has moved to December 2027, but transparency and good records still apply.
Anichur Rahaman is a software architect and the creator of StoreConsole. He designs commerce and ERP systems for growing businesses, with a focus on event-driven architecture, data integrity and self-hosted operations.