One self-hosted console to run your entire business — commerce, ERP, HRM, CRM & manufacturing

AI Agents Inside the ERP: What They Can Safely Do in 2026 and What They Shouldn't

Agents now take actions inside business systems, not just answer questions. Use a four-level risk ladder, five safe first use cases, a vendor checklist and a 90-day pilot plan to decide what to hand over.

Author

Anichur Rahaman

2 months ago12 min read2 views
AI Agents Inside the ERP: What They Can Safely Do in 2026 and What They Shouldn't

Picture the purchasing lead of a ten-outlet retailer walking in at 8:40 on a Monday. Over the weekend an AI agent, running on the vendor's default settings, read a stock report in which a box counted as 1 piece in one table and 12 in another. It raised 14 purchase orders at twelve times the intended quantity, and the first one was emailed to a supplier at 6:15. (This is an illustrative scenario, not a real customer.)

The same agent, set up differently, would have saved those 14 orders as drafts in a queue. Nothing leaves the building, the buyer spots the quantities in a minute, and the Monday carries on. The difference is not how clever the AI is. It is how much the agent may do with no person in between.

Two years ago, "AI in the ERP" meant a chat box that answered questions about your data. In 2026 software vendors sell agents: AI that does not just talk, but looks things up, makes decisions and takes actions inside your business systems. It can raise a purchase order, send a payment reminder or change a price.

That is useful, and it is also the point where mistakes stop being embarrassing and start costing money. A chatbot that gives a wrong answer wastes a minute. An agent that sends a wrong order to a supplier wastes a pallet of stock.

This article gives you a practical way to decide what an AI agent may do in your ERP today, what it should only draft, and what it should not touch yet. It ends with a vendor checklist and a 30-60-90 day pilot plan you can start next week.

This is part 1 of a three-part series on AI in operations. Part 2 explains how MCP connects AI to business data safely: MCP explained. Part 3 covers approvals, audit trails and human-in-the-loop design: AI guardrails.

From copilot to agent: what actually changed

A copilot answers when you ask. It reads data and replies in words. If it is wrong, you notice before anything happens, because a person reads the answer first.

An agent is given a goal, a set of tools and permission to use them. It decides which tool to call, calls it, looks at the result and continues until the goal is met. The tools are the important part: "create draft order", "update stock", "send email", "issue refund".

This changes where the risk lives. With a copilot the risk is a wrong answer. With an agent the risk is a wrong action, repeated quickly and at scale, often without anyone watching each step.

Early forecasts reflect this. In June 2025, Gartner predicted that over 40% of agentic AI projects will be cancelled by the end of 2027, citing rising costs, unclear business value and inadequate risk controls (Gartner press release). The same release warns about "agent washing": vendors relabelling ordinary chatbots and automation as agents. Gartner estimated that only about 130 of the thousands of vendors claiming agentic features were genuine.

The risk ladder: four levels of agent action

The simplest way to stay safe is to stop asking "can the AI do this?" and ask "what happens if it does this wrongly?" Actions fall onto a ladder of four levels, from harmless to hard to reverse.

Four-step risk ladder for AI agent actions: answer questions, draft documents, take reversible actions, move money or make irreversible changes, with the level of human control for each
The further up the ladder, the more control a human must keep.
LevelWhat the agent doesExample in an ERPWhere it belongs in 2026
1. AnswerReads data, explains, summarises"Which SKUs will run out in 10 days?"Safe to run freely, with read-only access
2. DraftPrepares a document a person reviewsA draft purchase order, a product description, a reminder emailSafe if nothing is sent or posted until a person approves
3. Reversible actionChanges data that can be cleanly undoneTagging orders, reserving stock, rescheduling a taskAllowed for narrow, well-tested cases, with a log and an undo
4. Money or irreversiblePays, refunds, deletes, sends to outsidersPaying a supplier, issuing a refund, deleting recordsA person decides, every time, for now

Two details make the ladder work. First, the level is decided by the action, not the AI. "Send an email" is level 2 if a person clicks send and level 4 if the agent sends it to your customers on its own. Second, an agent may be allowed to climb a level only after it has proven itself on the level below, with numbers to show it.

The ladder turns into a routing rule once you ask three questions in order. The first answer that applies decides where the task lands.

Flowchart: a task first asks whether it can be undone (no: keep with a person), then whether money moves or an outsider sees it (yes: draft for approval), then whether the data is clean and the error rate measured (no: fix the data and stay on drafts; yes: automate with a log and an undo)
Three questions, asked in order, send any task to a person, a draft queue or automation.

Five strong first use cases for a growing business

The best starting points sit on levels 1 and 2. They save real time, a human stays in the loop, and a failure is cheap. Here are five that work well for small and mid-sized businesses.

1. Stock and order questions

"How many of this item are left across all outlets?" "Which orders from yesterday are still unpaid?" Staff lose hours each week hunting through screens and exports for answers like these. A read-only agent that queries the live data removes the hunt. The risk is low because it cannot change anything.

2. Draft purchase orders

The agent looks at sales speed, current stock and supplier lead times, then prepares a draft purchase order per supplier. The buyer reviews quantities, edits and approves. The agent does the arithmetic and the typing; the human keeps the judgement and the money.

3. Invoice-to-PO matching

When a supplier invoice arrives, the agent compares it with the purchase order and the goods receipt, then flags differences: a higher unit price, a quantity not received, a duplicate invoice. It suggests, the accountant decides. This is the classic three-way match, and it is tedious work that suits a machine.

4. Overdue follow-ups

The agent lists overdue customer invoices and drafts a polite reminder in the customer's language, adjusted for how late the payment is and how long you have known the customer. Someone reads and sends. Over time, once drafts are accepted unchanged for weeks, you can let it send the first gentle reminder by itself.

5. Product copy and translations

Descriptions, titles and meta text for hundreds of products are slow to write by hand. An agent drafts them from your product attributes, and an editor skims and publishes. Keep a person on anything with legal weight, such as ingredient lists, sizing claims or warranty terms.

What not to automate yet

Some actions look attractive because they are repetitive. Leave them with people for now, or keep the agent to a draft-only role:

  • Paying suppliers or releasing payroll. Money leaves your account and is hard to bring back.
  • Refunds and credit notes above a small limit. These are a favourite target for manipulation by clever customer messages.
  • Price changes across the catalogue. One wrong rule can undercut every margin overnight.
  • Deleting or merging records: customers, products, accounts.
  • Posting to the general ledger. Accounting entries should have a named human behind them, for audit reasons.
  • Anything that tells customers something binding, such as delivery promises, contract terms or compensation.

The reason is documented. The OWASP project lists excessive agency (an AI system holding more functions, permissions or autonomy than its task needs) as a named risk in its Top 10 for LLM applications, and in December 2025 it published a separate Top 10 for Agentic Applications covering goal hijacking, tool misuse and rogue agents. Customer text, supplier emails and uploaded PDFs can contain instructions aimed at the agent. If the agent can move money, those instructions become dangerous.

Data quality and permissions come first

An agent is only as reliable as the data it reads and the rights it holds. Both are easier to fix before you switch it on.

Clean data

If stock counts are wrong, the draft purchase order will be wrong, confidently and politely. Before a pilot, check the basics:

  • Stock levels match what is on the shelf, within a tolerance you have measured.
  • Every product has one SKU, a supplier, a cost and a lead time.
  • Duplicate customers and suppliers are merged.
  • Units of measure are consistent (a box is not sometimes a unit and sometimes twelve).

Narrow permissions

Give the agent its own account, never a shared admin login. Grant it the minimum: read access to the modules it needs, create-draft rights only where it drafts, and nothing else. If the ERP cannot restrict an agent separately from the user who invoked it, treat that as a serious gap. Ideally the agent acts with the permissions of the person asking, so it can never see or do more than that person can.

Part 3 of this series goes deeper on approvals and audit trails. For now, remember the rule: if a person is not allowed to do it, neither is their agent.

How to judge a vendor's "AI agent" claims

Because of agent washing, a demo that looks impressive proves very little. Ask these questions and expect specific answers, not slides.

  1. Permissions. Can I limit the agent per role, per module and per action? Is it read-only by default?
  2. Audit trail. Is every action recorded with who asked, what the agent saw, what it did and when? Can I export the log?
  3. Explainability. For any action, can I see the reasoning and the data it used, in plain language?
  4. Undo. Which actions can be reversed, and how? What happens to the actions that cannot?
  5. Approval steps. Can I require human approval for chosen actions or amounts, and is it enforced by the system, not just by a prompt instruction?
  6. Data handling. Where does my data go? Is it used to train models? Can the model run where I choose?
  7. Cost control. Can I see usage per user and set limits so a looping agent cannot run up a bill?
  8. Failure behaviour. What does it do when it is unsure: stop and ask, or guess?

If your customers talk to an AI directly, check the rules where you sell. In the European Union, the AI Act's transparency obligations, which apply from 2 August 2026, require that people are told when they are interacting with an AI system rather than a human.

To see what the first two levels look like in a working product, here is a short tour of an AI assistant built into an ERP, answering questions and preparing drafts from live business data.

An AI assistant working on live ERP data: asking, drafting, reviewing.

A 30-60-90 day pilot plan

Do not roll an agent out to the whole company. Run a small, measured loop, and let the numbers decide each step.

Ninety-day pilot loop: days 1 to 30 read-only questions, days 31 to 60 drafts with human approval, days 61 to 90 one reversible action, with a measure and review step in every phase
Each phase repeats the same loop: run, measure, review, decide whether to go up a level.

Days 1 to 30: read-only

Pick one team and one process, for example purchasing. Fix the data issues listed above. Turn on level 1 only: questions and summaries. Measure time saved and, more importantly, how often the answers are wrong. Have staff mark each answer correct or incorrect.

Days 31 to 60: drafts with approval

Add level 2 for one use case, such as draft purchase orders. Every draft goes through a human. Track three numbers: the share of drafts accepted unchanged, the share edited, and the share rejected. Write down why drafts were rejected; those reasons become your rules.

Days 61 to 90: one reversible action

If acceptance is high and errors are rare, allow one level 3 action in a narrow case, for example auto-reserving stock for orders under a set value. Keep the audit log open, test the undo, and set an alarm for anything unusual. At day 90, review honestly. If the numbers are not good, stay where you are. That is a result, not a failure.

A simple scorecard keeps the review honest:

  • Hours saved per week, measured, not estimated.
  • Acceptance rate of drafts, and the main reasons for rejection.
  • Errors that reached a customer or supplier (the target is zero).
  • Cost of the AI usage compared with the time saved.

Where this is heading

Back to that Monday. With the ladder applied, the weekend run leaves 14 drafts in the buyer's queue, not 14 orders at suppliers. The buyer sees that every quantity is twelve times the sales history, rejects the batch in a few minutes and fixes the unit of measure in the item master. The rejection reason joins the rules, and the next weekend's run comes out clean.

Agents will get better, and the ladder will move: tasks that need a human today will be routine in two years, once logs, undo and approval steps have earned trust. The businesses that benefit most will be the ones with clean data, narrow permissions and a habit of measuring. Those same foundations also decide how safely an agent can reach your data in the first place, which is the subject of part 2: the Model Context Protocol.

Key takeaways

  • An agent takes actions, so judge it by what happens when it is wrong, not by how clever it sounds.
  • Use the four-level ladder and three routing questions: can it be undone, does money move or an outsider see it, is the data clean. Start on levels 1 and 2.
  • Good first uses: stock questions, draft purchase orders, invoice matching, overdue reminders and product copy.
  • Do not automate payments, large refunds, catalogue-wide price changes, deletions or ledger postings yet.
  • Fix data quality first and give the agent its own narrow permissions, never more than the asking person has.
  • Test vendor claims on permissions, audit, explainability, undo and approvals, then pilot for 90 days with measured results.

Anichur Rahaman is a software architect and the creator of StoreConsole. He designs commerce and ERP systems for growing businesses, with a focus on event-driven architecture, data integrity and self-hosted operations.

About the Author

Anichur Rahaman

Continue Reading