Skip to content
All blog
AI ERP Agents

Agentic ERP — what happens when the software starts acting on its own

David 8 min read

There is a real difference between software that answers and software that acts, and the word "agent" is doing a lot of unearned work in the market right now.

An assistant waits. You ask, it answers, you decide. Every consequence passes through a human.

An agent has a goal and takes steps toward it without being asked each time. "Chase overdue invoices" becomes: check what is overdue, decide who to contact, write the message, send it, note the reply, escalate what needs escalating, do it again tomorrow.

That second thing is genuinely different, and most of the caution people apply to AI was calibrated for the first.

Why it is a bigger jump than it sounds

Every control in your business quietly assumes a human initiates action. The approval sits between a person deciding and the thing happening. Remove the person from the initiating end and the control is still there — it is just no longer where the decision is being made.

Three assumptions break:

"Someone chose to do this." With an agent, no one chose this action. Someone chose a goal, once, possibly months ago. The link between intent and action stretches until it is hard to see.

"Errors are one-off." A person who makes a mistake makes one mistake. An agent with a flawed rule makes the same mistake four hundred times before lunch, consistently and without hesitation. Automation does not just move faster — it moves faster in whatever direction it is pointed, including the wrong one.

"Someone will notice." Noticing depends on friction. The value of an agent is removing friction. You have to deliberately reintroduce the noticing, because it will not survive on its own.

Where agents genuinely work

Not hypothetical — these are working now, and the pattern is consistent.

Chasing. Overdue invoices, missing documents, unapproved requests. Nobody wants to do it, everyone forgets, and the cost of a mistake is embarrassment rather than money. Almost the perfect agent task.

Monitoring and escalating. Watching for the conditions a human would spot if they had time to look — stock below reorder, margin drift on a product, an integration silently failing. The agent watches; a human decides. Low risk, high value.

Chasing down the exception. Not deciding an exception — gathering what a human needs to decide it. Pull the PO, the delivery note, the email thread, the supplier's history, and present them together. That is often twenty minutes of a person's time, reduced to nothing, with the judgement untouched.

Routine reconciliation with a report at the end. Match what matches, list what does not, hand the list to a person.

Notice the pattern: agents are safest when they gather, watch and prepare, and a human still decides. The value is in the legwork, which is most of the time anyway.

Where they are dangerous

Anything that moves money out. Payment runs, refunds, credit notes. The blast radius is real and it is not recoverable by an apology.

Anything that talks to customers unsupervised. An agent that emails your best client something confidently wrong has caused damage no rollback fixes. Reputational actions are not transactional — you cannot un-send.

Anything with a compounding effect. An agent adjusting reorder points based on demand it partly caused by adjusting reorder points is a feedback loop. These fail slowly, then all at once, and they are hard to spot because each individual step looks reasonable.

Anything where being wrong is invisible. If a mistake surfaces at month-end, your control loop is weeks long and the agent has done it hundreds of times.

The guardrails agents actually need

Different from assistant-era controls. Assistants need review of the answer. Agents need constraints on the action space.

A budget, literally. An agent should have limits it cannot exceed regardless of what it concludes. Ringgit per day. Emails per hour. Records touched per run. Not because you expect it to go wrong, but because the cost of a runaway is bounded by whatever you set here, and unbounded if you set nothing.

A blast radius. What can it touch? An agent chasing invoices needs to read the ledger and send emails. It does not need to edit the ledger. Give it the narrowest permissions that let it do the job — the ordinary principle of least privilege, which the industry keeps rediscovering after each incident.

A kill switch someone knows about. Obvious, routinely missing. And "someone knows about" is the load-bearing part — a stop button nobody can find at 5pm on a Friday is decoration.

A log a human reads. Not a log that exists. One somebody actually looks at, on a schedule. Agents fail quietly; the whole design intent is that you do not have to watch. That is exactly why you must deliberately watch, and why it must be somebody's named job.

Reversibility by default. Prefer actions that can be undone. Draft rather than send. Flag rather than write off. Propose rather than pay. Where an action is irreversible, that is precisely where the human belongs — and the more capable the agent, the more tempting it is to skip that, which is the trap.

The honest state of it

Agents are real and useful for a narrower band of work than the marketing implies.

They are excellent at the loop of watch, gather, prepare, remind. That loop is a large share of what junior operational roles actually consist of, and automating it is genuinely valuable.

They are not ready — and may not be soon — to be trusted with unsupervised, irreversible, money-moving decisions in a small business that cannot absorb a bad week. Not because the technology cannot do it. Because the downside is asymmetric: the upside of an agent paying invoices unsupervised is a few saved hours; the downside is paying a fraudster four hundred times. Those are not comparable, and no accuracy rate makes them comparable.

The correct posture is not "agents when they are good enough". It is agents wherever the action is reversible, and humans wherever it is not — a rule that does not change as the models improve, because it was never about model quality. It is about which mistakes you can undo.

That is the same argument as AI hallucinations in financial data, one level up: the guardrail is structural, not statistical.

Common questions

What is the difference between an AI assistant and an AI agent?

An assistant waits to be asked. You ask, it answers, you decide, and every consequence passes through a human. An agent has a goal and takes steps toward it without being asked each time: check what is overdue, decide who to contact, write the message, send it, note the reply, escalate what needs escalating, then do it again tomorrow. That second thing is genuinely different, and most of the caution people apply to AI was calibrated for the first.

Should I let an agent pay suppliers without a human approving?

No. Anything that moves money out — payment runs, refunds, credit notes — has a blast radius that an apology does not fix. The downside is asymmetric: the upside of an agent paying invoices unsupervised is a few saved hours, and the downside is paying a fraudster over and over. No accuracy rate makes those comparable, which is why the rule is agents wherever the action is reversible and humans wherever it is not.

What happens when an agent gets something wrong?

It gets it wrong repeatedly. A person who makes a mistake makes one mistake; an agent with a flawed rule makes the same mistake hundreds of times before lunch, consistently and without hesitation. It may also go unnoticed, because noticing depends on friction and removing friction is the whole point of the agent. If the error only surfaces at month-end, your control loop is weeks long, so the watching has to be deliberately reinstated as somebody's named job.

Which task should I give an agent first?

Pick one reversible, annoying task. Chasing is the obvious candidate — high tedium, low blast radius, immediately measurable — along with monitoring and escalating, and gathering the context a human needs to decide an exception. Give it a budget it cannot exceed, the narrowest permissions that let it do the job, a kill switch someone can actually find, a log and an owner. Run it for a month and read the log properly before extending it.

What to do about it now

Start with one agent on one reversible, annoying task. Chasing is the obvious candidate — high tedium, low blast radius, immediately measurable.

Give it a budget, a narrow permission set, a log, and an owner. Run it for a month. Read the log properly, not once.

Then decide whether you trust it with the next thing. That is the entire method, and it is deliberately unexciting. The businesses that get hurt here will be the ones who went from zero agents to an agent with the company card, because a demo was impressive.

This is the kind of work SmartB Studio is built for. Get in touch and we will go through it against your actual processes rather than a generic demo.


Related: what AI still cannot do in ERP and the new shadow IT.


See what you could build

Start a free trial and describe what your business needs in plain language — SmartB Studio builds the module for you.

Start free trial
Get started

No credit card · Cancel anytime · Your data stays yours