Skip to content
All blog
AI Accounting Limits

Where AI in accounting still falls short

Chong 8 min read

We build accounting automation. It is in our interest for you to believe it does everything. It does not, and the limits are structural rather than temporary — a better model next year does not remove most of them.

Here is where it stops.

It does not know why you bought something

An invoice from a hardware supplier for RM 4,200 of pipe fittings. Is that repairs, a capital project, or stock for resale?

The document does not say. The answer lives in why the purchase was made, and that information exists in somebody's head or in a project code that may or may not have been entered. A model can learn that this supplier is usually repairs and propose accordingly, and it will be right most of the time — but "usually" is doing real work in that sentence, and the case where it is wrong is exactly the case that matters at year end.

This does not close. It is not a modelling problem. The information is genuinely absent from the input.

It cannot tell you what nobody recorded

The most expensive things in a set of accounts are frequently the things that were never entered.

The invoice still sitting in someone's email. The verbal agreement to a discount. The stock written off in practice but not in the system. The customer who has quietly stopped ordering.

Automation makes recorded data faster and cleaner. It has nothing to say about the unrecorded, and it can make the gap worse by producing books that look complete. A pile of unprocessed paperwork is at least visibly a pile.

It is confidently wrong in a specific, dangerous way

Ask a language model a question about your figures and you get fluent prose with numbers in it. The fluency is constant whether the numbers are right or invented.

This is unlike every previous generation of accounting software, which failed obviously — it crashed, or refused, or produced something clearly nonsensical. A system that produces a wrong number in a well-argued paragraph passes casual review, which is the only review most numbers get.

The mitigation is architectural rather than a matter of model quality: figures must be retrieved and computed by the system, then presented by the model, never produced by the model's own reasoning. Ask any vendor which of those two their product does. We wrote about the failure mode in AI hallucinations in financial data.

It inherits your inconsistencies

A model trained on your history learns your history — including that two people have coded the same cost to different accounts for three years, that a category has drifted in meaning, that someone's workaround became the convention.

It will then apply that inconsistency faithfully, consistently, and at high speed.

This is why data quality is the real constraint and why the least glamorous part of an implementation — cleaning the chart of accounts and the supplier master — determines more of the outcome than the software choice does.

It does not carry accountability

If an automated entry is wrong, "the system posted it" is not an answer that survives an audit, a tax authority, or a board.

Someone owns the process. Someone set the thresholds. Someone reviewed the exceptions or failed to. Automation redistributes work; it does not redistribute responsibility, and any vendor implying otherwise is selling you a problem. See who is accountable for an automated entry.

It struggles at the edges of periods

Cut-off is where accounting gets genuinely hard, and it is where pattern-matching is least reliable.

Whether a delivery on the 30th belongs in this month, whether an accrual is still required, whether an invoice dated the 2nd relates to work done in the previous period — these depend on facts about events, not patterns in transactions. Systems can flag candidates. They cannot make the call, and a system that quietly makes it anyway is a problem you will discover at year end.

It cannot judge whether the numbers mean anything

Revenue is up eleven per cent. Is that good?

Depends on whether it came from one customer who may not repeat, whether margin held, whether it was bought with discounting that has trained customers to wait, whether it is genuine growth or a timing difference. That analysis needs context about the business, the market and intent — most of which is not in the ledger.

This is the durable core of the work, and it is a larger share of the job than it used to be.

What this means practically

None of the above is an argument against automating. It is an argument for automating with your eyes open:

  • Automate the mechanical, review the judgemental
  • Insist every figure is traceable to source records
  • Fix your data before you scale the automation, not after
  • Keep a named human accountable for each automated process
  • Treat period-end and unusual transactions as human territory by default

The businesses that get the most out of this are not the ones with the highest expectations. They are the ones that know precisely where the boundary sits and staff accordingly.

Common questions

What can't AI do in accounting?

It cannot determine intent that is absent from the source document, account for transactions nobody recorded, exercise judgement on estimates and period cut-off, carry professional accountability, or tell you whether the numbers are good news. It also cannot correct inconsistencies in your existing data — it learns and reproduces them, which is why data cleanup usually determines the outcome more than the software choice does.

Will these limits go away with better AI?

Most will not, because they are not modelling problems. A system cannot infer why a purchase was made when that information appears nowhere in its inputs, and it cannot account for a transaction that was never recorded. Accountability is a legal and professional matter rather than a technical one, so it will not shift regardless of how capable the software becomes.

What is the most dangerous failure mode?

A confident, fluent, wrong answer. Older accounting software failed visibly by crashing or refusing, whereas a language model produces a well-argued paragraph containing an invented figure, which passes the casual review most numbers receive. The defence is requiring that every figure be computed by the system and traceable to source records rather than generated by the model.

Should this stop me automating?

No, but it should shape what you automate. Mechanical, high-volume, checkable work is where automation pays; judgement, period cut-off and unusual transactions should remain human by default. The businesses that do best are those that know exactly where the boundary sits rather than those with the highest expectations.


Related: AI in accounting · data quality is the real constraint · when not to automate an accounting process


See what you could build

Start a free trial and describe what your business needs in plain language — SmartB Studio builds the module for you.

Start free trial
Get started

No credit card · Cancel anytime · Your data stays yours