How accurate is AI in accounting?
Every vendor quotes an accuracy figure. Ninety-eight per cent, ninety-nine point five, sometimes higher.
The figures are usually true and mostly useless, because accuracy in accounting automation is a family of different measurements and the one being quoted is rarely the one that affects you.
The four different things "accuracy" means
Extraction accuracy. The share of fields read correctly from a document. Highest number, easiest to achieve, least consequential — most extraction errors are visually obvious.
Classification accuracy. The share of transactions coded to the correct account. Lower, and much more consequential, because a coding error is not obvious and flows into your reporting.
Match accuracy. The share of payments correctly matched to invoices. Consequential because a wrong match creates two errors — one invoice wrongly settled, another wrongly outstanding.
Straight-through rate. The share of transactions completing with no human involvement. Not an accuracy measure at all, but frequently quoted as though it were.
A vendor saying "99.2% accurate" is usually quoting the first. Your experience is determined by the second and third.
The question that matters more
Here is the framing that changes the conversation:
Of the transactions the system got wrong, how many did it flag as uncertain?
Consider two systems:
- System A: 99% accurate. The 1% it gets wrong, it is confident about. Those errors post silently and are discovered later, if ever.
- System B: 96% accurate. Of the 4% it gets wrong, it flags 3.5% as uncertain. Half a per cent posts silently.
System A has the better accuracy figure. System B lets through a quarter as many undetected errors, which is the number that determines what actually reaches your accounts.
Undetected error rate is the metric. Accuracy is a proxy for it, and a poor one. This is why confidence scores are worth more than a percentage point of headline accuracy.
Why we design to 98% rather than 100%
Our own design target for reconciliation is around 98% of lines matched automatically. That is a deliberate ceiling, not a limitation we hope to overcome.
The last two per cent is not a technical problem. It is genuine ambiguity — a payment that could plausibly settle two different combinations of invoices, a deduction with no stated reason, a receipt from a party nobody recognises. Those need a decision by someone who can find out what happened.
A system pushed to match 100% is not resolving that ambiguity. It is guessing, or absorbing the residue into a balancing figure. Both produce a reconciliation that always agrees and therefore tells you nothing.
A vendor promising 100% is describing a system that has stopped raising exceptions. That is a worse product, marketed as a better one.
Separately, full traceability means every line is accounted for whether or not it matched automatically. The 98% is about how much runs without a person; traceability is about nothing being unexplained.
What accuracy depends on
Largely your data, not the vendor:
- Consistency of history. A model learning from inconsistent coding cannot exceed the consistency it was taught. See data quality is the real constraint.
- Clean master data. Duplicate suppliers cap matching accuracy structurally.
- Volume of similar transactions. Recurring patterns are learned well; genuinely varied one-offs are not.
- How much context lives outside the system. If the correct treatment depends on why a purchase was made, no accuracy figure will reach it.
Two businesses running identical software will get materially different results. That is why a vendor's benchmark figure tells you about their reference customers rather than about you.
How to measure it yourself
The only figure worth having is your own, and it takes one exercise:
- Take a sample of transactions the system processed confidently — 30 to 50 is enough
- Check each properly against source documents
- Count how many are wrong
- Repeat quarterly
That gives you the undetected error rate, which is the thing that matters. It also gives you an answer to your auditor's question about how you know the system is working, which is otherwise difficult to evidence. See sampling automated transactions.
The realistic expectation
Well-implemented, on clean data, with a year or more of consistent history: the large majority of routine transactions processed without a person, a short exception queue containing genuinely ambiguous cases, and an undetected error rate low enough that quarterly sampling finds few or none.
Poorly implemented, on messy data: a permanent long exception queue, coding that has to be corrected downstream, and a team that concludes automation does not work — when what did not work was the preparation.
The software is rarely the variable.
Common questions
How accurate is AI in accounting?
It depends on which accuracy is being measured. Document extraction is highest and least consequential, transaction coding is lower and matters more, and payment matching sits between them. More useful than any of these is the undetected error rate — of the transactions the system gets wrong, how many it failed to flag as uncertain — because a system that is slightly less accurate but honest about its uncertainty lets fewer errors reach your accounts.
Why do vendors target 98% rather than 100%?
Because the residue is genuine ambiguity rather than a technical shortfall. A payment that could plausibly settle two different combinations of invoices, or a deduction with no stated reason, needs a person who can find out what happened. A system matching 100% is either guessing or absorbing the difference into a balancing figure, which produces a reconciliation that always agrees and therefore cannot signal that anything is wrong.
What determines how accurate the system will be for my business?
Mostly your data rather than the software: how consistently your history has been coded, whether your supplier and customer masters contain duplicates, how much of your transaction volume is recurring rather than one-off, and how much of the correct treatment depends on context that exists nowhere in your records. Two businesses running identical software commonly get materially different results.
How do I measure accuracy myself?
Take a sample of thirty to fifty transactions the system processed with high confidence, check them properly against source documents, and count the errors. That measures the undetected error rate, which is what actually reaches your accounts, and repeating it quarterly also produces the evidence an auditor will ask for when they want to know how you monitor automated processing.
Related: why confidence scores matter in finance · what an AI accounting error looks like · what 98 percent automated reconciliation means
Read next
See what you could build
Start a free trial and describe what your business needs in plain language — SmartB Studio builds the module for you.
Start free trial