Skip to content
All blog
AI Accounting Buying

How to evaluate AI accounting software

Chong 8 min read

Every accounting product now lists document capture, automated reconciliation, AI coding and natural-language reporting. The feature grid is useless for choosing, because everyone ticks every box.

What differentiates is how each handles the awkward cases, and that only shows up when you test rather than watch.

The general vendor questions — data ownership, security, commercial terms — are covered in how to evaluate an AI ERP vendor. This is the accounting capability test.

Test 1: give it your ugliest documents

Not a sample invoice. Take ten real documents from your worst suppliers: the photographed receipt, the invoice inside a forwarded email chain, the statement with a running balance, the one in two languages, the handwritten delivery note.

Watch for: what it does with the ones it cannot read. Does it fail clearly, or produce a confident wrong extraction? The second is far worse and you will only see it if you know what the document actually said.

Test 2: the multi-invoice payment

Give it a payment settling several invoices, less a credit note, less a deduction the customer applied without explanation.

Watch for: whether it proposes the combination and shows the arithmetic, or simply reports no match. This single case separates real matching from exact-reference matching more reliably than any other test, and it is the case that consumes most reconciliation time in practice.

Test 3: ask to see the exception queue

Populated with real messy data, not a demo set.

Watch for: grouping by reason rather than date, plain-language explanations rather than error codes, the proposed treatment with its confidence, one-click access to the source document, ageing, and routing to someone other than finance.

An undifferentiated list of unmatched items is a much weaker product than one that classifies them. See the exception queue as a control.

Test 4: the new supplier

An invoice from a supplier the system has never seen, in a category where the treatment is ambiguous.

Watch for: a proposal with a confidence level and the reasoning behind it. If it either refuses or codes confidently with no expressed uncertainty, there is probably no model underneath — see rules versus machine learning in accounting.

Test 5: trace a figure backwards

Pick any number on any report and try to reach the source document.

Watch for: how many clicks, and whether you can then get a plain-language statement of why that transaction was treated that way. Traceability without explainability is common and only half of what you need — explainability in accounting automation covers the distinction.

Test 6: break it deliberately

Try to do something the system should not allow. Post to a restricted account. Exceed an approval limit. Submit a duplicate. Submit an invoice from an unapproved supplier.

Watch for: whether the constraint is enforced or merely configured. A limit that exists in a settings screen and not in behaviour is not a limit, and reading the documentation will never tell you which you have.

The five questions that go with the tests

Short, accounting-specific, and revealing. The general commercial questions are in the article linked above; these are about the work.

"What happens at period cut-off?" The most-tested area in any audit and where automation is least reliable. A vendor who says it is handled automatically has either solved something difficult or not understood the question.

"How do I identify every transaction processed under a particular rule version?" Determines whether fixing a systematic error takes hours or weeks.

"What gets absorbed into a balancing figure?" If anything is written off automatically to make a reconciliation agree, the reconciliation has stopped being a control.

"Are explanations recorded at the time, or generated when I ask?" The difference between evidence and a plausible reconstruction.

"Show me a customer's exception queue." Anonymised. A vendor with confident customers can arrange this.

Scoring it

Weight the tests by what you actually do. A business with heavy marketplace settlements should weight test 2 and the traceability test most. One drowning in supplier paperwork should weight test 1.

What not to weight heavily: breadth of features you will not use, integrations with systems you do not have, and anything demonstrated rather than tested.

The meta-signal

Across all six tests, notice how the vendor responds to a case that does not work.

Every product has cases it handles badly. A vendor who says "that one would go to your exception queue, here is what the reviewer would see" is describing a product built by people who know what accounting is like. One who deflects, or insists the case is unusual, is telling you what support will be like when you meet it in production.

Common questions

How do you evaluate AI accounting software properly?

By testing with your own difficult data rather than watching a demonstration. Give it your worst documents and see whether failures are clear or confidently wrong, give it a payment settling several invoices less a credit note, ask to see a real exception queue, submit an invoice from an unknown supplier and check for an expressed confidence level, trace a report figure back to a source document, and attempt actions the system should block.

What separates a real matching engine from a basic one?

A payment covering multiple invoices, less a credit note and less an unexplained deduction. A basic engine reports no match; a capable one proposes the combination that reconciles and shows the arithmetic. That case consumes most reconciliation time in practice, which is why it discriminates better than any feature list.

What should I ask about period cut-off?

Simply what happens at it. Cut-off is the most tested area in an audit and the place automated processing is least reliable, because the decision depends on facts about events rather than patterns in transactions. A vendor claiming it is fully automatic has either solved something genuinely difficult or has not understood the question, and it is worth finding out which.

What is the most revealing thing in an evaluation?

How the vendor responds to a case their product handles badly. Every product has them, and a vendor who explains that it would route to the exception queue and shows what the reviewer sees is describing something built by people who understand accounting. Deflection, or insisting the case is unusual, indicates what support will be like once you meet it in production.


Related: questions to ask an AI accounting vendor · red flags in an AI accounting demo · how to evaluate an AI ERP vendor


See what you could build

Start a free trial and describe what your business needs in plain language — SmartB Studio builds the module for you.

Start free trial
Get started

No credit card · Cancel anytime · Your data stays yours