Running an AI accounting pilot properly
A pilot is supposed to reduce risk by testing a decision cheaply. Most accounting pilots do not, because they are designed to succeed rather than to find out.
Run on a hand-picked sample of clean transactions, for two weeks, evaluated on impressions. They demonstrate that the software works on easy cases, which was never in doubt.
Here is what a pilot has to establish, and how to run one that would actually change your mind.
What the pilot must answer
Four questions. If the design cannot answer these, it is a demonstration rather than a pilot.
How accurate is it on our data? Not the vendor's benchmark — yours, with your suppliers, your coding history, your mess.
What does the exception queue look like at steady state? Volume, composition, and how long clearing it takes.
Where does it get things wrong? The pattern of errors matters more than the count, because a systematic error class is a configuration problem and a scatter is a data problem.
How much work is this actually going to be? Including the parts nobody counts — correcting, investigating, handling the client or supplier queries it generates.
The design that works
Use real, messy, unselected data. A full period of everything, not a chosen sample. The entire point is to meet the awkward cases, and selecting the data removes them.
Run one complete cycle. A month, including period end. Two weeks misses cut-off, accruals and the close, which is where the informative disagreements are.
Run in parallel. The existing process continues. This is the expensive part and it is what makes the pilot a test rather than a leap — you have a correct answer to compare against, produced independently.
Log every disagreement. For each difference between the automated and manual treatment, record: what the system did, what the person did, who was right, and why. That log is the actual output of the pilot; everything else is impressions.
Have the eventual users run it. Not the person who chose the software. The people who will handle exceptions daily are the ones whose experience predicts the outcome.
What the disagreement log tells you
Reviewing it at the end, disagreements fall into four categories, and the proportions tell you what to do next.
System wrong, cause identifiable. A threshold, a rule, a supplier not yet learned. Fixable, and expected in a pilot. A high proportion here is fine.
System wrong, cause not identifiable. The concerning category. If you cannot say why it did something, you will not be able to fix it or explain it to an auditor. A few of these warrant a serious conversation with the vendor.
Person wrong. More common than anticipated, and valuable. It means your manual process was inconsistent, which is worth knowing regardless of what you decide about the software.
Both defensible. The treatment was genuinely ambiguous. These point at decisions your business has never made, and making them improves the process either way.
A pilot producing mostly the first and third categories is a good result. One producing the second is a warning.
What to measure
Numbers to have at the end, per process:
- Transactions processed
- Proportion that would have posted without human review
- Disagreements, by the four categories above
- Exception queue volume, and trend across the month
- Average time to clear an exception
- Time spent by the team on the pilot, honestly recorded
That last one is routinely omitted and is the basis of any real business case. A pilot that saved processing time while consuming more in correction and investigation has told you something important.
The three pilot mistakes
Choosing the data. Every awkward case removed is a question unanswered.
Stopping early because it looks good. Week two enthusiasm is not evidence. The difficult cases cluster at period end and the queue composition changes as the system learns.
Evaluating on impressions. "The team liked it" and "it seemed accurate" are not findings. The disagreement log is.
What a good result looks like
Not zero errors. A good pilot result is:
- Accuracy acceptable on your data, with errors you can explain
- Exception volume falling across the month rather than flat
- Disagreements mostly attributable to fixable causes
- A team that found the exception work manageable
- A clear list of what needs configuring differently before go-live
And one more, frequently the most valuable: a list of things your manual process was doing inconsistently, discovered by comparing it against something that does the same thing every time.
When to stop the pilot
Two conditions warrant stopping rather than persisting:
Errors you cannot explain. If the vendor cannot tell you why a treatment was applied, that problem does not improve after go-live — it becomes an audit problem instead.
An exception queue that is flat or rising. It should fall as patterns are learned. If it does not after a full cycle, either the data has an inconsistency the system cannot resolve or a threshold is badly wrong, and going live means going live with an unworkable queue.
Common questions
How long should an AI accounting pilot run?
One complete cycle, meaning a full month including period end. Two-week pilots miss cut-off, accruals and the close, which is precisely where the informative disagreements between automated and manual treatment occur, and the exception queue composition changes over a month as recurring patterns are learned.
What data should a pilot use?
Real, unselected data — a full period of everything, including the awkward cases. Choosing a clean sample removes exactly the transactions the pilot exists to test, which is why so many pilots demonstrate that the software works on easy cases while answering nothing that was actually in question.
How do you evaluate a pilot?
By a log of every disagreement between the automated and manual treatment, recording what each did, which was right and why. Sorting those into system wrong with identifiable cause, system wrong with no identifiable cause, person wrong, and both defensible tells you what to do next. Impressions such as "the team liked it" are not findings.
What is a good pilot outcome?
Errors that are explainable rather than absent, an exception queue falling across the month rather than staying flat, disagreements mostly traceable to fixable causes, a team that found the exception work manageable, and a specific list of what to configure differently. A frequently valuable by-product is a list of things the manual process had been doing inconsistently, revealed by comparison with something that behaves the same way every time.
Related: choosing the first accounting process to automate · starting with AI in accounting · measuring whether AI accounting worked
Read next
See what you could build
Start a free trial and describe what your business needs in plain language — SmartB Studio builds the module for you.
Start free trial