Skip to content
All blog
AI Accounting Measurement

Measuring whether AI accounting actually worked

David 7 min read

The general question of whether an automation is helping is covered in is your automation actually working. This is the accounting version — what to measure specifically when the process is AP, reconciliation or the close.

The complication: three of the six measures deteriorate at first, and a business that judges at week three on those will switch off something that was working.

Take a baseline before you start

The most common measurement failure is having nothing to compare against. Before go-live, record:

  • Working days from period end to accounts issued
  • Hours per week on transaction processing, honestly estimated
  • Unreconciled items at period end, count and value
  • Transactions processed per month
  • Errors found after the fact, in the last few periods

Half an hour, and without it every later claim is an impression.

The six measures

1. Days from period end to accounts issued

The clearest single indicator, and the one management already cares about.

Expect: worse or unchanged for the first month or two, then a clear improvement. See what month two of automation looks like.

2. Straight-through rate

The proportion of transactions posting without human review.

Expect: low at first while thresholds are deliberately conservative, rising as evidence accumulates. A rate that is not rising after three months means nobody is reviewing the thresholds — which is a management gap rather than a software one.

3. Exception queue volume and trend

More informative than the straight-through rate, because it is a leading indicator.

Expect: high initially — partly genuine, partly historical backlog — then falling sharply. Flat or rising after a full cycle is the clearest warning sign available, and it means a threshold is wrong or the underlying data has an inconsistency the system cannot resolve.

4. Undetected error rate

From your quarterly sampling of confidently-processed transactions. This is the accuracy measure that matters, and the only way to get it is to look — see sampling automated transactions.

Expect: a small number of errors. Finding none in the first sample usually means the sample was too small or not random.

5. Unreconciled items at period end

Count and value. Tests whether reconciliation is genuinely current or whether items are accumulating quietly.

Expect: falling, and importantly, nothing being written off to make it fall. A count that drops because differences are being cleared to suspense is not an improvement.

6. Time spent on processing

Honestly recorded. The measure most people assume is the whole story and which is actually the least reliable, because time freed tends to be absorbed by other work rather than showing up as a saving.

Expect: an increase for the first month. This is the one that generates "it is making more work", and it is true and temporary.

The number that decides it

Beyond the six, one question tells you whether anything really changed:

What does the finance team now do that it could not do before?

If the answer is "the same work, faster", the benefit is efficiency and it is real but limited. If the answer includes things that were previously impossible — reconciling every supplier statement rather than the largest three, chasing debtors on time every month, looking at what moved before anyone asks — that is where the value actually is.

Automation that produces only speed usually gets absorbed. Automation that produces capability changes the function. The difference is largely whether anyone decided in advance what the freed capacity was for.

What not to measure

Headcount reduction, unless that was genuinely the objective. Most finance teams stay the same size and absorb more, which is a better outcome and does not show up as a saving.

Vendor-reported accuracy. Their figure, their definition, their reference customers.

Satisfaction in the first month. The team is doing more work during transition and their view will be correspondingly negative. Ask at month three.

The review points

Month one: are we still on track? Expect worse numbers. Look at exception queue trend, not level.

Month three: the real assessment. Days to close, exception volume, first sampling results, team view.

Month six: is the straight-through rate still rising, and has anyone reviewed the thresholds?

Annually: do the thresholds still match how the business now operates?

The month-six and annual reviews are the ones nobody schedules, and their absence is why automated finance functions drift.

Common questions

How do you measure whether AI accounting worked?

Take a baseline first, then track days from period end to accounts issued, the proportion of transactions posting without review, exception queue volume and trend, undetected error rate from sampling, unreconciled items at period end, and time spent on processing. Three of these get worse before they improve, so judging at week three will produce the wrong conclusion.

Which measure is the most useful early on?

Exception queue trend rather than level. Volume is expected to be high initially because it contains both genuine exceptions and historical backlog, but it should fall sharply as patterns are learned. A queue that is flat or rising after a full cycle is the clearest warning that a threshold is wrong or the underlying data contains an inconsistency the system cannot resolve.

Why does automation seem to create more work at first?

Because the team is handling a long exception queue containing historical items, correcting proposals so the system learns, and in a parallel run also comparing two sets of results. The increase is real and typically lasts a month or two, which is worth telling people in advance since the natural conclusion otherwise is that it does not work.

What is the strongest sign automation succeeded?

That the finance team is doing something it could not do before — reconciling every supplier statement rather than the largest few, chasing debtors on time every month, examining what moved before anyone asks. Automation that produces only speed tends to have the freed time absorbed by other reactive work, whereas automation that produces new capability changes what the function delivers.


Related: what month two of automation looks like · sampling automated transactions · keeping automated books healthy


See what you could build

Start a free trial and describe what your business needs in plain language — SmartB Studio builds the module for you.

Start free trial
Get started

No credit card · Cancel anytime · Your data stays yours