Skip to content
All blog
AI Accounting Controls

Sampling automated transactions — the control nobody schedules

David 7 min read

Finance teams running automation review their exceptions diligently. The queue is visible, it demands attention, and clearing it feels like the job.

Meanwhile the ninety-something per cent of transactions that processed confidently are never examined by anyone.

That is the gap. The exception queue contains what the system knew it was unsure about. The errors that matter are the ones it was sure about and wrong — and no amount of exception review will surface them, because by definition they were never flagged.

What sampling finds that nothing else does

Systematic miscoding. A supplier consistently coded to the wrong account. The model is confident because your history taught it this, and it is wrong every time. Nothing flags it — it is the pattern.

Thresholds that stopped fitting. A tolerance set when your average invoice was RM 3,000, still applied now it is RM 12,000. Everything passes, including things that should not.

Rules that outlived their reason. A treatment mandated for a situation that no longer exists, applied faithfully for two years.

Drift after a business change. You added a product line, changed suppliers, restructured departments. The automation reflects the previous shape, confidently.

Each is invisible to exception review and obvious under sampling. All four share a property: they are consistent, which is exactly what anomaly detection cannot catch either.

How to actually do it

It is a small exercise, and its value comes from being repeated rather than being thorough.

Take 30 to 50 transactions that processed with high confidence in the period. Random, not selected — selecting is how you accidentally sample only the ones you already understand.

Check each properly. Not that it balances, but that the treatment is right: open the source document, confirm the account, the period, the amount, the tax treatment. This is the part that takes time, and skipping it makes the whole exercise theatre.

Record what you find. How many checked, how many wrong, what kind of wrong. The record matters as much as the finding.

Fix the cause, not the transaction. A miscoded supplier means correcting the pattern so future invoices are right, not just fixing the four you found.

Repeat quarterly. Monthly if volumes are large or the system is new.

An hour or two per quarter for most small finance teams.

Where to concentrate

Random sampling is the baseline. Two refinements make it more productive:

Stratify by amount. A purely random sample from a population of mostly small transactions will mostly contain small transactions. Sample separately from the largest transactions, because that is where an error costs most.

Sample around changes. After a new supplier group is onboarded, after a threshold is adjusted, after a business change, after a system update. Errors cluster around change, and a sample taken immediately afterwards is worth several taken during a stable period.

Why this is also your audit answer

When an auditor asks how you know your automated processing is working, there are two possible responses.

One is a description of the system, which is a statement of intent.

The other is a file showing that you sampled 40 transactions each quarter, found three errors, identified the causes, and corrected them. That is evidence of a functioning control, and it is considerably more persuasive.

Crucially, it cannot be produced retrospectively. A sample taken during the year, recorded at the time, is evidence. The same exercise done during the audit is not. This is one of the specific things automated finance functions get asked about and frequently cannot answer — see year-end adjustments and the audit file.

Finding nothing is a result

Teams sometimes stop sampling because it keeps coming back clean.

That is the wrong conclusion. Clean samples are evidence the system is working, and they are what lets you justify raising automation levels — lowering a confidence threshold is defensible when you have three quarters of clean sampling behind it, and a guess when you do not.

It is also the case that the value arrives unevenly. Several quarters find nothing, then one finds a supplier that has been miscoded since a rate change. The exercise is cheap insurance, and its return is lumpy by nature.

The minimum version

If a full sampling programme will not happen, do the reduced version:

Once a quarter, take your ten largest automated transactions and check them properly.

Fifteen minutes. It does not measure your error rate, but it catches the errors that cost the most, and it is far better than the common alternative of never looking at anything the system was confident about.

Common questions

Why sample transactions the system processed correctly?

Because you cannot tell in advance which of them were processed correctly. Exception review only covers what the system flagged as uncertain, so errors it was confident about — systematic miscoding, thresholds that no longer fit, rules that outlived their purpose — are invisible to every other control. Those errors are consistent rather than anomalous, which also means anomaly detection will not find them.

How many automated transactions should I sample?

Thirty to fifty per quarter is sufficient for most small finance teams, drawn randomly rather than selected, and checked properly against source documents rather than merely confirmed to balance. Sampling separately from the largest transactions is worth adding, since a random sample of a mostly-small population will mostly contain small items.

When is the best time to sample?

Quarterly as a baseline, and immediately after any change — a new supplier group, an adjusted threshold, a business restructure, a system update. Errors cluster around change, so a sample taken just afterwards is worth several taken during a stable period.

What if the sample keeps coming back clean?

That is a useful result rather than a reason to stop. Clean samples are the evidence that justifies increasing automation levels, since lowering a confidence threshold is defensible with several quarters of clean sampling behind it and a guess without. The return on sampling is lumpy by nature — several quarters find nothing and then one finds a supplier that has been miscoded since a rate change.


Related: when the accountant becomes the reviewer · how accurate is AI in accounting · preparing for an audit with automated books


See what you could build

Start a free trial and describe what your business needs in plain language — SmartB Studio builds the module for you.

Start free trial
Get started

No credit card · Cancel anytime · Your data stays yours