Skip to content
All blog
AI Accounting Controls

Keeping a human in the loop — and making it mean something

Chong 7 min read

Every vendor says it. Every implementation plan includes it. And in a large share of finance functions it means a person clicking approve on a queue they are not reading.

A human in the loop is only a control if four things are true. Worth testing yours against them.

One: which loop, exactly?

"A human reviews it" is not a design. The useful version names the specific steps.

Does a person see every transaction, or only exceptions? Before posting or after? Before payment or after? Does the review block the transaction or run alongside it?

Those are different controls with different strengths. A review that happens after posting catches errors before they reach the accounts but not before they reach the ledger. A review that blocks payment is stronger than one that blocks posting, because payment is the irreversible step.

The test: can you write down, for each automated process, exactly which step a human occupies and what they are attesting to?

Two: does the human have what they need?

A reviewer presented with a proposed entry and nothing else cannot review it. They can only agree with it.

Meaningful review requires the source document, the reason the system proposed this treatment, what happened last time with this supplier, and what the system was unsure about. If those take four clicks to assemble, they will not be assembled — under time pressure the reviewer approves on plausibility.

The test: how long does it take to properly review one exception? If it is more than a minute for a routine case, the design is pushing people toward rubber-stamping.

Three: does the human have a real alternative?

A reviewer with an approve button and no practical way to say "I do not know" will approve.

Working designs offer more than binary: approve, reject, hold pending information, escalate, ask the originator. If the only paths are approve or reject, and rejecting creates a problem someone will complain about, the outcome is predictable.

The test: what does a reviewer do when they are unsure? If the answer is "approve it and check later", the control is not functioning.

Four: is the human accountable?

If nobody is named, nobody is responsible, and the review becomes a step rather than a control.

This means a named person per process — not "finance" — and a record showing who reviewed what and when. It also means the reviewer must be able to decline without career consequences, which is a cultural matter rather than a system setting and is the reason some approval controls fail in ways no software can address.

The test: for the last exception approved yesterday, can you name who approved it and what they saw?

The volume problem

The most common way a human loop degrades is not a design flaw. It is arithmetic.

A reviewer given eight exceptions a day reviews them. Given three hundred a week, they cannot — there is not enough attention available, so they skim. Nothing changed in the process; the volume changed, and the control quietly stopped working while continuing to look identical.

This is why exception volume is a control metric rather than an efficiency metric. A rising queue means the control is degrading, and the response is fixing the upstream cause rather than adding reviewer hours. See the exception queue as a control.

Where the human genuinely must be

Some steps should never be fully automated regardless of confidence:

  • Period cut-off decisions — judgement about events, not patterns
  • Estimates and provisions — requiring a view
  • First transactions with a new counterparty — no pattern exists yet
  • Anything above a materiality threshold — the saving is not worth the tail risk
  • Anything irreversible — payments in particular
  • Anything where being wrong damages a relationship — automated dunning on a disputed invoice

These are worth writing down explicitly, because someone will eventually ask why the system cannot handle them, and the answer should exist before the question.

The honest version

A human in the loop is expensive. It is the main constraint on how much automation pays back, and there is real pressure to reduce it.

Reducing it is legitimate — but it should be a decision made on evidence, with sampling behind it, not a drift where the loop technically exists and nobody is really looking. The second is the worst outcome available: the cost of the control without the benefit, plus the false confidence.

If you are going to have a human in the loop, make it real. If it is not real, be honest that you are running unsupervised and manage the risk elsewhere — through sampling, tighter limits, and faster detection.

This is also why a fast, capable assistant still needs a person in the loop for the part it was never built to do — see AI finds the pattern, someone still explains the business.

Common questions

What does "human in the loop" actually require?

Four things: a defined step where the person sits and a clear statement of what they are attesting to, enough context presented to them to make review possible rather than merely agreement, real alternatives to approval such as hold, escalate or query, and a named accountable individual with a record of who reviewed what. Without all four it is a step in a workflow rather than a control.

Why do approval steps stop working?

Usually volume rather than design. A reviewer handling eight exceptions a day reviews them properly, while the same person facing three hundred a week can only skim — nothing in the process changed, but the control degraded while continuing to look identical. This is why exception volume should be monitored as a control metric rather than an efficiency one.

Which accounting steps should always involve a person?

Period cut-off decisions, estimates and provisions, first transactions with a new counterparty, anything above a materiality threshold, anything irreversible such as payments, and anything where being wrong would damage a relationship. These are worth documenting explicitly, since someone will eventually propose automating them and the reasoning should already exist.

Is it ever acceptable to remove the human review?

Yes, for high-volume low-judgement transactions with an independent check, provided the decision is based on evidence — sampling showing the error rate in that category is acceptable — rather than on drift. The dangerous position is a review that technically exists while nobody is genuinely looking, because it carries the cost of the control, none of the benefit, and false confidence on top.


Related: the exception queue as a control · when not to automate an accounting process · segregation of duties when software posts


See what you could build

Start a free trial and describe what your business needs in plain language — SmartB Studio builds the module for you.

Start free trial
Get started

No credit card · Cancel anytime · Your data stays yours