Indirect tax coding with AI
Tax coding appears to be a clean automation candidate. Rules exist, they are written down, and the same transaction types recur.
In practice it is one of the trickier areas, for a reason that is structural rather than technical: the correct treatment often depends on facts that do not appear on the invoice.
Why the document is not enough
The treatment of a transaction can turn on what the goods or services actually were, what they will be used for, where supply took place, the status of the counterparty, and whether an exemption or relief applies.
An invoice reliably tells you the amount and, usually, a description. It frequently does not tell you the rest — and where it states a treatment, that is the supplier's view, which is not automatically your correct treatment for input purposes.
So a model coding tax from the document alone is inferring from partial information. It will be right most of the time, because most transactions are ordinary, and the cases where it is wrong are the ones that matter.
Why the errors are worse than elsewhere
Two properties make tax coding errors more serious than general coding errors.
They are systematic. A supplier coded incorrectly is coded incorrectly on every invoice, from the moment the pattern was learned. By the time anyone looks, the population is large.
They accumulate in a place with an external counterparty. A miscoded expense is an internal reporting issue you can correct. A misstated tax position involves an authority, and correction has a process attached to it.
Combined with the fact that automated errors are consistent and invisible, this is an area where the sampling discipline matters more than almost anywhere else.
The design that works
Learn the ordinary, mandate the rest.
Recurring, unambiguous transactions from known suppliers with stable treatment: automate. That is the large majority of volume, and a model with your history handles it well.
Everything else should be a rule or a human decision, not a proposal:
- Transactions above a materiality threshold
- Anything involving a foreign counterparty
- New suppliers, until a treatment is confirmed
- Categories where treatment depends on use rather than description
- Anything where an exemption or relief is being claimed
- Anything the system's confidence puts below threshold
The point is that these categories should not be automated even when the system is confident, because confidence reflects resemblance to history rather than correctness, and history may contain the error.
Where the treatment depends on use
The hardest category, and worth calling out specifically.
Where the correct treatment depends on what something will be used for — a purchase that could be for business or partly otherwise, an item whose treatment differs by application — no document-based inference reaches the answer. The information exists in somebody's intent.
Two responses, both better than hoping:
Capture it upstream. If the purchase order or requisition records the purpose, the invoice inherits it. This is generally the better project, and it improves more than the tax coding.
Route it to a person by category. If a category is known to be use-dependent, mandate human classification for it regardless of confidence.
The controls to run
Sample by tax code, not just overall. A general sample of transactions may contain few of the tax treatments you are most exposed on. Sample within the treatments that carry risk.
Reconcile the tax accounts to the returns. Movement in the control accounts should agree to what was submitted. A persistent difference indicates a systematic coding issue, and it is one of the few places a genuine problem surfaces as an arithmetic one.
Review new suppliers' first treatment. The moment a pattern is set is the moment to check it, because everything afterwards inherits it.
Check after any change in what you buy or sell. A new product line or supplier type is exactly when learned patterns stop applying.
What to keep
For each transaction, the treatment applied and the basis for it. Where a position is anything other than routine, the reasoning should be recorded at the time.
Reconstructing why a treatment was applied, two years later, from a system that recorded only the outcome, is the situation to avoid. It is also entirely avoidable — see evidence an auditor will accept.
The honest summary
Tax coding automates well for the ordinary majority and badly for the cases that carry exposure. The correct posture is therefore not "automate tax coding" or "do not", but automate the routine, mandate human treatment for defined categories, and sample the automated portion by treatment rather than at random.
Specific rates, thresholds and reliefs change and vary by jurisdiction. This piece is about the mechanics of automating the coding, not about what any particular treatment should be — that is a question for your tax adviser, and the design above is what makes their advice enforceable in your system.
Common questions
Can AI handle tax coding?
It handles recurring, unambiguous transactions from known suppliers well, which is most of the volume. It handles poorly the cases where the correct treatment depends on facts absent from the invoice — what the item will be used for, the counterparty's status, where supply occurred — because it is inferring from partial information and will be confidently wrong on exactly the transactions that carry exposure.
Why are automated tax coding errors particularly serious?
Because they are systematic and external. A supplier coded incorrectly is coded incorrectly on every subsequent invoice, so by the time anyone notices the affected population is large. And unlike an internal misclassification that can simply be corrected, a misstated tax position involves an authority and a correction process.
Which tax transactions should never be coded automatically?
Anything above a materiality threshold, transactions involving foreign counterparties, new suppliers until a treatment has been confirmed, categories where treatment depends on use rather than description, and any position claiming an exemption or relief. These should be excluded even when the system is confident, since confidence reflects resemblance to past treatment rather than correctness.
How do I check automated tax coding is right?
Sample within tax treatments rather than at random, since a general sample may contain few of the treatments carrying most risk. Reconcile the tax control accounts to what was submitted, because a persistent difference is one of the few places a systematic coding problem surfaces arithmetically. Review the first treatment applied to each new supplier, as everything afterwards inherits it.
Related: e-invoicing and AI accounting · sampling automated transactions · SST readiness for Malaysian businesses
Read next
See what you could build
Start a free trial and describe what your business needs in plain language — SmartB Studio builds the module for you.
Start free trial