Skip to content
All blog
AI Security Data

Where does your business data go when your ERP has AI in it?

Masni 7 min read

"If the AI reads our invoices, where do those invoices actually go?"

It is the most sensible question a finance director asks, and it is usually answered with a wave of the hand and the word "enterprise-grade". Here is the real version, including the parts that are genuinely uncertain.

What actually happens when AI processes your data

Most business software does not train its own models. It calls someone else's — a model provider, over an API. So your invoice is sent to a third party's servers, processed, and a result comes back.

That sentence alone alarms people, and it should prompt questions. But note that it is not new. Your data already leaves your building constantly: your accounting software is in someone's cloud, your email is on someone's servers, your backups sit in a data centre you have never seen. The AI call is one more third party in a chain that was already long.

What is genuinely new is that the third party is processing content rather than storing it — reading the invoice, not just holding the file. That is a different kind of exposure and it deserves different questions.

The questions that matter

Not "is it secure?" — everyone says yes. These:

Is my data used to train models?

The single most important question. If your data trains a model, information from it can, in principle, influence outputs to other people. For business records this is usually unacceptable, and it is easy to get a straight answer.

Enterprise API tiers from major providers generally do not train on customer data by default — that is a standard commercial term, because no enterprise would buy otherwise. But "generally" and "by default" are doing work in that sentence. Get it in writing, for your tier. Consumer products often have different terms from business ones, and the difference has caught out plenty of employees pasting company data into a chat window.

Ask your ERP vendor: which provider, which tier, and what does the contract say about training? A vendor who cannot answer immediately has not read their own supply chain.

How long is it retained?

Providers commonly retain API inputs briefly for abuse monitoring, then delete. Periods vary and terms change. What matters is that your vendor knows the number and will tell you.

Where is it processed, geographically?

This is where it gets real for businesses in this region. Many model providers process in the US or EU. If you have data-residency obligations — sector rules, a contract with a customer, or a regulator with opinions — "somewhere in Virginia" may be a problem.

Be precise about what you actually need here. Some businesses have hard residency requirements. Many believe they do and have never checked. Both errors are expensive: one is a compliance failure, the other is ruling out useful tools for no reason.

What is sent — everything, or the minimum?

A well-built system sends the smallest thing that does the job. Reading an invoice? Send the invoice. Not the invoice plus your customer list plus the last two years of transactions "for context".

This is a design question and it distinguishes thoughtful vendors from ones who wired an API in and shipped. Ask what gets sent on a given operation. A good vendor answers concretely; a vague answer means nobody drew the boundary.

Who at the vendor can see it?

Your ERP vendor — us included — can typically see your data. That is true of every SaaS product you already use, and it is worth asking anyway: who internally, under what controls, logged how?

What we do, and where the limits are

For fairness, our own answer.

We use third-party model providers on enterprise terms that do not permit training on customer data. We send the minimum required for an operation rather than bulk context. Your records live in your instance; a model call is transient processing, not a copy handed over permanently.

The honest limits: we are a third party, and so are our providers. If your requirement is that business data never leaves infrastructure you control, no cloud AI product meets that bar — ours included. That is a real requirement for some organisations and they should self-host something, and it will cost more and do less. We would rather say so than pretend.

Geographic processing is the constraint we get asked about most in this region, and the honest position is that it depends on which capability and which provider. If you have a hard residency requirement, ask us before you buy, not after.

The risk nobody talks about

Everything above is the outbound question. There is an inbound one that is arguably more likely to hurt you.

Your staff are already pasting company data into consumer AI tools. The supplier contract into a chatbot to summarise. The customer list into something to clean up formatting. The salary spreadsheet into a tool to make a chart.

That is happening in most companies right now, on consumer terms — which frequently do permit training — with no logging, no controls, and no idea what went where.

The irony is sharp: organisations spend months evaluating a vendor's data handling while an unmonitored channel leaks the same information daily. A sanctioned tool with known terms is a security improvement over an unsanctioned one with unknown terms, and that comparison — not the comparison against a hypothetical zero-AI world — is the one that reflects your actual situation.

If you have not asked what your team is pasting where, do that before your next vendor security review. The answer is usually instructive.

A workable position

For most businesses in this region:

  • Insist on no-training terms in writing. Non-negotiable and easy to get.
  • Ask what is sent per operation. Tests whether the vendor designed or just wired.
  • Check whether you truly have a residency requirement rather than assuming. If you do, raise it before buying. If you do not, do not invent one.
  • Ask who at the vendor can see your data, and what is logged.
  • Deal with the consumer-tool channel, which is probably your live exposure.
  • Match caution to sensitivity. Your invoices are not your crown jewels. Your customer list and pricing might be. Not all data deserves the same paranoia, and treating it all as maximally sensitive means you will run out of energy before you get to what matters.

The goal is not zero risk — you gave that up when you started using email. It is knowing your exposure and choosing it deliberately, rather than discovering it in an incident.

Common questions

Will my business data be used to train AI models?

It should not be, and this is the single most important question to ask. Enterprise API tiers from major providers generally do not train on customer data by default, because no enterprise would buy otherwise — but "generally" and "by default" are doing real work in that sentence, so get it in writing for your tier. Consumer products often carry different terms from business ones. Ask your vendor which provider, which tier, and what the contract says.

Where is my data processed, and does the location matter?

Many model providers process in the US or EU, and whether that matters depends on your own obligations — sector rules, a contract with a customer, or a regulator with opinions. Be precise about what you actually need rather than assuming. Some businesses have hard data-residency requirements; many believe they do and have never checked. Both errors are expensive: one is a compliance failure, the other rules out useful tools for no reason. Raise it before you buy.

What if our data must never leave infrastructure we control?

Then no cloud AI product meets that bar, and we would rather say so than pretend otherwise. Any vendor is a third party, and so are the model providers it calls. A model call is transient processing rather than a permanent copy handed over, but it still leaves your infrastructure. Organisations with that hard requirement should self-host something, and it will cost more and do less.

Is adopting an AI tool riskier than not adopting one?

The comparison that reflects your actual situation is not against a hypothetical zero-AI world. Your staff are very likely already pasting supplier contracts, customer lists and salary spreadsheets into consumer AI tools, on terms that frequently do permit training, with no logging and no controls. A sanctioned tool with known terms is a security improvement over an unsanctioned one with unknown terms. Ask what your team is pasting where before your next vendor review.

This is the kind of work SmartB Studio is built for. Get in touch and we will go through it against your actual processes rather than a generic demo.


Related: AI hallucinations in financial data for the correctness side of trust, and what AI still cannot do in ERP.


See what you could build

Start a free trial and describe what your business needs in plain language — SmartB Studio builds the module for you.

Start free trial
Get started

No credit card · Cancel anytime · Your data stays yours