Control over AI
Blog
AI data leakage 8 min read

Is AI safe for company data?

Your board wants a yes or a no. The honest answer is that this is four questions wearing one coat, and three of them were settled before anyone started typing.

Leadership team weighing whether company data belongs in an AI tool
Quick answer

There is no general answer, because four separate questions are hiding inside this one. Which account is being used (a free account usually has no processing agreement and trains on input by default). What the vendor contract says, and where the vendor actually sits. What goes into the prompt, which is the only one of the four that can still change at the moment it matters. And who can reach it afterwards: admins, integrations, and whoever receives the output. Three of the four are settled by procurement and IT. The fourth is answered a hundred times a day by someone on a deadline.

01

"Is AI safe" is four questions with four different answers and four different owners

02

Account type decides more than tool choice does

03

A contract governs the vendor, not what your people paste in

04

Retention promises hold until a court says otherwise

05

For special category data and privileged information, the answer stays no

Two people at the same company summarise the same quote for a customer. One is logged into a personal ChatGPT account, the other into the company tenant. Same document, same prompt, same customer name and margin in the text. For one of them this is a reportable incident and for the other it is Tuesday, and the difference has nothing to do with what either of them did.

That is what makes "is AI safe for company data" a bad question to put to a board. It sounds like one question and it is four. They have different answers, different owners, and three of them were settled long before anyone opened a chat window.

One: which account is this?

The largest variable, and it is almost never about the tool. Nearly every major AI vendor runs two products under one brand: a consumer version and a business version.

On the business tier, the vendor does not train on your input by default, a data processing agreement exists, and retention becomes configurable. OpenAI states that split explicitly for Team, Enterprise and the API. On the consumer tier none of that applies, and the training toggle is usually on out of the gate.

Same person, same prompt, same tool. The only difference is which button they logged in with. If you fix one thing this quarter, fix this one, because it is the only part of the four that you can enforce centrally.

More in free vs business AI accounts.

Two: what does the contract say, and where does the vendor sit?

A processing agreement governs what the vendor may do with your data. Necessary, and not the same thing as safe. Three clauses worth reading properly:

  • Retention. How long do conversations persist, including after deletion? Thirty days post-deletion is the common default, and it is a promise rather than a property. In the New York Times litigation, OpenAI was ordered to preserve output logs that users had deleted, including temporary chats. Retention policies hold until a court decides otherwise.
  • Human review. Almost every provider reserves the right to have staff review conversations on suspicion of abuse. That is separate from training, and it survives the training opt-out.
  • Transfers. Is the vendor in the EU, or only its data centre? A US company with a Frankfurt region is still a US company. See data sovereignty starts in the prompt field.

Three: what goes into the prompt?

Here is the gap. The first two questions get answered by procurement, legal and IT: once, deliberately, usually well. The third gets answered dozens of times a day by someone finishing something before a meeting.

It is also the only one still open at the moment it matters. The account is fixed, the contract is fixed, but whether the customer name and the margin travel with that quote is a live decision right up until the send.

The useful question there is not "am I allowed to" but "does the model need this". Shortening a quote works without the customer name. Rewriting a complaint response works without the account number. Reviewing a cover letter works without the date of birth. Ask that question and half the sensitive data usually falls out on its own, with no loss to the task.

Four: who can reach it afterwards?

The forgotten one. Your data does not only sit with the vendor. It sits in a conversation that persists.

On a business tier, an administrator can in principle reach it through the compliance and export tooling that comes with the plan. That is usually the reason you bought the plan, and it is still worth knowing before anyone types something personal into it. See can an admin read your AI chats.

Then there are integrations. Assistants that sit in your mail, your documents, your project tooling. Each connection widens what the tool can see, and that widening is rarely reassessed after the day it was switched on.

When the answer is simply no

Some cases do not survive the breakdown, because the answer comes out no regardless.

Special category data under Article 9 (health, ethnicity, beliefs, criminal matters) does not belong in a general-purpose assistant, business tier or not. The same holds for anything covered by legal privilege or a professional duty of confidence: that is not a data protection rule you can settle with a processing agreement, it is an obligation attached to the professional.

And if you cannot answer the four questions above, that is itself the answer. Not because something is certainly going wrong, but because you would not find out if it were.

So the honest summary for a board is that three quarters of this question is procurement and configuration, and the remaining quarter is a moment. That moment belongs to someone who is finishing a task and not thinking about policy, which is why the only intervention that changes anything happens there: showing what is in the text before it leaves, so the decision sits with a person rather than with nobody.

FAQ

Common questions

Is ChatGPT safe for company data?

It depends on the account. Business tiers (Team, Enterprise, API) do not train on your input by default and come with a data processing agreement. Free and personal accounts run under consumer terms, where the training setting is on by default and no processing agreement exists. Same tool, materially different situations.

Does a business plan make us compliant?

No. It resolves the vendor side of the question: training, retention, the processing agreement. It says nothing about lawful basis, data minimisation, or whether the specific data your staff paste in should have been shared at all. Those obligations stay with you as controller.

Is a European AI tool automatically safer?

For the international transfer question it genuinely helps: a European vendor with no US parent is not exposed to the CLOUD Act. For the question of what people paste in, nothing changes. The data still leaves your organisation and reaches a third party.

What is the biggest risk with AI and company data?

Not the model and not the vendor, but the habit of pasting a whole document because that is faster than extracting the relevant paragraph. That is how most data travels that nobody consciously decided to share.

Should we just block AI?

Blocking moves the usage to personal accounts and personal phones, where you have no visibility at all. Research by KPMG and the University of Melbourne found that 57 per cent of workers hide their AI use from their employer. A ban raises that number rather than lowering it.