Control over AI
Blog
AI data leakage 8 min read

What "we don't train on your data" actually means

The promise covers training. It does not cover storage, human review, logging, or retention. What happens to text that is not used for training.

DPO reading an AI vendor's terms
Quick answer

"We don't train on your data" means one thing: your input is not used as material to improve the model. It says nothing about four other things that do happen. Your text is stored, in conversation history and in technical logs. Vendor staff may access it under defined conditions, typically on suspicion of abuse. A retention period applies that you rarely control. And that period can be overridden by a court. The training commitment is real. It is simply much narrower than it sounds, and it is the wrong clause to build an assessment on.

01

The promise is about model training only, not storage or access

02

Human review on suspicion of abuse survives the training opt-out

03

Retention is vendor policy, not a property of the system

04

A preservation order can override retention, and one already has

05

Ask about retention, access, logs and jurisdiction, not just training

A vendor assessment lands on a DPO's desk. Somewhere in the security questionnaire is the line "we do not use your business data to train our models", the box gets ticked, and the assessment moves on.

The line is true. It is also the answer to one question out of five.

What the promise actually covers

"We don't train on your data" means your input is not used as learning material to improve the model. On business tiers this is now the default at the major providers, and it is a real commitment with contractual weight behind it.

It describes one path your text can take: the route into a training pipeline. Your text has four other paths, and the promise is silent on all of them.

One: storage

The conversation is retained. It has to be, otherwise you could not see your own history and the service could not hold context across turns. So even with training off, your text sits in a database somewhere, tied to your account.

None of this is hidden and all of it is in the terms. But it is a different answer from "your text is nowhere". For a DPIA the distinction is the whole point: storage is processing, and processing needs a basis, a purpose and a retention period.

Two: human review

Nearly every provider reserves the right to have conversations reviewed by staff. Usually on suspicion of abuse, in response to a safety signal, or for quality assurance.

That right is entirely separate from the training setting. You can opt out of training and still fall under this clause. In practice it applies to a small fraction of conversations, but "rarely" is not "never", and a risk assessment is precisely the document where that distinction belongs.

Three: technical logs

Alongside the conversation itself, services log usage: which account, what time, how much text, which errors. For API traffic that often includes request content, unless you hold a zero-retention arrangement.

Those logs sit outside the conversation history you can see, and they frequently carry their own retention period. Deleting a conversation in the interface does not remove the corresponding log entry.

Four: retention, and who controls it

The common arrangement is that deleted conversations persist for roughly thirty days before permanent removal. Business tiers sometimes make this configurable.

The detail that matters: this is policy, not a property of the system. In the New York Times litigation, OpenAI was ordered to preserve output logs users had deleted, including temporary chats and API requests. Enterprise customers and API customers with zero-retention agreements were outside the order. Everyone else was inside it.

That is the sharpest illustration of this article's point. The training promise was not broken. Those conversations still were not used for training. And text that users believed was gone was sitting in preservation.

Why the distinction matters under the GDPR

The regulation does not care whether an activity is called training. What matters is whether personal data is processed, by whom, for what, and for how long. Each of the four points above is a processing activity.

For your Article 28 agreement that means the training clause is one article. Retention, vendor staff access, sub-processors and transfers are four more, and collectively they carry more weight.

There is a second reason to be careful with the word. Whether a model trained on personal data itself contains personal data is not a settled question. In their opinion on AI models the European data protection authorities concluded that such a model cannot simply be treated as anonymous and that the assessment has to be made case by case. So even on the far side of the training question, the answer is less tidy than the marketing line suggests.

What to ask instead

Four questions that together produce a picture:

  1. How long do you retain my input, and can I configure that?
  2. Who inside your organisation can access my conversations, and under what conditions?
  3. What happens to technical logs, and do the same periods apply?
  4. Which jurisdiction do you fall under, and what is your process when you receive legal demands?

A vendor with clear answers to all four makes the training question almost incidental. A vendor who can only answer the training question has told you exactly one thing.

Which leaves the part no contract reaches. The setting governs what the vendor may do with your text. It has nothing to say about what is in that text, and that is the only variable still open at the moment someone is about to hit send.

FAQ

Common questions

Does "no training" mean my text is not stored?

No. Training and storage are separate. Almost every AI service retains conversations regardless of the training setting, if only to show you your history and to investigate abuse. Turning training off does not change that.

Can vendor staff read my conversations?

Under defined conditions, yes. Nearly every provider reserves the right to have conversations reviewed on suspicion of abuse or in response to a safety signal. That right is independent of the training setting and does not disappear when you opt out.

How long are conversations kept?

Deleted conversations typically remain in vendor systems for around thirty days before permanent removal. Business tiers sometimes let you configure this. Consumer tiers almost never do.

Can a retention period be overridden?

Yes. In the New York Times litigation against OpenAI, the company was ordered to preserve output logs users had deleted, including temporary chats. Retention is a vendor policy operating inside a legal system, not a guarantee.

What should I ask an AI vendor instead?

Four things alongside training: how long you retain my input, who inside your organisation can access it and under what conditions, what happens to technical logs, and which jurisdiction you fall under. Those four determine the actual risk. Training is one input among them.