Does Your AI Vendor Train on Your Data?

“Does this vendor train on my data” sounds like a yes or no question. It is not. The honest answer for most major AI vendors is “it depends which account you are using”, and that is the finding that matters.

The tier is the answer

At the large providers, commercial and enterprise tiers contractually prohibit training on customer data. Consumer tiers frequently permit it, by default, at the same company under the same brand.

So the question your security review answered is not the question your organisation is exposed to. Your review covered the enterprise agreement. Your exposure includes every employee who signed up for a free account with a work email address, and that account runs under consumer terms.

Same interface. Same model. Opposite default. Nothing on screen tells the user which side of the line they are on.

This is why shadow AI and data governance are the same problem rather than two adjacent ones. An unsanctioned free account is not merely untracked software. It is software operating under the permissive version of the terms you negotiated away.

Defaults move

The second thing people get wrong is treating this as a fact to establish once.

Both OpenAI and Anthropic have changed consumer training defaults within roughly the last eighteen months. Not obscure policy footnotes. The default behaviour of the products your staff use most.

A verification from last year is a historical record. If your process is an annual questionnaire, your answer is out of date for most of the year, and nothing in the questionnaire process will tell you when it changed.

What to actually check

Ignore the marketing page. It will say something reassuring that is true of one tier.

  1. Find the terms that govern each tier separately. Consumer terms, business terms, enterprise agreement, API terms. Four documents, frequently four different answers.
  2. Look for the opt-out and who holds it. On consumer tiers, training is often on by default with a per-user setting buried in preferences. A control that lives with the individual user is not a control you have.
  3. Separate training from human review. Distinct commitments. A vendor can truthfully say it does not train on your data while staff review conversations for abuse or quality, which is a different disclosure with different implications.
  4. Check retention alongside it. “We do not train on it” and “we do not keep it” are unrelated statements. API retention typically runs 7 to 30 days, and zero data retention usually requires approval rather than a toggle.
  5. Ask what governs a work email on a free account. The question that produces the useful answer, and the one nobody asks.

The practical fix

You will not solve this with policy language alone, because the people creating the exposure are not reading the policy.

What works is removing the reason to use the consumer tier. Provision enterprise accounts for anyone who wants one, make getting one faster than signing up personally, and say plainly which account to use for work. Most people are not attached to their free account. They are attached to not waiting three weeks.

Where you can, use domain capture or SSO enforcement so that work email addresses cannot create standalone consumer accounts in the first place. That converts a policy problem into a configuration problem, which is a much better class of problem.

Then keep watching

Whatever you verify today has a shelf life. Consumer defaults change, enterprise terms get revised, and new features ship with their own data handling that the original agreement never contemplated.

Tracking those changes is what CopperFeed is for. The full diligence checklist is in AI Vendor Due Diligence, and the chain of companies behind your vendor is in Subprocessors.

Vendor defaults described here reflect published comparisons at the time of writing and change without notice. Verify against current terms for the specific tier you use. General guidance, not legal advice.