Category: Vendor Due Diligence

Training defaults, retention, subprocessors and DPAs. What to verify before signing, and what changes after you do.

  • Subprocessors: The Vendors Behind Your Vendor

    A subprocessor is a company your vendor uses to deliver its service, which means it is a company handling your data because of a decision you did not make.

    You approved one vendor. That vendor approved five more. Those five have their own. The list is longer than it used to be, and at AI vendors it is growing faster than anywhere else in software.

    Why the chain got longer

    Legacy SaaS could plausibly run on its own infrastructure with a cloud provider underneath. An AI feature typically cannot.

    A single assistant inside a product you already license routinely involves the model provider, the cloud that model runs on, a content moderation service, an analytics provider and security tooling. Five parties for one feature, most of them invisible from your side, none of them chosen by you.

    None of that is improper. It is how the stack is assembled, and a vendor building all five in-house would be worse at four of them. The problem is not the existence of the chain. It is that your diligence stops at the first link while your data does not.

    The notice window is the real issue

    Well-drafted DPAs commit the vendor to publishing its subprocessor list, giving notice before adding to it, and providing a mechanism for you to object.

    The mechanism has been hollowed out by the calendar. At most AI vendors the realistic notice window has compressed to somewhere between 14 and 30 days.

    Consider what that asks of you. You receive notification that a new company will process your data. You have two to four weeks to assess an organisation you have never evaluated, form a view, and object, with objection typically meaning you may terminate. In practice nobody does this, and everybody knows nobody does this. The clause is satisfied and the diligence is theatre.

    Negotiate the window if you have leverage. Thirty days is a floor worth holding, sixty is worth asking for, and a vendor that will not commit to publishing the list at all has told you something useful.

    What to actually do about it

    The realistic goal is not evaluating every subprocessor. You will not, and pretending otherwise produces a process that gets abandoned. The goal is knowing which links matter and noticing when they change.

    1. Get the list, dated. Save a copy rather than bookmarking the page. Vendors update these quietly and you want to be able to diff.
    2. Subscribe to the notification list. Most vendors offer one and it is not on by default. Route it to a shared inbox rather than one person, because the one person changes jobs.
    3. Care about two categories. Anyone processing content rather than metadata, and anyone in a jurisdiction that changes your transfer position. The analytics provider counting page views is not your risk. The model provider reading everything your staff type is.
    4. Check the model provider specifically. Frequently the most consequential name on the list, and frequently the one buyers assume is the vendor itself. A product with its own branded assistant may be routing to somebody else’s model entirely.
    5. Re-read after any acquisition. Your vendor being acquired can change its subprocessor chain wholesale, and the notice you get will be about the acquisition rather than about the data.

    The version nobody catches

    Everything above assumes a change you are told about. The harder case is the one that arrives as a product announcement.

    A vendor ships an AI feature into a tool you have used for years. To deliver it they add a model provider. Your contract has not changed, your vendor list has not changed, and the number of companies processing your data has gone up. If the subprocessor page was updated, it was updated quietly, and the release notes described a feature rather than a data flow.

    This is the same failure that runs through everything on this blog. Your records stayed accurate and the product moved. Watching vendor releases and changelogs for the ones that change data handling is what CopperFeed exists to record.

    The broader checklist is in AI Vendor Due Diligence, and the training question in Does Your AI Vendor Train on Your Data.

    Notice windows and subprocessor practices described here reflect published guidance at the time of writing and vary by vendor and contract. General guidance, not legal advice. Your own DPA governs.

  • Does Your AI Vendor Train on Your Data?

    “Does this vendor train on my data” sounds like a yes or no question. It is not. The honest answer for most major AI vendors is “it depends which account you are using”, and that is the finding that matters.

    The tier is the answer

    At the large providers, commercial and enterprise tiers contractually prohibit training on customer data. Consumer tiers frequently permit it, by default, at the same company under the same brand.

    So the question your security review answered is not the question your organisation is exposed to. Your review covered the enterprise agreement. Your exposure includes every employee who signed up for a free account with a work email address, and that account runs under consumer terms.

    Same interface. Same model. Opposite default. Nothing on screen tells the user which side of the line they are on.

    This is why shadow AI and data governance are the same problem rather than two adjacent ones. An unsanctioned free account is not merely untracked software. It is software operating under the permissive version of the terms you negotiated away.

    Defaults move

    The second thing people get wrong is treating this as a fact to establish once.

    Both OpenAI and Anthropic have changed consumer training defaults within roughly the last eighteen months. Not obscure policy footnotes. The default behaviour of the products your staff use most.

    A verification from last year is a historical record. If your process is an annual questionnaire, your answer is out of date for most of the year, and nothing in the questionnaire process will tell you when it changed.

    What to actually check

    Ignore the marketing page. It will say something reassuring that is true of one tier.

    1. Find the terms that govern each tier separately. Consumer terms, business terms, enterprise agreement, API terms. Four documents, frequently four different answers.
    2. Look for the opt-out and who holds it. On consumer tiers, training is often on by default with a per-user setting buried in preferences. A control that lives with the individual user is not a control you have.
    3. Separate training from human review. Distinct commitments. A vendor can truthfully say it does not train on your data while staff review conversations for abuse or quality, which is a different disclosure with different implications.
    4. Check retention alongside it. “We do not train on it” and “we do not keep it” are unrelated statements. API retention typically runs 7 to 30 days, and zero data retention usually requires approval rather than a toggle.
    5. Ask what governs a work email on a free account. The question that produces the useful answer, and the one nobody asks.

    The practical fix

    You will not solve this with policy language alone, because the people creating the exposure are not reading the policy.

    What works is removing the reason to use the consumer tier. Provision enterprise accounts for anyone who wants one, make getting one faster than signing up personally, and say plainly which account to use for work. Most people are not attached to their free account. They are attached to not waiting three weeks.

    Where you can, use domain capture or SSO enforcement so that work email addresses cannot create standalone consumer accounts in the first place. That converts a policy problem into a configuration problem, which is a much better class of problem.

    Then keep watching

    Whatever you verify today has a shelf life. Consumer defaults change, enterprise terms get revised, and new features ship with their own data handling that the original agreement never contemplated.

    Tracking those changes is what CopperFeed is for. The full diligence checklist is in AI Vendor Due Diligence, and the chain of companies behind your vendor is in Subprocessors.

    Vendor defaults described here reflect published comparisons at the time of writing and change without notice. Verify against current terms for the specific tier you use. General guidance, not legal advice.

  • AI Vendor Due Diligence: What to Check Before You Sign

    AI vendor due diligence is the same exercise you already run on any processor, plus four questions that did not exist five years ago and that most security questionnaires still do not ask.

    The general parts are well covered elsewhere. This is the delta.

    The four AI-specific questions

    Any AI vendor processing personal data on your behalf is a processor under GDPR Article 28, so a valid DPA is a legal requirement rather than a nice-to-have. Assume that part is table stakes and go looking for these.

    1. Training defaults, per tier

    The one that catches people. Commercial tiers at the major vendors contractually prohibit training on customer data. Consumer tiers frequently permit it, at the same brand, under the same logo.

    Your diligence covers the enterprise agreement you signed. It says nothing about the free account an employee opened with a work email, which runs under consumer terms and a different default. Two people at your company can paste the same document into what looks like the same product and get opposite outcomes.

    These defaults also move. OpenAI and Anthropic have both changed consumer defaults within roughly the last eighteen months. A policy you verified last year is not a policy you have verified.

    2. Retention

    Ask how long inputs and outputs are held, and separate the API answer from the product answer, because they are usually different.

    API retention commonly runs somewhere between 7 and 30 days. Anthropic has been at 7 days, described in vendor comparisons as the strongest default in the market; OpenAI has been at 30. Zero data retention exists at several vendors but generally requires approval rather than a checkbox, which means it needs to be asked for during procurement rather than discovered afterwards.

    3. Subprocessor depth

    AI vendors stack subprocessors faster than legacy SaaS did, and the chain is longer than most buyers expect. A single AI feature routinely involves the model provider, the cloud it runs on, a content moderation service, an analytics provider and security tooling.

    Each is a place your data goes. This deserves its own treatment, in Subprocessors: The Vendors Behind Your Vendor.

    4. Transfer mechanism, read properly

    Standard Contractual Clauses in the 2021 form, the UK Addendum where relevant, and a transfer impact assessment. None of that is new.

    What is worth your attention is the carve-outs. Vendors advertise EU-only processing and then except support access, abuse monitoring, or a subprocessor that is not regional. An EU-only claim needs reading line by line, because the exceptions are where the transfers actually happen.

    Why the paperwork alone is not enough

    A DPA describes an intended arrangement at a moment in time. It does not tell you what is running today.

    Three things change after signature and none of them require your consent. Subprocessors get added, and the notice window at most AI vendors has compressed to somewhere between 14 and 30 days, which is not long enough for a real assessment. Defaults get revised, as the consumer flips demonstrate. And the vendor ships a new AI feature into a product you already licensed, which changes what the software does with your data without changing a word of your contract.

    That last one is the gap between due diligence as practised and due diligence as needed. Your review was accurate. The product moved.

    A workable process

    1. Ask the tier question explicitly. Not “do you train on customer data” but “which of your tiers permit training, and what governs an employee using a free account with a work email”. The second question gets a much more interesting answer.
    2. Get zero data retention decided during procurement. It is an approval, not a setting, and asking later means asking without leverage.
    3. Require the published subprocessor list and an objection mechanism, and negotiate the notice window upward if you can. Fourteen days is not a diligence period.
    4. Diff the DPA annually. Vendors update these pages quietly. Keep a dated copy of what you agreed to so you can see what moved.
    5. Watch the product, not just the contract. Vendor changelogs are where new data processing shows up first.

    That last step is the one nobody staffs, and it is what CopperFeed records: dated entries for the releases and repricings that change what your software does.

    The training question in detail is in Does Your AI Vendor Train on Your Data. And none of this reaches tools nobody told you about, which is shadow AI.

    Vendor-specific retention and training defaults cited here reflect published comparisons at the time of writing and change without notice. General guidance, not legal advice. Verify current terms and take advice on your own obligations.