Governed AI
The AI Vendor Data Agreement: Four Clauses a No-Training Promise Does Not Cover
Christian Pillat · September 9, 2026 · 5 min read
An AI vendor data agreement needs to answer more than whether your content trains a model. Four clauses do the work: what counts as your content, whether anything derived from it can be used against you, whether the promise flows down to the model provider, and what leaves when you do.
Most procurement conversations get the first question right and stop there. This week gave everyone a reason to keep reading.
What happened, in the shortest honest version
A dispute broke into public view on 8 September over one of the seven Millennium Prize problems. Two mathematicians working toward a proof say information about their specific approach reached a foundational model provider, which then produced a proof of its own along the same unusual line. The provider says it did not see their work "through any means until they released it publicly", that its own effort began on 1 September, and that the two proofs differ significantly. The researchers dispute that account. It is being argued in public, and I have no view on who is right.
What makes it a procurement story rather than a mathematics story is one sentence in the provider's own response. It said it cannot rule out that de-identified data derived from the researchers' usage of its products helped improve its models.
Read that as a franchisor. Not an accusation, not a leak, not a breach — a company saying, accurately and in good faith, that it cannot fully account for what its systems learned from the way somebody used them.
Why a no-training clause did not answer it
The clause every buyer asks for says the vendor will not use your content to train or improve models. It is the right clause, it is genuinely load-bearing, and with a reputable enterprise vendor it holds — which is why the answer to does AI train on your franchise data is still no. Nothing this week changes that.
What the sentence above is about is a different category. Three things routinely sit outside the definition of "your content" in an AI vendor agreement:
- Telemetry and metadata — what was asked, how often, from where, how long the session ran.
- Aggregate patterns — behaviour across many customers, from which no individual customer is identifiable.
- De-identified derivatives — material computed from your content that is no longer your content.
None of those three is sinister. All three are how software gets better, and you have accepted the same trade with every SaaS product you own. The point is narrower: a no-training term written around "customer content" does not necessarily reach them, and until this week most buyers had no concrete reason to ask where the boundary sat.
What an AI vendor data agreement has to say out loud
Four clauses. Ask for them in writing, before the second demo, and read the answers for whether they cite a provision or a value.
- A definition of "your content", with the exclusions listed. Not the term itself — the carve-outs. If telemetry, aggregates and de-identified derivatives are excluded, say so in the agreement, so nobody is surprised by a sentence in a press statement two years from now.
- A use restriction on derivatives, not just a training restriction. Training is one use. The clause you want covers what may be done with anything derived from your account: whether it can inform product development, benchmarks, published research, or a competitive offering.
- Flow-down to the model provider, and notice when their terms change. Your vendor's promise is worth exactly what its own upstream agreement is worth, and upstream terms are revised on a schedule nobody consults you about. Require that the term is passed down, that you are told when it moves, and that a material change is an exit right rather than an email.
- An exit with a date on it. What you can export, in what format, and how long anything retained upstream survives your cancellation. A retention window that can be set to zero is a feature; a retention window nobody can quote is an answer.
A vendor who has done this work replies to all four in one message. A vendor who has not sends a page about how seriously they take your trust, which is the same information in a slower form.
The clause franchisors specifically need
There is a franchise-shaped version of clause two, and it is the one I would not sign without.
Your operating playbook is the asset. Not your customer list and not your recipes — the accumulated method for running a location, which is the thing a competing brand cannot buy and cannot reverse-engineer from the outside. Every question your network asks an AI system is a fragment of that method. In aggregate, the question stream is the playbook, described from the inside.
So the clause reads roughly: nothing learned from this account, in any form, aggregated or otherwise, may be used to build, improve or inform a product or service offered to another franchise system. That is a field-of-use restriction rather than a privacy term, and it is a normal thing to ask for.
It matters more in franchising than in most industries because you are not the only party living under these terms. 61.9% of franchisors charge franchisees a technology fee, on the IFA's analysis of FDD technology-fee data — which means the agreement you sign becomes a set of conditions your franchisees pay for and operate under, without having read it. That is the same accountability the ownership question raises in franchise AI content ownership: you are contracting on behalf of people who will hear about the terms only if something goes wrong.
What this is not a reason to do
It is not a reason to stop. A week of headlines about a model provider will produce a round of franchisors deciding to wait, and waiting has its own bill — the one I costed out in cost of ungoverned AI in franchising. While the brand waits, the network does not: the manual still gets pasted into personal accounts under consumer terms that permit exactly the thing everybody is now worried about.
The lesson runs the other way. The episode is an argument for a sanctioned environment with a real agreement behind it, because the alternative is not an absence of exposure but an unbounded and undocumented version of it. That is the whole case for a private AI environment for franchises, and it got stronger this week rather than weaker.
What did change is the reading list. Four clauses instead of one, and a habit of asking what a term excludes rather than what it promises.
The narrower question, answered properly: does AI train on your franchise data.
Get new posts weekly