Back to all posts

Governed AI

Does AI Train on Your Franchise Data? Training and Retrieval Are Not the Same Thing

Christian Pillat · December 8, 2025 · 5 min read

Does AI train on your franchise data? With a reputable enterprise vendor, no. Training is how a model was built, months before you bought anything. Retrieval is a system reading your documents at the moment of a question and discarding them afterwards. Enterprise agreements normally prohibit the first.

Almost every AI-privacy conversation I have with a franchisor collapses those two things into one, and the collapse is expensive. It produces brands that reject a governed system over a risk they were never exposed to, while a hundred franchisees paste the manual into free accounts nobody vetted.

So the answer to does AI train on your franchise data depends less on the technology than on which account somebody typed the question into.

The two things the word "AI" is doing in that sentence

Think about a district manager you hired last month.

Everything she knows about how businesses work in general — reading a P&L, running a meeting, writing a professional email — she learned before she met you. That is training. It happened elsewhere, over years, and none of it was yours.

Everything she knows about your brand she gets by opening your manual when a question comes up. That is retrieval. She reads the page, answers, and closes the binder.

The two work in opposite directions:

  • Training changes the model permanently. It adjusts the model's internal settings, it affects everyone who ever uses that model, and it cannot be undone for one customer afterwards.
  • Retrieval changes nothing. Your document is pulled up, the relevant passage is attached to the question, an answer is written, and the passage is dropped. Delete the document and the next answer no longer knows it existed.
  • Only one is a decision you control. The training run finished before the product existed. Retrieval is a contract you sign.

When someone asks whether AI is "learning from" their manual, they are almost always picturing training and describing retrieval.

So, does AI train on your franchise data?

Split the question by the account it is asked in.

In a consumer or free tier, conversations are commonly used to improve the provider's models unless a setting is changed, and the setting sits with whoever opened the account rather than with your brand. That is a genuine exposure, and it is the one already happening in most networks.

In an enterprise or business agreement, the terms run the other way: your content is not used to train or improve models, it is processed to answer your questions and nothing else, and the commitment is contractual rather than a preference somebody can toggle. Reputable vendors pass that term down to the model provider underneath.

So the honest framing is not "AI versus no AI." It is a sanctioned account with terms, or a hundred unsanctioned accounts without them — the argument I set out at length in governed AI for franchises.

The franchise-specific problem is capacity to check. Half of US franchise systems operate in fewer than 10 states and only 16% reach 35 or more, on FRANdata's count of system footprints. A brand that size has nobody whose full-time job is reading a data-processing addendum, so the questions go unasked — not from indifference, but because nobody owns the task.

What reputable enterprise vendors actually do

None of it is heroic. It is table stakes, which is why a vendor who has it confirms all of it in one reply, and a vendor who does not sends you a page about how much they value your trust.

  • A contractual no-training term. In the agreement, not the marketing FAQ, and extended to any model provider underneath.
  • Separation between customers. Your content is retrievable by your network and nobody else's — and in franchising that has a second layer, because one franchisee's numbers should not surface in another's answer.
  • A retention window you can see. How long content sits with the model provider before deletion, and whether that can be zero.
  • Named subprocessors. Whoever runs the model is handling your text. A vendor unwilling to name them is asking you to accept a party you cannot evaluate.
  • Bounded human access. Which employees can read customer content, under what trigger, and whether it is logged.

And the honest residual: retrieval still sends your text to a third party for processing — the same category of decision you already made with your payroll provider. Not zero. Bounded, contractual and auditable, which is a different thing from safe.

The questions to ask, and the answers that should worry you

Six questions, in writing, before the second demo.

  1. Is our content used to train or improve any model — yours, or your provider's? A good answer cites a clause. A worrying answer cites a philosophy.
  2. Who are your subprocessors, and which of them sees our text? "We use industry-leading providers" is a refusal wearing a suit.
  3. How long does that provider retain it, and can retention be set to zero?
  4. Can one franchisee's data appear in another's answers? Ask for the mechanism, not the assurance.
  5. Who at your company can read our content, and is that access logged?
  6. If we leave, what is deleted, when, and what proof do we get?

Post the answers where your franchise advisory council can read them. Some brands get asked by someone with more weight than a council: FRANdata puts more than 12.4% of active US franchise brands under some level of private-equity ownership or backing, reported via Franchising.com, and those owners run data questionnaires on a schedule. Everyone else gets asked eventually by a large franchisee's counsel or an insurer.

Why the confusion costs money

The conflation produces two bad outcomes and I see both regularly. The first is a blanket prohibition: AI gets banned because "it trains on our data," the policy circulates, and nothing else changes. Usage moves into personal accounts, which is where training exposure actually lives — so the policy raises the risk it was written to remove and removes the visibility that would have shown you.

The second is treating every vendor as equivalent. If training and retrieval are one word in your head, all AI products are one product, and the decision defaults to whoever demoed best. That is how brands end up with a system that cannot say where an answer came from — the ground condition for the confident wrong answers operators have to live with, and for a franchise network reporting view nobody can trace to a source.

Under the technical question, what franchisors are really asking is whether the operating manual can end up in a competitor's answer. Through a governed retrieval system with an enterprise agreement, essentially no. Through a general manager's own free account, opened on a phone and mentioned to nobody, it may already have — and no vendor questionnaire will find that, because there is no vendor to send it to.


Answers a franchisee can check need something underneath them: a governed AI architecture for a franchise network.

Get new posts weekly

Weekly at most. Unsubscribe any time.

Back to all articles

See this working on your own content

Bring one operations document and the questions it should answer. We will show you the answers and the citations live.

Schedule Demo