Governed AI
The Economics of AI Usage Pricing: Franchise AI Cost Per Location, Explained
Christian Pillat · June 10, 2026 · 5 min read
Franchise AI cost per location is driven by tokens — the chunks of text a model reads and writes — so a task costs roughly what it costs to take in and produce. Policy lookups are cheap. Document-heavy and comparative work is not. Budget for the shape of consumption rather than an average.
Almost every AI product a franchise brand is shown this year is priced, underneath, on consumption. The vendor may hide that behind a seat price or a bundle, but the bill they pay their model provider moves with usage, and eventually yours does too.
The budget is arriving regardless. Three in four franchisors expect to increase capital spending on technology and innovation, though a narrower 28% mentioned incorporating AI and increased automation in their plans, on the FRANdata and IFA franchisor survey. Few of them can yet say what a unit of that spend buys.
Which makes token economics an operator's subject. You need enough vocabulary to tell a fair pricing model from an unfalsifiable one.
What a token actually is
A token is a fragment of text. Common words are usually one; longer or unusual words break into several; punctuation counts. A page of prose is some hundreds of them.
Four properties explain most of what shows up on a bill:
- Both directions count. The model is billed for what it reads (input) and what it writes (output), usually priced separately, output the dearer of the pair.
- Context is re-read. A long back-and-forth re-sends the conversation so far with each turn, so the tenth message in a thread costs more than the first even if you typed three words.
- Retrieval is input. Grounding an answer in your manual means passing the relevant sections into the model — a real cost with a real return.
- A long answer is a cost, not a courtesy. A drafted policy or a rewritten job description bills as output the whole way.
None of it is exotic, and all of it is invisible in a demo where somebody asks three short questions and everything looks free.
Why one question costs many times another
Task shape, not question difficulty, moves the number. Roughly in ascending order:
A policy lookup. A short question, a retrieved passage or two, a short cited answer. The cheapest thing the system does, and the thing your network does most.
A document summary. A supplier agreement or inspection report in, a paragraph out. Large input, small output — cheaper than people assume. A drafted document is the inverse shape, and dearer for it.
A meeting. Audio is transcribed on its own basis, and the transcript then becomes a large input for the summary and the follow-ups. Two costs, one activity.
A comparison across locations. Many rows of financial data in, several passes over them, a paragraph out. The expensive one, and the one that finds money.
Two effects vendors rarely raise. A model that reasons at length before answering produces tokens you never see and somebody pays for. And a retry — a failed tool call, a re-asked question, an answer regenerated because the first was wrong — bills again.
Franchise AI cost per location: the shape, not the average
Here is where operators go wrong. They ask what a location costs a month, get a number, multiply by their unit count and budget that. Consumption arrives skewed, though, in three ways.
It is bursty. A rollout, a health inspection, a bad month at four sites, a supplier switch — each produces a cluster of asking, in the weeks you most want people asking.
It concentrates in a few users. In most locations one or two people drive the majority of it, usually a manager who found a use nobody wrote down.
A handful of heavy tasks outweigh everything else. Take an illustrative profile rather than any customer's: a location asks 200 questions in a month, four of them full comparisons of its P&L against its volume band. Those four can account for more consumption than the other 196 combined — and they are the four worth paying for.
So the mean location is a fiction. What matters for budgeting is the spread: what the tenth-busiest location does against the median one.
The spread has a governance side. Metering per named individual is itself a personal signal, and a consumption view showing headquarters who asked what, by name, is a monitoring feature with an invoice attached — franchisee monitoring vs privacy arriving where nobody expects it.
How to budget a line that moves every week
The problem is never the amount. It is that this line does not behave like the others, so last year plus a bit produces a number nobody can defend.
Four practical moves.
- Derive the allowance from a pilot's distribution, not a vendor's average. Run six to ten locations for a quarter, then set the per-location allowance near the upper end of what the useful ones consumed.
- Budget the heavy tasks separately. Comparative and document work is lumpy and schedulable. If you know a supplier review is coming, you know where a burst is coming from.
- Decide who pays past the allowance before anybody reaches it. Allowances, ceilings and overage are a control surface, and I have set ours out under franchise AI cost control.
- Watch both variables move. Model prices per token have fallen steadily; the work people hand to a model has risen faster. Budget against usage growth, not last year's rate card.
One connection changes the arithmetic more than any of that. The moment a location's ledger is connected, comparative work becomes possible — and comparative work is the expensive kind, which argues for sequencing it rather than avoiding it: franchise POS accounting integration.
Six questions that expose a pricing model
Ask these of any vendor, in writing.
- What is the billed unit, in your own words? If nobody can define it in a sentence, what you are being sold is a lever with a unit's name on it.
- Are input and output priced the same? A vendor bundling them is absorbing a risk or hiding a markup; either is fine as long as it is the answer.
- Does the history of a conversation bill again on each turn?
- Do retries and failed answers bill?
- Is the rate a pass-through or a markup, and what happens when the model price changes? Falling costs should reach you.
- What happens at the ceiling — stop, degrade, or keep billing?
None of that requires technical knowledge. It requires refusing a bundle whose contents nobody will itemise — which is all spend metering means as a governance term: a number somebody owns.
And be blunt about what the evasions protect. A vendor who will not itemise is rarely guarding a trade secret. They are guarding an assumption about how much your network will use, made before they met you — which is the actual price.
Metering is one control among several, and here is where it sits: naming what a network actually governs.
Get new posts weekly