Franchise Tech
Score It Yourself: Evaluating Franchise AI Features After the Demo
Christian Pillat · July 12, 2026 · 5 min read
Evaluating franchise AI features means separating what a chatbot with your manual already does from what needs data a public account cannot reach. Score each claim on five axes — source, scope, memory, action and frequency — then weigh it against how often a real person would meet it.
Vendors rarely lie. Two capabilities of wildly different value simply get described in the same sentence, and the sentence is true both times.
The two claims that sound identical
"It answers questions from your operating manual" and "it answers questions about your business" are one word apart and a category apart.
The first needs a document and a model. Anyone can build it, several people have, and your own team could stand up a passable version in an afternoon. Worth having; not worth choosing a supplier over.
The second needs what a public account structurally cannot reach: the ledger, the roster, the thread where a decision was made, the record of what happened last time. Assembling those is unglamorous integration work, which is why it stays scarce.
The distinction is hard to hold in a demo room because both look like a chat window and both produce a paragraph in four seconds. The difference sits upstream of the screen, and the screen is what you are shown.
There is a second reader for whatever you buy. IFA's analysis of FDD disclosures found 61.9% of franchisors charging franchisees a technology fee, published as Tech Fees by the Numbers. Whatever scores well on your sheet lands on somebody's P&L as a line they did not choose, and they will ask what it does on an ordinary Wednesday.
A rubric for evaluating franchise AI features
Score every claim zero, one or two on each axis, alone, after the meeting, from your notes.
- Source. Where the answer's context comes from. Zero: general knowledge — fluent, plausible, about businesses like yours. One: documents you supplied. Two: documents plus something that is not a document — a ledger, a conversation, a record of a decision.
- Scope. Whether the answer changes with who asks. Zero: everyone gets the same reply. One: screens differ by permission. Two: the answer differs, because scope governs what the system will say — the distinction drawn in role based access control franchise networks.
- Memory. What it knows about time. Zero: every question starts from nothing. One: it holds a readable history. Two: it notices something changed and says so unasked.
- Action. What it leaves behind. Zero: text. One: it creates a task or notifies a named person. Two: it closes the loop, recording the acknowledgement and whether the thing happened.
- Frequency. How often a real person in your network would meet this. Not a capability score but a usage estimate, and you make it, not the vendor.
The binary version of this test — could somebody with a chatbot and a copy of your manual reproduce it in an afternoon — is a good first cut and too blunt to buy on, because it collapses four axes into one yes. The five-axis version tells you where a product is actually strong, which is the choice you face once two shortlisted vendors both clear the first cut.
Resist adding the scores up. A total flatters a product mediocre at everything and punishes one exceptional at the axis you needed.
The frequency column decides more than the capability column
This is the axis buyers leave out, and leaving it out is how brands end up owning impressive software nobody opens.
A capability scoring two on source, memory and action, met by one person at headquarters once a quarter, is a slide. One scoring a flat set of ones, met by every general manager at Friday close, is a habit. The second changes how a network runs; the first gets mentioned at renewal as something you meant to use.
So write the frequency estimate as a sentence with a person in it. "Every location manager, at close, most days." "Our two field coaches, before each visit." "Whoever runs the advisory council, twice a year." Vagueness is the tell: if you cannot name who meets it and when, nobody has a reason to open it.
Then check that estimate against the demo. A vendor whose product genuinely gets used daily describes the routine without prompting, which is one of the questions worth asking in the demo before you sit down to score anything.
Three claims, scored
"Ask the manual anything and get a cited answer." Source one, scope zero or one, memory zero, action zero, frequency high. Verdict: table stakes, genuinely useful, no reason to pick anybody. Buy it as part of something, never as the reason.
"An AI-written weekly summary of network performance." Source one or two depending on whether it reaches the financials, scope one, memory one, action zero. Frequency: one reader, once a week. Verdict: demos beautifully, and the person receiving it usually already knew.
"It flags a location whose waste line has drifted, names the supplier change explaining it, and puts both in the coach's brief the morning of the visit." Source two — it needs the ledger, the invoices and the visit schedule at once. Memory two, action one or two, scope two. Frequency: a field team, weekly. Verdict: this is what the second category looks like, and it is less exciting to watch than the summary that scored lower.
The pattern is worth naming. Features that score highest need your data already in place, so they demo worst and matter most.
What the rubric cannot see
Three honest limits, because a scoring sheet invites more confidence than it earns.
It cannot score your side of the work. A capability scoring two on source assumes your ledger is mapped and your manual is current. Most of the gap between a good demo and a disappointing quarter lives there, not in the product.
It over-rewards ambition. A vendor describing a two on every axis may be describing a roadmap. Score what you watched, not what you were told, and mark the difference.
It says nothing about the company. Pricing, support, ownership, whether they will be here in three years — none of that is on this sheet, and all of it belongs in what to ask the vendor.
Use the sheet twice. Once alone, from your notes. Then send the five axes to the vendor and ask them to score their own product. A supplier who marks themselves a one somewhere and explains why has told you they know their product's shape. One returning twos across the board has told you something less flattering, and it cost you an email.
Everything scored here ends up in the same place: one more system a franchisee is asked to hold in their franchise technology stack, on top of what is already there. No feature clears that hurdle on cleverness. It gets cleared by being met often enough, by enough people, to become how the work is done.
Scoring starts with better notes, and those start in the meeting: franchise software demo questions.
Get new posts weekly