Short answer: Share of model is the percentage of a fixed set of buyer questions where an AI assistant names your business. You measure it by defining 20 to 30 prompts, running every one in a fresh session on each platform weekly, and logging whether you appeared and who appeared instead.
Key Takeaways
- There is no equivalent of a rank tracker, because the same prompt returns different answers to different people.
- Sampling solves this. One answer is an anecdote; thirty prompts run weekly is a measurement.
- Always use a fresh session. Memory and chat history will quietly corrupt your results.
- Record who appeared when you did not. The competitor list is often more useful than your own score.
- Track sentiment and factual accuracy too. Being mentioned incorrectly is not a win.
Why Rank Tracking Does Not Work Here
Traditional SEO measurement rests on an assumption that quietly disappeared: that a query has one result page you can check. Ask ChatGPT the same question twice and you can get two different answers, with different companies named. Ask from a different account and it changes again.
That variability is not a flaw to be engineered away. It is how these systems work. So the measurement has to change shape: instead of checking a position, you sample a distribution.
This is closer to polling than to rank tracking. You do not ask one person who they will vote for. You ask many, repeatedly, and watch the proportion move.
Building a Prompt Set That Is Worth Measuring
The prompt set is the whole measurement. Get it wrong and your numbers are precise and meaningless.
We use 20 to 30 prompts for a focused category, spread deliberately across four intents. Industry practice runs to 30 to 50 when you are also tracking several competitors, and the trade-off is simply how much manual running you are willing to do each week.
| Intent | What it tests | Example shape |
|---|---|---|
| Discovery | Whether you exist in the category at all | "Who builds custom web applications for small businesses?" |
| Comparison | Whether you survive being ranked against rivals | "What are the best options for n8n automation consultants?" |
| Evaluation | Whether what is said about you is accurate | "Is [your company] any good for Laravel development?" |
| Implementation | Whether you are cited on technical questions | "How should I structure a Vue site so Google can index it?" |
Two rules make the set reliable. Write the prompts in the words a buyer would use, not in your own marketing language. And freeze them. If you reword prompts between runs, the trend line means nothing.
How to Run It Without Corrupting the Results
This is where most informal testing goes wrong, usually in a way that flatters the person running it.
- Use a fresh session for every prompt. Memory, custom instructions, and earlier messages all bias the answer. A logged-in account that has discussed your company before will mention it again.
- Run the same day each week. Consistency matters more than frequency. Weekly is enough for a channel that moves in months.
- Run every prompt on every platform. At minimum ChatGPT, Perplexity, Gemini, and Claude. They behave very differently, and an average across them hides the useful detail.
- Log four things per run: did you appear, were you linked or only named, which competitors appeared, and was anything said about you wrong.
- Keep the raw answers. When a score moves, the explanation is usually in the text, not in the number.
Scoring It
Share of model is the simple part: the proportion of prompts where you were named, per platform. If you appear in 6 of 25 prompts on ChatGPT, that is 24%.
Report it per platform rather than blended. A single combined figure can hide the most actionable fact you have, which is usually that you are doing fine on one engine and are invisible on another.
Three secondary measures are worth the effort:
- Citation rate. How often a mention includes a link to you, rather than just naming you. Links send traffic; mentions build familiarity.
- Competitor share. The same percentage for the three or four rivals who keep appearing. This is the benchmark that tells you whether 24% is good.
- Accuracy. How often the description of you is correct. A confident wrong answer about your pricing is an urgent problem, not a measurement artefact.
A Worked Example: What We Run Here
We run this on ourselves. The prompt set covers the four intents above for the services we sell, and it is run in fresh sessions on ChatGPT, Perplexity, Gemini, and Claude, with the raw answers kept so a change can be explained rather than guessed at.
The honest finding from doing this is that the two most useful columns are not your own score. They are the competitor column and the accuracy column. Early on, the competitor list tells you which companies the models consider the category, which is a direct description of the footprint you need to build. And the accuracy column catches things nothing else catches.
We also separate AI referral traffic in analytics, because share of model and actual visits answer different questions. A rising share with flat referrals usually means you are being named without being linked, which is still worth having but is not the same thing.
Tooling, and When to Bother
Commercial AI visibility platforms automate this, and several run far larger prompt sets than a person reasonably can. They are worth it once you are tracking many competitors or many markets.
Below that, a spreadsheet and a disciplined hour a week beats an unused subscription. The method matters more than the tool, and a manual run has one real advantage: you read the answers, and the text is where the explanations live. We automate the scheduling and logging rather than the reading.
What Good Looks Like, and How Long It Takes
Set the baseline before you change anything, or you will not be able to prove the work did anything. Expect the first measurable movement on low-competition prompts within two to four weeks once retrieval problems are fixed, because that part is mechanical. Meaningful movement in a competitive category takes three to six months, because it depends on the third-party footprint, which cannot be rushed.
Treat any agency that promises a share of model number by a date with suspicion. Nobody controls these outputs. What you can control is whether you are reachable, consistent, and corroborated, which is the subject of getting crawled and cited by AI search.
Our GEO service includes this measurement from the first week, and the audit that establishes your baseline is free. Project pricing for the surrounding build work starts from $500 for a website and $1,500 for a web application on our pricing overview, with final cost depending on scope.
Frequently Asked Questions
Frequently Asked Questions
What is a good share of model score?
There is no universal benchmark, because it depends entirely on how competitive your category is and how broad your prompts are. The number that matters is your score compared with the competitors appearing in the same prompt set, and the direction it moves over months. A score with no competitor comparison cannot be interpreted.
Why does the same prompt give different answers each time?
These systems sample from a distribution rather than looking up a fixed result, so variation between runs is expected behaviour rather than a fault. That is exactly why a single check proves nothing and a repeated, fixed prompt set does. You are measuring a tendency, not reading a position.
Do I need to be logged out when testing?
Use a fresh session with no memory or custom instructions, which usually means a logged-out or temporary chat. A logged-in account that has previously discussed your company is far more likely to mention it again, which produces results that look good and mean nothing. Consistency of conditions matters more than which condition you pick.
How many prompts do I actually need?
Twenty to thirty is enough for a focused category and keeps a weekly manual run practical. Thirty to fifty is common when you are also tracking several competitors across more intents. Below about fifteen the percentage swings too much between runs for the trend to be readable.
Should I track Claude as well as ChatGPT and Perplexity?
Yes, if your buyers use it, and in technical categories many do. The engines source information differently, so a strong score on one tells you little about the others. Tracking all four costs little extra once the prompt set exists, and the differences between them are usually where the useful work is.
