How the score is calculated
The whole method, including the parts we know are weak. It is our method — not an industry standard, and not a measure of whether an AI is any good.
Last updated 18 September 2026
What the number is
An AI visibility score answers one question: when somebody asks an AI assistant a question a buyer would ask, does it name you — and how prominently? It is measured, not estimated. We ask real questions of real models through their APIs and read the answers.
It is not a measure of AI quality, not a ranking of the models, and not a claim about how much traffic you get. A high score means you are the brand those models reach for. What that is worth in revenue depends on your market, and we do not pretend to know it from outside your business.
The formula
Every score is a weighted sum of four things, on a 0–100 scale:
score = 100 × (0.45·mention + 0.25·rank + 0.2·citation + 0.1·sentiment)| Component | Weight | What it measures |
|---|---|---|
| Mention | 45% | The share of questions where you appear at all, in any of the answers to that question. A question counts once, however many of its answers name you. |
| Rank | 25% | How high you sit inside the recommendation lists, averaged over the questions where you appear. First position scores 1.0 and each further position loses 20%. Named but never put in a list scores 0.35 — better than nothing, worse than being recommended. |
| Citation | 20% | The share of questions where an answer links to your domain as a source. This is the only component that requires the model to actually browse. |
| Sentiment | 10% | Whether the surrounding text is positive, neutral or negative about you: 1.0, 0.6 and 0.1 respectively, averaged over the answers that mention you. |
A question counts towards mention, rank and sentiment only if the model actually answered it. An engine call that fails is recorded as not measured, never as not mentioned — otherwise a bad network day would look like losing visibility.
Reading the band
| Score | Band | In plain words |
|---|---|---|
| 0–14 | Invisible | These models do not name you. |
| 15–39 | Emerging | They occasionally name you. |
| 40–69 | Competitive | They regularly recommend you. |
| 70–100 | Dominant | You are the default answer. |
What counts as a mention
The answer text is lower-cased and normalised, then each of your names and aliases is matched with word boundaries — so a brand called Core is not counted inside the word score. Multi-word names tolerate punctuation and spacing between their words, so Acme Analytics also matches acme-analytics.
Chinese names are matched literally, because there are no spaces to anchor a word boundary to. That means a short Chinese brand name can match inside a longer unrelated word; if that happens to you, tell us and we will add the surrounding form as a distinct term to exclude.
How the questions are sampled
Each scan runs a fixed set of buyer questions, on every engine your plan covers. Each question is asked more than once, and every answer is stored, because the same question does not get the same answer twice. That variance is the main thing most visibility tools hide, and it is why a single draw of a single question is not evidence of anything.
Two names in the report mean different things: questions is the size of your question set, and samples per question is how many times each one was asked.
How much to trust one number
Every score is reported with a 95% confidence interval, computed on the question level with the Wilson method. If a report says 33% with an interval of 10–70% over 6 questions, the honest reading is "somewhere between one in ten and seven in ten, and we cannot narrow it down yet".
It is worth being precise about where that width comes from, because the obvious fix does not work. The interval is driven by how many questions were asked, not by how often each was repeated. Twenty questions asked once give a much tighter interval than one question asked twenty times — repeating a question samples the same topic, so the answers agree with each other for reasons that have nothing to do with the truth. In our own testing, 60 draws of a single question produced a 24-point interval, while a 20-question set produced 40 points with one draw each: more questions, fewer draws.
The practical rule: if the interval is too wide to act on, you need a larger question set, not a bigger sample count.
Comparing two scans
Two scans are only comparable if they rest on the same basis: the same engines, the same samples per question, and the same question set. Change any of those and the difference you see is mostly the change you made.
So the dashboard does not compare a scan with whatever came before it. It walks back to the most recent scan with an identical basis and compares against that, and it says so when it cannot. A comparison that spans a question-set edit is not a trend, it is an artefact — we found this in our own history, where an alert claimed a rival had overtaken us when the real cause was that we had started measuring on a second engine.
Known weaknesses
These are real, they are in the current version, and they are listed here rather than discovered by you later.
- On an engine that does not browse, citation is unreachable. Citation is worth 20% of the score, and a model that answers from memory has no sources to cite. On a scan whose engines return no sources, the arithmetic ceiling is 80, not 100. It does not change the ranking between brands in the same scan — everyone loses the same 20% — but a score measured with a browsing engine and a score measured without one are not the same scale, and we do not present them as if they were.
- Sentiment is a word list. A fixed set of positive and negative words is scanned in a window around each mention. It does not understand negation, so "not expensive" reads as a negative mention. It is 10% of the score, its failures are visible, and you should treat it as a hint rather than a verdict.
- Rank is per answer, not per model. Where a question is asked several times, the rank used is your best position across those answers. Best-position scoring is generous by design: it answers "can this model put you first?" rather than "does it usually?". The mention component, which is the largest weight, uses the same question-counted grace: a question counts as a mention if any of its answers named you.
- Engine scope varies by plan. The cheapest plans measure on one engine, and the report always names which engines ran, so two scans on different plans are not compared against each other.
This is our method, not a standard
There is no accredited way to measure brand visibility in AI answers. Anyone claiming one is describing their own method with more confidence than the field supports. What we offer instead is a method you can read in full above, applied consistently, with the raw answers stored so you can check any number against the text that produced it.
If you think a weight is wrong, that is a fair argument to make, and it is the reason the formula is on this page instead of in a footnote. The intended use of this version is to compare you against your competitors under one fixed method over time — which is what actually tells you whether your work is paying off.
You can see the method applied to our own site, unflatteringly, in the category we monitor.