Four different observations wearing one word
Buyers and vendors use these four terms as if they were interchangeable, which is how a weak measurement gets sold as a strong one. They are not synonyms, and OpenAI's own documentation is careful about the distinction between what a search-backed answer shows and what it proves.1
| Term | What it means | What it does not mean |
|---|---|---|
| Mention | The brand name appears somewhere in the answer | That the answer offered it as an option |
| Recommendation | The answer presents the brand as a choice for the question that was asked | That anyone will act on it |
| Citation | A URL is shown with the answer and can be opened | That the page caused the answer |
| Visibility score | A summary of how often brands appeared within a defined question set | Anything about questions outside that set |
A record you can re-check contains nine things
If you cannot reconstruct why a record was counted, you are holding a number with a story attached. Ask to see one complete raw record before you buy anything. It should carry all of the following.
- The question, word for word, including any constraints attached to it.
- The model, with its version or mode.
- Whether web search was available for that specific run.
- Date and time, with a timezone.
- Language and market, because both change answers materially.
- The complete answer text, not a summary or a screenshot.
- Which brands appeared, and how each one was classified.
- Every source URL the provider returned with that answer.
- A way to export or print the record so someone outside the tool can read it.
Twelve questions to ask before you buy
These are the questions that separate a measurement instrument from a dashboard. Ask them in writing and keep the answers, because you will need them again in six months.
- 1. Can I edit the question set myself?
- 2. Is the complete answer text stored, or only a score?
- 3. Are mentions, recommendations, and citations reported as separate observations?
- 4. Are the returned source URLs shown for each answer?
- 5. Are competitor results broken out per question, not only in aggregate?
- 6. Are the model, language, market, and search setting recorded with every run?
- 7. Can I export the underlying records?
- 8. How many repeated runs are included, and what does an extra repeat cost?
- 9. Does the tool report uncertainty, or only a single point estimate?
- 10. How long is history retained, and what happens to it if I cancel?
- 11. Which claims on your own site are independently verified?
- 12. What will you not do? A vendor with no stated limit has not thought about one.
Re-read a quarter of the records by hand
Automated classification drifts, especially on ambiguous sentences where a brand is named approvingly but not offered as a choice. A sampling check catches most of that drift, and it takes less time than people expect.
Two habits make the difference. First, draw the sample deliberately rather than reading whichever records are easiest to reach — at least one per question type, and at least one per competitor. Second, set the pass threshold before you start reading, so the decision is not made by whoever is defending the dashboard.
- Re-read the answer and confirm each brand was classified the way a careful human would classify it.
- Open every source URL and confirm it resolves and supports the sentence it is attached to.
- Confirm the run conditions were recorded at the time, not reconstructed afterwards.
- Read the answer for a caveat: 'good for', 'although', 'cheaper but' — those clauses are where a recommendation turns into something weaker.
The number moves even when your website does not
This is the most common source of a false win. Repeated-sampling research shows that citation distributions are heavily skewed and that rankings across sources are unstable from sample to sample, throughout the frequently cited set rather than only at the top.2 Many apparent differences between two domains sit inside the measurement's own noise floor.
A protocol paper on repeated-query auditing quantifies the cost of skipping repeats: reliability of roughly 0.58 at five iterations, 0.74 at ten, and 0.81 at fifteen.3 The practical reading is that five runs is a quick look and fifteen is a defensible record, and that a move of a few percentage points is usually not a finding.
There is a second moving part that no number of repeats fixes: the model itself changes. Pin the model and mode where you can, and treat a model change as a new baseline rather than a continuation of the old trend line.
Red flags worth acting on
Any one of these is worth a direct question. Three of them together usually means the product cannot tell you what it claims to.
- Only a total score, with no way to see the answers behind it.
- Mentions described as recommendations, or a nearby citation described as proof.
- Run conditions that are not disclosed, or a 'last updated' label with no date attached.
- No history, so every report is a snapshot with nothing to compare against.
- No export, which makes independent checking impossible by design.
- Rankings presented as fixed positions, with no indication of sampling variance.
- Language implying a guaranteed recommendation if you follow the roadmap.
What AI Cite Who records, and what we do not
The same standard has to apply here, so this is field by field. Where something is not offered, it says so rather than leaving it implied.
| Field | Status | Detail |
|---|---|---|
| Question set | Recorded | Generated for review or entered manually; enabled questions drive a run |
| Model and mode | Recorded | One fixed search-capable model, noted per run, with a capped number of search calls |
| Language | Recorded | Audits follow the account locale and the question set is written in it |
| Complete answer text | Recorded | Stored with each audit run and used to derive the classifications |
| Mention versus recommendation | Separated | Classified as distinct observations, with provider-returned source URLs kept alongside |
| Source URLs per answer | Recorded | Stored exactly as the provider returned them; competitor-owned pages are labelled rather than counted as opportunity |
| Competitor comparison | Recorded | Per question, and discovered candidates only count once you accept them |
| History | Plan dependent | Free keeps none, Plus keeps 90 days, and Pro keeps 730 days. Older runs are locked rather than deleted. |
| Raw data export | Pro only | A readable HTML report with the summary, per-question results, cited sources, observable evidence paths and actions. It does not include the full answer text. |
| Uncertainty interval | Not offered | We report what a run observed; we do not publish a confidence interval across repeated runs |
| Platforms beyond the search-capable model | Not offered | We do not measure assistants we do not operate |
Run the same yardstick afterwards
When you do make a change, repeat the identical question set under the identical conditions, then put the two records side by side. If the conditions moved, say so, and treat the difference as inconclusive rather than as a win.
What you end up with is not a promise. It is a defensible account of what was observed, what it was based on, and what is still unknown — which is exactly what a team needs in order to decide what to fix next.