What these tools can and cannot see
A recommendation monitoring tool asks a language model a defined set of questions, saves the answers, and turns them into a report. The questions come from you or from a vendor's prompt library. The answers come from a model with web search available. Nothing in that loop touches another person's chat history.
That limit is worth stating plainly, because a dashboard invites the opposite assumption. A visibility percentage is not an eavesdropping device. It is a summary of repeated tests, run under conditions that should be written down and kept with the numbers.
- Measured: how often a brand appears in answers to a defined question set, and which public sources the provider returned alongside those answers.
- Not measured: private conversations, a buyer's intent, or a causal explanation of why one brand appeared instead of another.
- Not measurable by anyone: whether a future answer will include you.
Four jobs hide behind one product category
Teams buy these products for different reasons, and each reason needs different evidence. Decide which job is yours before comparing prices, because the cheap product for one job is usually the wrong product for another.
| The job | What you actually need | What you can skip |
|---|---|---|
| First look | One baseline of how your public pages read to a machine | History, competitor benchmarking, quotas |
| Question-level tracking | A stable question set, repeat runs, and a record of what moved | Multi-brand rollups |
| Competitor and citation diagnosis | Per-question answers with the competitor set and the returned source URLs | Long history |
| Multi-brand or multi-market programmes | Project separation, quotas, per-market question sets | Very little, which is why this costs the most |
Seven dimensions that survive a re-check
Feature lists converge fast, so compare on what you can inspect after the demo ends.
- 1. Question ownership. Can you edit the question set, or are you limited to a vendor library? Your buyer's wording matters more than someone else's keyword database.
- 2. Complete answer retention. Is the full answer text stored, or only a score? A score with no answer underneath cannot be checked, quoted, or defended.
- 3. Separation of mention, recommendation, and citation. A brand name inside a list is not a recommendation, and a URL shown next to an answer is not proof that the page caused that answer.1
- 4. Returned source URLs. If the tool never shows them, you cannot tell a first-party fact gap from a third-party coverage gap.
- 5. Competitor comparison per question. Company-level share of voice hides the pattern that matters: a rival that owns comparison questions and loses discovery questions is a different problem from one that leads everywhere.
- 6. History under recorded conditions. A trend line is only a trend if the model, language, market, and search setting stayed the same across runs.
- 7. Export and auditability. Ask for a raw export and read three rows by hand. That single test resolves most of the ambiguity in a demo.
Where AI Cite Who sits
Free is a website readiness baseline built from public page evidence: access, identity, offering, proof, indexability prerequisites, and machine-readable signals. Each project includes one website analysis per week. It deliberately excludes a recommendation audit and competitor tracking, because neither can be done defensibly from a single crawl.
Plus is $49.90 per month or $499 per year, with 300 monthly credits, 2 projects and 50 saved questions per project; a run checks up to 10. Pro is $99.90 per month or $999 per year, with 900 monthly credits, 5 projects and 100 saved questions per project; a run checks up to 20.
Paid audits sample ChatGPT API web-search answers, separating direct recommendations, neutral mentions, final-answer citations and retrieved candidates. Answer text is retained internally for analysis. Discovered competitors remain candidates until confirmed; comparisons still require checking questions, language and search conditions.
Pro downloads include readable HTML reports and Markdown handoff reports with structured evidence and tasks, not a complete raw-answer export. Downloading an existing report uses no credits. Checks currently cover ChatGPT only; saved-question capacity is not a per-run allowance.
Names that come up in this category
The vendor descriptions below retain their September 16, 2026 verification date. They are not a ranking and we are not affiliated with these vendors. For AI Cite Who, OtterlyAI and Profound, use the linked workflow comparison with official pages checked on September 28, 2026.
- Semrush AI Visibility (semrush.com/ai-seo) documents prompt and source tracking with competitor gap analysis across AI platforms.2
- Profound (tryprofound.com) documents prompt volumes, answer-engine insights, and crawler analytics across several named assistants.3
- Peec AI (peec.ai) documents per-prompt visibility, sentiment and position tracking, source discovery, CSV export, and an API.4
- Ahrefs Brand Radar (ahrefs.com/brand-radar) documents mention, citation, and share-of-voice metrics built on a large pre-collected prompt index.5
A trial you can run in a week
Two patterns stand out once you start testing. Most of these products are built for brands that already have search demand, which is precisely the assumption a smaller company cannot make. And every capability list is a marketing page rather than an independent test, ours included.
So run the trial yourself. Seven steps, one week, no purchase required.
- Write 10-20 real buyer questions. Mix discovery, comparison, alternatives, budget, and at least one scenario question.
- Run the set twice, a day apart. Note how far the competitor set moves. That movement is your noise floor, not your performance.
- Open every returned source URL and check that the page supports the sentence it is attached to.
- Ask for complete answer text rather than a summary. If it cannot be provided, that is your answer.
- Confirm that mention, recommendation, and citation are reported as separate observations.
- Confirm the model, language, market, and search setting are recorded with each run.
- Ask what happens to your history if you stop paying.
The boundary no vendor can cross
Repeated-sampling research keeps landing on the same result: identical questions produce different answers and different citations, and many apparent gaps between two brands sit inside the noise of the measurement itself.6 One protocol paper reports reliability of roughly 0.58 at five iterations, 0.74 at ten, and 0.81 at fifteen, which is why a single run should never be presented as a verdict.7
The honest promise is therefore narrow. A monitoring tool can give you a repeatable record of what a defined model answered, which sources came back, and how that changed under identical conditions. It cannot promise a recommendation, a ranking, a visit, or a sale. When a vendor implies otherwise, that is a reason to keep looking.