Why the demo cannot help you
Every tool in this category does the same thing in a demo. You type a question about your category, an assistant answers, and a dashboard lights up with your brand name and a percentage. On a single run they are telling you almost the same thing, which is why they are so hard to tell apart.
The differences live in what happens on the second run, and in what the tool kept from the first one. That is not something a sales call can show you, because it takes time to produce. It is also the reason the trial matters more than the feature list.
- What a demo shows: the tool can reach a model and render a chart.
- What a demo cannot show: whether two runs agree with each other.
- What only your own trial shows: whether answers are stable enough in your category to support a decision.
Before you open any trial, write 20 to 30 questions
The most useful preparation happens before you see anyone's interface. Write down the questions you would actually want answered, in your customers' words, and keep the list somewhere you control. Twenty to thirty is enough to start. The point is that the list is yours, not a vendor library tuned to produce flattering charts.
Write the messy ones too. "Best tool for a small team" and "is this worth it for a two-person agency" are different questions, and they can produce different answers. If you only test the tidy version, you will not learn how the tool handles the way people really type.
- Discovery: what should I use for this job?
- Comparison: this option against that one, for a specific situation.
- Validation: is this any good, or does it have a known problem with something?
- Budget: what is available under a specific monthly figure?
- One question about your own brand, phrased exactly the way you fear a customer would type it.
Run the set once and save the raw answers
Run all of them, and before you look at any score, export or copy the full answer text. A percentage with no answer underneath it cannot be checked later, and the answer is where the useful detail lives: which competitors were named, in what order, and which pages were cited alongside them. Those pages come from the search step the provider documents, and they are the part of the answer a human can actually open and read.3
Watch for whether the tool separates three different observations. A brand can be mentioned in an answer, recommended by it, or cited as a source on the page. These are not the same event, and a tool that collapses them into one visibility figure is telling you less than it appears to.
- Mentioned: your name appears somewhere in the answer.
- Recommended: the answer puts you forward as an option to consider.
- Cited: one of your pages is returned as a source for the answer.
- A single number covering all three has made a choice on your behalf.
Run it again the next day, then compare the two lists
This is the step almost nobody does during a trial, and it produces the most information. Same questions, same conditions, one day apart. Then look at what moved.
Some movement is expected and it has been measured. A published protocol for repeated-query auditing of brand recommendations puts numbers on it: the generalizability coefficient is about 0.58 at five iterations, 0.74 at ten, and 0.81 at fifteen.1 Five runs is a quick read, not a conclusion.
If two runs produce completely different competitor sets, that is not automatically a reason to reject the tool. It is a reason to stop treating a single number as a fact, and to ask the vendor what they do about the variance.
Open the citations and check three of them by hand
Whatever the tool claims about sources, open a few and look at them. A cited page should actually contain the claim it is attached to. This takes ten minutes and it separates tools that store real URLs from tools that display domain names as decoration.
- Does the URL load, and is it a page you would want a buyer to read?
- Does the page actually support the sentence it was attached to?
- Is the source your own site, a competitor's, or an independent third party?
- If a competitor was cited and you were not, is there a page on your site that should have been the better answer?
The number that will mislead you
A single-run visibility percentage looks precise and is not. Research sampling citation behaviour across three generative search platforms found citation distributions following a power-law form, with rank instability not only at the top but throughout the frequently cited set, and many apparent differences between two domains falling inside the noise floor of the measurement itself.2
So when a dashboard shows 38% against a competitor's 41%, the honest reading is that you cannot separate those two from one run. What you can compare is movement across runs you controlled yourself, on questions you wrote.
Five things you should be able to say at the end of the trial
If you can answer all five of these, you have enough to decide. If you cannot answer two or more, the trial tested the demo rather than the tool.
- The questions were mine, and I can run the same set again without asking anyone.
- I have the full answer text, not only scores.
- I know whether this tool treats a mention, a recommendation and a citation as different things.
- I opened at least three cited pages and they held up.
- I know how much the competitor list moved between two identical runs.
What the trial cannot tell you
It cannot tell you that fixing the gaps will produce a recommendation. No tool can promise that, and any vendor who does is selling something other than measurement.
It cannot tell you why a competitor was recommended. You can see that they were, and which pages appeared near the answer, but the causal step is not observable from outside the model.
What it can do is replace a guess with a measurement you ran yourself. That is a smaller claim than the category usually makes, and it is the one that survives a second look.