DIAGNOSIS METHOD

Why ChatGPT recommends your competitors

One answer is a sample, not a verdict. A repeatable method for turning an uncomfortable result into observations you can inspect, compare, and act on.

In short

  • Ask the same question twice and you will often get a different shortlist. That is the system working as designed, not a hidden penalty.
  • Sort every answer into four buckets — recommended, listed, mentioned with a caveat, absent — before you count anything.
  • A citation tells you which page was shown next to an answer. It does not tell you which page caused it.3
  • The first defensible move is on pages you own. Third-party work comes later, and never as a purchase.4

One answer is a sample

If you have ever asked an assistant about your own category and watched a competitor get named while your product was not, the instinct is to treat that answer as a verdict. It is not. It is one sample from a stochastic system, produced in one language, in one market, at one moment, against whatever search index existed that second.

OpenAI's search documentation says it plainly: search results and citations can be incomplete, outdated, or incorrect.3 Independent measurement reaches the same conclusion from the outside. Studies of repeated sampling across generative search platforms find that identical queries return different answers and different citations, that citation distributions are heavily skewed, and that source rankings are unstable across samples — across the frequently cited set, not only at the top.1

How unstable is it in practice? A protocol paper on repeated-query auditing of brand recommendations reports a generalizability coefficient of about 0.58 at five iterations, 0.74 at ten, and 0.81 at fifteen.2 A single run is a quick read at best. Everything below exists to turn several samples into a small number of claims you can actually defend.

Write the questions before you look at the answers

The fastest way to fool yourself is to write questions after seeing which ones you already win, or to write one broad question such as 'best CRM for startups' and treat its answer as the whole market. Buyer questions come in kinds, and each kind surfaces a different competitor set.

Question typeExampleWhat it tells you
DiscoveryWhat is the best CRM for a 15-person software company?Who gets named while the buyer still has no shortlist
ComparisonProduct A versus product B for a team of eightWho wins a head-to-head framing
AlternativesAlternatives to a market-leading platform for a small agencyWho owns the switching conversation
Budget or constraintCheapest e-signature tool with a real APIWhether you are even eligible at the buyer's price point
ScenarioInvoicing tool for a freelancer billing in two currenciesWhether you are attached to a specific job to be done

Sort every answer into four buckets

Before counting anything, classify what actually happened in the answer. The gap between a recommendation and a mention is where most dashboards quietly mislead.

  • Recommended — the answer presents the brand as an option for the question that was asked.
  • Listed — the name appears in a list or an aside without being offered as a choice.
  • Mentioned with a caveat — the brand appears alongside a limitation, a warning, or a 'though' clause.
  • Absent — the brand does not appear at all.

Find where the competitor actually wins

Company-level share of voice hides the structure of the problem. A competitor that sweeps comparison questions and never appears in discovery questions has a different weakness — and a different fix — than one that leads everywhere.

Group your results by question type first, then read the language the answer used. Answers tend to attach a specific attribute to a specific brand: integration depth, onboarding time, a price floor, a compliance stance, an ecosystem. That attribute is what to go looking for on your own pages.

This is the step where teams jump straight to 'we need more content'. Wait one more step. If the attribute an answer rewards is already stated on your site, but only inside an image or only in a downloadable PDF, your problem is format rather than volume.

Classify the sources you can actually open

Only answers that used web search hand you URLs you can inspect. An answer with no links gives you nothing to verify, and its wording cannot be attributed to a specific training page.3

For the links you can open, group them before deciding what to do. Each class calls for a different response, and two of the five are not worth chasing at all.

Source classWhat it can supportWhat it cannot
A page you ownFacts you publish: category, customer, limits, proofWhether anybody else noticed
Independent directory or review siteThat an operator classified you in a categoryThat the listing is accurate or current
Community threadThat people discuss this problemAny sentiment you can direct
Editorial or trade pressThat a publication judged the subject newsworthyAnything you can buy honestly
Competitor-owned comparison pageThat a rival is framing the categoryAnything you control

Split what you saw from what you are guessing

Keep two columns and never let them merge. The left column holds an observation: quoted text, the source URL, the run conditions, the date. The right column holds your explanation of it, labelled as an explanation.

ObservationInference
In answer 7 of 20 the competitor was named first, described as best for small teams.The model associates that brand with small teams.
Their help-centre page sits above ours for the setup question.Their documentation is treated as more complete.
No answer in any run mentioned our pricing page.Our pricing approach may be hard to extract, or absent from the pages that were crawled.

Order the work by evidence strength and controllability

Sort the candidate actions into three tiers and work them in order. The temptation is to start with the tier that feels most strategic; it is also the slowest and the least controllable.

  • First, pages you own. If an answer rewards an attribute and your page contradicts it, omits it, or buries it in an image, fix that before anything else. It is fast, it is yours, and the change is verifiable.
  • Second, third-party facts you are eligible to correct — a directory listing with the wrong category, an outdated price note, a product name that changed. Read the publisher's rules and submit accurate corrections only.4
  • Third, independent editorial or expert material. This takes the longest, cannot be bought, and only works if there is something genuinely worth describing.

When not to conclude anything

A method that never tells you to stop is not a method. These are the conditions under which the honest output is 'inconclusive'.

  • No citations at all. You have a sample of one model's behaviour with nothing to verify. Report it that way.
  • A sample too small for the volatility you observed. If the competitor set moved every time you ran the question, you have measured noise.
  • Changed conditions. A model change, a new market, or search switched on or off resets the baseline rather than showing improvement.
  • Commissioned as a story. If the request is to confirm a conclusion rather than to test one, the answer is no.

Then run it again

After you have made a change, run the identical question set under identical conditions. If the numbers moved, ask first whether the measurement changed. A controlled comparison can support a real decision about clarity and proof; it cannot promise that an assistant will recommend you.

Our own reports organise the first tier of work around three directions: state the category and the intended customer explicitly, publish genuinely useful comparison content that includes your own limits, and strengthen public proof a stranger can verify. Treat those as directions to weigh, not a prescription. The evidence you collect should decide which one matters in your market.