PUBLIC SOURCES

What a citation beside a ChatGPT answer can and cannot prove

A source link is worth opening because you can check it. It is not a receipt for how the answer was produced. Here is where that line sits.

In short

  • When a search-backed answer shows a link, you can open it and check whether the page supports the sentence it was attached to.
  • When an answer shows no link, there is nothing to verify. You cannot attribute that text to a specific training page.1
  • Whether a site can appear in ChatGPT's search answers is governed by one crawler's robots rules, which a site owner can actually control.2
  • Buying reviews or seeding community threads to manufacture citations is a search-spam pattern, not a tactic.4

When there is something to check

OpenAI's search documentation is specific here. Answers that used web search may include citations; clicking one opens the source, and on desktop you can hover to preview it. Where the interface offers a Sources button, it lists the cited sources along with related links.1

That is the whole of the reliable case, and it is worth being precise about it. A citation is an invitation to check a claim. It is not a summary of how the answer was assembled. Open it, find the sentence it appears to support, and read the page around that sentence.

  • Does the page support the claim as written, or a weaker version of it?
  • Does it describe the same product, plan, or version the question asked about?
  • Is it current, and is the publisher in a position to know?
  • Would you put this page in front of a customer as your evidence?

When there is nothing to check

Many answers arrive with no links at all. In that case there is no source to open and nothing to inspect. The absence of a citation is not evidence that a particular page was or was not part of a model's training data, and no tool can supply that attribution honestly.

The opposite trap is just as common. A real citation can still be weak: the link may open a page that mentions your category without naming you, a page that has not been updated in three years, or a page that does not distinguish between your product and the one it describes. A source does not upgrade the sentence it sits beside.

What you seeWhat it supportsWhat it does not support
A link shown with the answerThe page exists and was returned alongside this answerThat the page caused the answer
No link at allNothing you can verifyThat a specific page was or was not in training data
A link to a directory or community threadThat a third-party page discusses this categoryThat it recommends your product
A link to a competitor's comparison pageThat a competitor published a comparisonThat the comparison is fair, current, or complete

Who decides whether a site can appear at all

The mechanism a site owner actually controls is smaller and more concrete than most of this discussion suggests. OpenAI documents two relevant agents. OAI-SearchBot is the one used to surface sites inside ChatGPT's search results, and a site that opts it out will not appear in search answers, although it may still be linked as a navigation result. A robots file change takes roughly 24 hours to take effect.2

ChatGPT-User is a different agent. It acts when someone asks ChatGPT or a custom GPT to open a page. It does not crawl the web automatically, and it is not used to determine whether content appears in search. Managing search visibility means managing OAI-SearchBot, not speculating about training data.

Google's guidance for its own AI surfaces reaches a compatible conclusion: appearing in AI Overviews or AI Mode requires no special technical setup beyond being crawlable and indexable, and any structured data you publish has to match what the page visibly says.3

Five kinds of public source, and what each is worth

Once you start opening citations, patterns appear quickly. Sorting them into classes speeds up every later decision, because two of the five classes are not worth pursuing at all.

Source classTypical pageWhat it can proveWhere it stops
OwnedYour own product, pricing, or documentation pageThat you published the factNothing about whether anyone else noticed
Independent directoryA review or software directory listingThat an operator classified you in a categoryNothing about whether that listing is accurate or current
CommunityA forum or subreddit threadThat people discuss this problemAny sentiment you can direct
EditorialA trade publication or news articleThat a publication judged the subject newsworthyAnything you can buy legitimately
Competitor-ownedA rival's comparison or alternatives pageThat a competitor is framing the categoryAnything you control

How to verify one citation in five minutes

This is the check that turns an appealing link into usable evidence. It is deliberately mechanical.

  • Copy the exact sentence the citation was attached to, not the whole answer.
  • Open the URL in a clean browser session and search the page for the subject of that sentence, not just your product name.
  • Check the page's own publication or last-updated date.
  • Classify the page using the table above, and note whether anything on it is yours to correct.
  • Save the URL, the date you checked it, and the run conditions of the answer you were checking.

Why the same question cites different sources next week

Repeated-sampling studies keep finding the same structure: identical queries produce different answers and cite different sources, citation distributions are heavily skewed, and the ranking order across cited domains is unstable from one sample to the next, across the frequently cited set rather than only at the top.5 Many apparent gaps between two domains fall inside the measurement's own noise.

Several things move a result at once: the wording of the question, the model, the language, the market or location context, whether search ran at all, and ordinary sampling variation. That is why a single citation is a lead rather than a finding. Record the conditions, re-run the question later, and compare only what was measured the same way.

A protocol paper on repeated-query auditing puts a number on the cost of skipping repeats: reliability of about 0.58 at five iterations, 0.74 at ten, and 0.81 at fifteen.6 Five runs is a quick read, not a conclusion.

Turn what you find into compliant work

Source discovery is most useful as a work queue, and it starts on pages you own. If an answer rewards a particular attribute, check whether your own page states that attribute in text a machine can read. Proof that exists only inside an image, a video, or a downloadable PDF is invisible to this process.

For third-party sources, contribute accurate information and follow the publisher's own rules. What you must not do is manufacture the appearance of independent coverage. Buying reviews, seeding community threads, or paying for placements written to look editorial are search-spam patterns in Google's policy, and the consequences land on the domains involved.4

What we mean by a citation gap

In our reports, a citation gap means a question where a competitor's answer was connected to public sources and yours was not, within the same run. It describes one measurement. It is not a claim that a missing citation caused anything.

The follow-up is narrow on purpose: classify the source, decide whether it is yours to correct, and fix the clearest factual gap you find on a page you own. Then run the same question again under the same conditions and compare like with like.

If you want the full boundary in one place, the pricing page sets out what the free baseline covers and what stays in the paid plans.