Not every ChatGPT answer has a citation, because not every ChatGPT answer touches the web. A citation only appears when the underlying search tool actually runs. Understanding when that happens, and what happens next, tells you exactly what a site needs to have in place to be eligible for a citation at all.
When ChatGPT actually searches
According to OpenAI's own help center article on searching the web with ChatGPT, "ChatGPT may search the web automatically when your question would benefit from current information," and a person can also trigger it manually by selecting Search from the tools menu. Answers that rely on the model's training data alone carry no citations, because nothing was retrieved. This matters for site owners: a page can be well written and well structured and still never get cited, simply because the query it would answer never triggers a live search in the first place.
What happens once a search runs
OpenAI's documentation describes the mechanism plainly: "ChatGPT search typically rewrites your query into one or more targeted queries that it sends" to a search provider, reviews the results, and — for complex questions — can send additional, more specific follow-up queries before writing an answer. A biotech question about a specific drug class, for example, might be rewritten into a targeted query, reviewed, then followed by an even more specific one before ChatGPT settles on which sources to name. The response it composes cites a handful of the pages retrieved during that process, not the training data.
The crawler distinction that decides eligibility
OpenAI runs more than one bot, and confusing them is the single most common mistake site owners make when managing robots.txt. GPTBot crawls pages to train OpenAI's models. OAI-SearchBot is a separate, distinct crawler that builds the index ChatGPT's search feature retrieves from. ChatGPT-User fetches a page live when a person explicitly asks ChatGPT to visit a specific URL. Blocking GPTBot affects training data only — it does not remove a page from ChatGPT search results. Blocking OAI-SearchBot does the opposite: it removes a site from ChatGPT's citable search index regardless of how well the page otherwise performs. OpenAI states this directly: to make a site "eligible for inclusion" in ChatGPT search, a site owner must "allow OAI-Searchbot to crawl the site and confirm that the website host or content delivery network allows traffic from OpenAI's published searchbot IP addresses," which OpenAI publishes at openai.com/searchbot.json.
In practice, eligibility usually breaks in one of three places: a robots.txt rule aimed at GPTBot that does not actually cover OAI-SearchBot, a host or CDN silently blocking OpenAI's published crawler IP ranges at the network layer where robots.txt alone would not reveal it, or a page that requires a login or renders its content only in JavaScript, so an unauthenticated crawler never retrieves anything to cite in the first place. Any one of the three produces the same symptom — a well-written page that never appears — which is why a proper audit checks all three rather than assuming robots.txt is the whole story.
What Google's index has to do with it
OpenAI has not published the exact composition of the index behind ChatGPT search, and multiple independent crawler trackers describe OAI-SearchBot as working alongside licensed search-provider data. Because OpenAI does not publish exact ranking weights, any specific numeric claim about "how much" a factor matters comes from a third party's own study, not from OpenAI. One such study worth naming directly: in April 2026, Search Engine Land reported on a case study finding that ChatGPT citations aligned with Bing's top-ranked results roughly 87% of the time across the queries tested — far more often than they aligned with Google's own top rankings for the same queries. If that pattern holds broadly, a page's standing in Bing's own index would matter more for ChatGPT citation odds than most SEO teams historically assumed, given how much attention has gone toward Google specifically. Treat this as one dated, attributed case study rather than a confirmed OpenAI mechanism — it is the kind of claim that should always carry its source, not be repeated as settled fact.
Being indexed is the floor, not the whole answer
Passing the crawler check only makes a page eligible to be considered — it says nothing about whether it gets picked over a competing page once retrieved. OpenAI's help center states that "ChatGPT ranks search results using multiple factors intended to help users find relevant, reliable information," and adds directly that "placement is not guaranteed." OpenAI does not publish the exact weighting of those factors, so any claim naming specific percentages beyond that statement is a third party's own research, not an OpenAI-confirmed mechanism. Independent research on adjacent generative-answer systems consistently points to the same practical pattern regardless of exact weighting: a query rewritten by the assistant needs a candidate page that answers it in a directly quotable sentence near the top of the content, because a synthesized answer is built by lifting short, specific spans of text rather than summarizing an entire article end to end. A page that buries its actual answer under three paragraphs of preamble is harder to extract from than one that states it plainly in the first two sentences, even when both pages are equally crawlable.
Sources and results are not the same as accuracy
OpenAI is explicit that citation is not a guarantee of correctness: "search results and citations can be incomplete, outdated, or incorrect," and it recommends users "open a cited source to check that it supports the answer" before relying on it. For a site owner, the practical implication is the same one that runs through every check in our own scoring: being retrievable and being cited is necessary but not sufficient. The content still has to say something specific, current and verifiable enough that, once retrieved, it earns the citation rather than getting summarized without attribution or skipped in favor of a clearer competing source.