How does ChatGPT choose the sources it cites?

When it needs current information, ChatGPT rewrites the question as search queries, retrieves candidates from the Bing index, reads the HTML of the most promising pages and picks 3 to 5 sources to synthesize an answer with citations. Pages win one of those slots when they answer directly at the top, carry high factual density, use clear semantic structure and show verifiable authorship. PDFs and JavaScript-dependent pages never get considered, because the crawler cannot read them in the first place. This article breaks down the five stages of the process, the five factors that decide the pick, and what changes across ChatGPT, Perplexity, Gemini and Claude.

What is the path from question to answer?

Five stages sit between the moment someone types "which asset managers publish the best analysis on interest rates?" and the moment an answer appears with three cited links. Knowing what happens in each one is the difference between optimizing in the right place and optimizing in the dark.

THE FIVE STAGES OF RETRIEVAL
STAGE WHAT HAPPENS WHAT DECIDES IT
1 · Rewriting The user's question becomes one or more search queries Covering the vocabulary users actually use
2 · Retrieval Candidates come from the Bing index Being indexed (Bing Webmaster Tools, IndexNow)
3 · Reading The system downloads and reads the candidate pages' HTML Content in the initial HTML, with no dependence on JavaScript
4 · Selection 3 to 5 sources make it into the final answer Direct answer, sourced facts, verifiable authority
5 · Synthesis The model writes the answer, citing the sources it selected Extractable, unambiguous passages

Source: OpenAI documentation on ChatGPT Search and retrieval analyses, 2025–2026.

Why stage 1 trips up so many firms

The question the investor asks is rarely the query the system runs. "Are inflation-linked government bonds worth buying now?" can become three separate searches: real yields, breakeven inflation, sovereign debt. A page built around one keyword covers one of them. A page that explains the instrument, the context and the decision covers all three. That is why optimizing for an isolated term pays less than answering the whole question.

Why stage 2 is the most neglected

ChatGPT keeps no web index of its own for real-time search. It uses Bing's. So an asset manager can publish the best content in the country and never get retrieved, for the single reason that nobody ever registered the domain in Bing Webmaster Tools. The step is free, it takes minutes, and most of the Brazilian financial industry still has not taken it.

Why stage 3 eliminates most asset managers

Here the system downloads the HTML and reads what is there. It does not open a browser, execute scripts or wait for components to load. If the text only appears after JavaScript runs, all the model receives is an empty shell. And if the content sits inside a PDF, extraction is expensive and the structure is gone. Both kinds of invisibility are covered in why AI can't read PDFs.

What decides the pick among the pages that survive?

By stage 4, the system has a handful of readable candidates and has to pick three to five. ReBo groups the criteria behind that pick into five factors, each broken into concrete variables. This is the checklist that separates content that gets cited from content that stays invisible.

THE FIVE FACTORS AND THEIR VARIABLES
FACTOR VARIABLE WHAT TO MEASURE OR APPLY
1 · Readable format Page type Content on an indexable HTML page, not a PDF, not an image, not behind a login
Selectable text Real text in the DOM, copyable, not rasterized or trapped inside <canvas>
Rendering Content in the served HTML, not injected only by heavy JavaScript
Unique, stable URL Every piece of content on its own permanent, descriptive URL
Robots and indexing meta robots index,follow; no block in robots.txt; present in the sitemap
Weight and access Light page, no paywall or pop-up standing in front of the content
2 · Semantic structure Heading hierarchy A single H1, with H2s and H3s in logical order describing each section
Headings as questions Subheadings phrased the way users ask the question
Self-contained blocks Each section answers on its own, without leaning on the previous paragraph
Navigable table of contents A contents list with anchor links to every section
Lists and tables Data in real <ul>, <ol> and <table> elements, never in an image
Structured FAQ A question-and-answer block at the end, marked up with FAQPage schema
Short paragraphs Ideas in direct blocks, easy to lift out as a quotable passage
3 · Metadata Title and meta description Unique, descriptive, carrying the central concept
Named author Real name of the portfolio manager or analyst, job title and a bio page
Visible dates Publication and last-updated dates stated explicitly
Open Graph and Twitter Card og:title, og:description and og:image all filled in
Schema.org (JSON-LD) Article, FAQPage, Organization and Person markup
Canonical Pointing to the definitive URL, avoiding duplicates
Institutional signals About page, corporate registration, regulatory licences, contact details, compliance
4 · Data density Numbers with units Precise values in %, BRL, USD or basis points, not "rose sharply"
Dates on the data Every figure anchored to a reference period
Primary sources cited The origin named in the body text (statistics agency, central bank, treasury)
Verifiable claims Factual, attributable sentences that are safe to quote
Defined terms Technical concepts explained within the text itself
Data tables Time series and comparisons in machine-readable tables
5 · Evergreen and linking Timeless content Pillar pages on subjects that stay relevant
Thematic internal links Each article links to 3 to 6 related pieces on the same site
Pillars and clusters A central guide connected to satellite articles on the same theme
Recurring updates Content reviewed on a schedule and re-dated
Consistent coverage Volume and regularity within the same field of expertise
External mentions Being cited by other sites and by the press reinforces authority

Source: ReBo citation-factor framework, June 2026.

Factor 1: why does format decide before merit does?

This is the knockout factor. Before it can judge whether your analysis is any good, the system has to be able to read it. A PDF is a stream of bytes designed for printing: text-positioning instructions, embedded fonts, tables that dissolve into loose cells with no header and no unit. The hierarchy a human sees exists only as font size, not as structure.

The test could not be simpler, and it needs no tooling. Open the page, hit "view source" and look for the text. If it is not there, it does not exist as far as the AI is concerned. The same applies to content in PDFs, to JavaScript-rendered sites, to text inside images and to anything behind a login or a form.

Factor 2: what does semantic structure change in practice?

Structure is what lets the model extract a passage and quote it with confidence. A two-thousand-character wall of prose may hold the best answer on the market, but the system still has to decide where that answer begins and ends, and the ambiguity costs the citation.

The most underrated device here is the question-form heading. An H2 that reads "Macro outlook" describes a topic. An H2 that reads "Why isn't the deficit coming down?" is an almost literal match for the query the system just formulated at stage 1. The second form wins consistently, and it costs nothing to adopt.

The second device is the self-contained paragraph. Text written for linear reading leans on "as we saw above" and "as mentioned earlier". For retrieval, every paragraph has to make sense in isolation, because in isolation is exactly how it will be read. Restating the subject instead of referring back to it costs five words and keeps the passage intact as a quotable unit.

Factor 3: why does metadata matter more in financial content?

Metadata answers who wrote this, when, and about what, without making the model infer any of it. In financial content it counts for more than in other sectors, because a wrong answer has consequences. Retrieval systems run conservative on money, health and legal topics, and they favor sources whose authorship and date they can verify.

In practice that means the portfolio manager's real name, with job title and bio page, rather than "Research Team". It means visible publication and last-updated dates. It means JSON-LD marking up Article, FAQPage, Organization and Person. And for a regulated firm it means surfacing the signals the market expects: an about page, corporate registration details, regulatory license numbers, a contact channel and compliance disclosures.

Factor 4: what is factual density, and why does it decide the outcome?

Factual density is the share of verifiable statements per paragraph. "Inflation eased" is a compressed opinion. "May's IPCA-15 came in at 0.62%, according to IBGE" is a fact with a number, a unit, a period and a primary source, and the model can quote it without taking on risk.

This is the difference between being read and being cited. A page can be perfectly legible and still never appear in an answer, because it offers nothing worth reproducing. Four elements raise that density immediately: a number with its unit, the date the figure refers to, the source named inside the text, and an explicit definition of every technical term used.

The higher the factual density, the more citable the content. That is why every block preserves numbers, dates and named entities together with their source.
— ReBo Method, structuring principle

Factor 5: why is a standalone page rarely cited?

The first four factors are about the page. The fifth is about the domain. Retrieval systems assess whether a site is a recurring source on the subject or a one-off appearance, and that reputation is built through consistent coverage, internal linking and external mentions.

Internal linking means each article points to 3 to 6 related pieces on the same site, forming pillars and clusters around a theme. External mentions mean being cited beyond your own domain: press, communities, aggregators. AI citation frequency tracks brand mentions off-site, which is why publishing well without distributing it rarely produces citations.

How does this show up in the Brazilian market today?

In June 2026 ReBo examined three Brazilian firms to observe the pattern. Two of them publish their analysis on structured web pages, and both showed up as ChatGPT citations on questions about the macro outlook. The third produces monthly commentary of comparable quality, with dense fiscal analysis, hard data and primary sources, but publishes it as a PDF in a list of links.

The third firm's material did not lose on merit. It lost on all five factors at once. The content sat inside a PDF with no HTML hierarchy and no anchors, on a listing page with no descriptive title, no author, no date and no per-issue schema. Excellent data, locked inside the file, with no thematic links between editions. It is a textbook Level 1 in the AI Visibility framework: invisible to AI.

The practical conclusion is uncomfortable and liberating at once. The contest for the three to five citations in each answer is not won by whoever has the best analysis. It is won by whoever has the readable analysis. A boutique with well-structured content beats a large house that keeps everything in PDF.

How do you write a paragraph an AI can cite?

Write for someone who will read exactly one paragraph of your text, out of context, and then has to decide whether it is safe to reproduce. Roughly 44% of citations come from the first 30% of the text, so the complete answer belongs at the top — not assembled piece by piece down the page.

THE SAME CONTENT, WRITTEN TWO WAYS
VERSION THAT NEVER GETS CITED CITABLE VERSION
Claim "Inflation surprised to the downside this month." "May's IPCA-15 came in at 0.62%, below the Focus survey median, according to IBGE."
Cross-reference "As we saw above, this changes the outlook." "The slowdown in IPCA-15 changes the interest rate outlook for the second half."
Heading "Macro outlook" "Why did inflation ease in May?"
Data point "The currency strengthened considerably." "The real appreciated 3.1% against the dollar over the month, closing at 5.16 per US dollar."
Technical term "We increased duration." "We increased duration, that is, the portfolio's sensitivity to changes in interest rates."

Source: ReBo AI-friendly writing guidelines, 2026. The figures illustrate the format and are not a recommendation.

Notice that the citable version is neither longer nor more technical. It just leaves nothing implicit. Every sentence carries the subject, the number, the unit, the period and the source that make it verifiable on its own.

Do Perplexity, Gemini and Claude choose the same way?

The general process is the same across all of them: search, rank candidates, synthesize with citations. What differs is where the candidates come from and how many sources make it into the answer — and that decides where your indexing effort is worth spending.

WHAT CHANGES BETWEEN THE ENGINES
ENGINE WHERE CANDIDATES COME FROM PRACTICAL CONSEQUENCE
ChatGPT The Bing index Registering the domain in Bing Webmaster Tools is a prerequisite, not a detail
Perplexity Its own index, with a dedicated crawler Tends to cite more sources per answer; allowing PerplexityBot matters
Gemini The Google index Search Console and Google's indexing best practices still apply
Claude Built-in search with its own crawler Requires static HTML; ClaudeBot does not execute JavaScript
Copilot The Microsoft ecosystem and the Bing index Shares the same indexing bottleneck as ChatGPT

Source: ReBo engine mapping, 2026.

The right way to read this table is that the fundamentals do not change. Readable HTML, semantic structure, factual density and verifiable authorship apply to all five. What varies is the indexing channel, which is why the diagnostic measures the five engines separately instead of treating "AI" as a single thing.

What knocks a page out of contention?

Being unreadable is the main one, and Factor 1 covered it. Three other disqualifiers come from configuration oversights rather than format choices, and they usually go unnoticed for months.

None of these three is a content problem. They are infrastructure settings that quietly wipe out months of editorial work, which is why the technical audit comes before production in the ReBo method.

How do you tell whether you are being cited?

AI citation is not deterministic. The same question produces different sources on different runs, because the system rewrites the query and re-scores the candidates every time. So a single answer measures nothing. Frequency measures everything.

Define a set of target questions — typically 30 to 80 — that clients, allocators, journalists and consultants genuinely ask about your market, then run them across the five engines on a fixed cadence. The metric that comes out is your citation rate: the share of those questions where your firm shows up. Measured against a baseline, it shows movement. Measured once, it shows nothing.

Two other signals back up the reading. The first is your server logs: hits from GPTBot, ClaudeBot and PerplexityBot are hard proof that the crawlers are reading you. The second is referral traffic arriving from the AI platforms, which shows the human effect of being cited.

To find out where your firm stands in that contest today, the ReBo diagnostic measures the real answers from all five engines and records who is showing up in your place.

Frequently asked questions

Do I need to be indexed in Bing to be cited by ChatGPT?

In practice, yes. ChatGPT search uses the Bing index as its retrieval base, and Copilot draws on the same ecosystem. Registering the site in Bing Webmaster Tools and enabling IndexNow speeds up indexing. It is free, takes minutes and is one of the steps asset managers most often skip.

Why does the same question return different sources each time?

Because the selection is not deterministic: the model rewrites the query and re-scores the candidates on every run. That is why visibility is measured as a citation rate across multiple runs and a defined set of target questions, never from a single answer.

Do Perplexity and Gemini choose sources the same way?

The general process is similar, but each platform has its own index and its own criteria: Perplexity runs its own index and tends to cite more sources per answer, while Gemini starts from the Google index. The fundamentals — readable HTML, semantic structure, factual density and verifiable authorship — apply to all of them.

Does new content stand a chance against large sites?

Yes. AI systems cite the page that best answers the specific question, not the biggest brand. Narrow pages with a direct answer and proprietary data routinely beat generic portals, especially on long-tail questions. And because almost the entire financial sector still publishes in PDF, the field is close to empty.

How many sources does ChatGPT cite per answer?

Typically 3 to 5 when it runs a live search. That small number is what makes the contest so sharp: there is no second page of results to fall back on, just three to five slots per question.

Will structured data alone get the AI to cite me?

No. JSON-LD helps the machine extract author, date, question and answer without ambiguity, which removes friction at the reading stage. But it is no substitute for readable, dense, sourced content. Schema on top of an empty page is still an empty page.

How long does it take for a new page to be cited?

Crawlers such as GPTBot and PerplexityBot revisit active sites within days or weeks, so the reading part usually happens fast. Consistent citation depends on volume, freshness and external signals, which is why it builds over months rather than with a single post.

Keep reading

For the two kinds of invisibility in detail, see why AI can't read PDFs. For the discipline that works on these factors, what GEO is and GEO vs SEO. For the umbrella concept, what AI Visibility is. For the view specific to the financial industry, AI Visibility for asset managers. The technical terms used in this article are defined under AI crawler, structured data, E-E-A-T and citation rate.