Ask ChatGPT a question about your product category and it will name three or four sources. Those citations decide who gets the traffic, the credibility, and increasingly the customer. Yet almost nothing has been published about how they are chosen.
Over six weeks we logged 1,200 answers across commerce, developer tools, and healthcare prompts, then traced every citation to the page it came from. We recorded position, freshness, structure, and authorship for each source, and compared them against the pages that ranked in classic search but never got cited.
Answer engines do not reward the best page. They reward the page that is easiest to quote.
Three patterns held across every category we tested. Cited pages committed to a direct claim inside the first 80 words. They carried a named author with a real footprint elsewhere on the web. And their key claim agreed with at least two other indexed sources, which the model appears to treat as a proxy for consensus.
Citation rate by engine
| Engine | Answers with citations | Avg sources |
|---|---|---|
| Perplexity | 100% | 5.4 |
| ChatGPT with browsing | 78% | 3.1 |
| Gemini | 64% | 2.7 |
If you want the full dataset, the methodology, and the per-category breakdowns, the complete tables are published below. Reproduce it, argue with it, or run it against your own pages with the tools on this site.