Which sources do AI engines use for answers?
AI engines pull from different places. Perplexity and Gemini lean on live web results with visible citations, ChatGPT and Claude mix training data with web search when a query triggers it, and DeepSeek leans mostly on training. Across all five, a handful of source types recur: Wikipedia and Wikidata, Reddit and forums, review platforms like G2 and Trustpilot, established news and editorial, and well-structured first-party pages.
What each engine leans on
The five engines do not answer from one shared index. Each has its own mix of training data and live retrieval, and that changes which sources you see cited.
- Perplexity is retrieval-first. Nearly every answer runs a live web search and shows numbered citations inline. If you want to see raw source behaviour, this is the clearest window.
- ChatGPT answers from training data by default and browses the live web only when the query looks time-sensitive or the model decides it needs to. When it browses, it surfaces links; when it does not, the sources are baked invisibly into training.
- Gemini is wired into Google's index and shares plumbing with AI Overviews, so its citations skew toward pages that already rank in Google Search.
- Claude answers from training data and runs web search when the question needs current information, citing the pages it pulls.
- DeepSeek leans most heavily on training data, with more limited live retrieval, so its answers reflect what was well-represented on the web at training time rather than today's results.
The practical takeaway: the more an engine retrieves live, the more your on-page content and fresh citations matter right now. The more it leans on training data, the more your long-standing presence across the web matters.
The source types AI pulls from repeatedly
Look across enough answers and the same handful of source types keep appearing, regardless of engine. If your brand is absent from these, you are absent from the answer.
- Wikipedia and Wikidata. The reference layer. Engines treat them as a trusted anchor for what an entity is, and Wikidata feeds knowledge panels that Gemini and Google draw on.
- Reddit and forums. Heavily cited for opinion, comparison, and real-user experience. "Best X for Y" answers frequently quote a Reddit thread.
- Review platforms. G2, Capterra, and Trustpilot for software and services; category-specific review sites elsewhere. Engines read these as third-party validation.
- Authoritative editorial and news. Established publications, industry outlets, and well-known listicles carry weight that a self-published blog post does not.
- Well-structured first-party pages. Clear, factual pages on your own domain, ideally with schema markup and direct answers to specific questions, get lifted when they are the cleanest source for a fact.
How to find which sources cite you
You do not need a tool to start. The vendor-neutral method is to read the citation panels the engines already show you.
- List the 5-10 prompts a buyer in your category would actually type, like "best CRM for small teams" or "alternatives to Mailchimp".
- Run each in Perplexity and in Gemini, the two engines with the clearest citation panels.
- Read the cited sources. Note which domains show up, whether you appear at all, and which competitors do.
- Repeat monthly, because live-retrieval answers shift as the web changes.
Here is what that looks like in practice. Run this in Perplexity:
best project management software for remote teams
The citation panel typically lists a G2 category page, one or two Reddit threads, a listicle from an established software-review site, and the product pages of the tools named. If your tool is not in that panel, it is not in the answer, and the gap tells you exactly which source types to go earn: a G2 presence, a mention in the ranking listicle, or a Reddit thread where people discuss you.
Doing this by hand across five engines and dozens of prompts every month is the tedious part. That is the work avisibli automates: it runs your category prompts across ChatGPT, Perplexity, Gemini, Claude, and DeepSeek on a schedule and shows you which sources each engine cited, so you can see whether you are in the answer and which sources to go win.
avisibli is the GEO platform that publishes this answer library. Self-references are limited to topics where a tool-based answer is genuinely useful to readers.