Where Does AI Get Its Information? The 4 Sources

Team AllAble logoWritten by Team AllAble
Back to blog
Dark four-source diagram on a laptop screen, four glowing source nodes feeding one AI answer card, navy background with orange and blue accents

When an AI assistant answers a question about your brand, it learned about you four different ways. Two of them you can change this quarter. Do you know which two?

Open ChatGPT, Claude, Perplexity and Google's AI Mode in four browser tabs and ask each one the same question about your category. You will get four different answers, and at least one of them will name a competitor you have never lost a Google position to. That variety is not noise. Each assistant reads from a different combination of four information sources, and the mix decides whether your brand appears at all. Here is where the asymmetry hides: two of those sources respond to work your team can start tomorrow, and two do not answer to marketing at all. I have watched clients spend an entire quarter optimizing the wrong two.

Where Does AI Get Its Data? The Four Sources, in Plain Language

Ask this question and you get two very different families of answer, depending on who is answering.

A research library will tell you it is about citation practice. The University of Maryland's library guide puts it as: "Typical contexts would be looking up information from the internet or internal company documents before generating a chatbot response." That is accurate and it answers a student's question. It does not tell a marketing team anything to act on.

A vendor will hand you a list. GWI's version is the four-source framing worth knowing, verbatim: "Web-scraped content · Licensed or proprietary data · Human-labeled training sets · Consumer panels and market research." It is a clean teaching list. Read it closely and you notice two of the four categories are what GWI sells, which makes it a good lesson and a poor operational map.

Here are the four sources a marketer actually needs, each with the answer that matters more than the definition: can you influence it, and how fast?

Four frosted glass source columns with flow lines converging into a single glowing AI answer node on a deep navy background

1. Training data: the frozen layer

Training data is everything a model absorbed before it was released. Vast Data's explainer describes the mechanism accurately: data, "mostly unstructured, is gathered from public online datasets (like Wikipedia and Common Crawl), customer data, industry-specific archives, file shares, object stores, and more."

That page sits at position 2 for this query, and it was published on 15 May 2025. It is 16 months old in a field that changed at least twice since. I am not raising that to score a point. Freshness is a large part of why the marketing angle is missing from this SERP.

OpenAI states its own version of this layer on its foundation-model documentation, first party and worth quoting in full:

"OpenAI's foundation models, including the models that power ChatGPT, are developed using three primary sources of information: (1) information that is publicly available on the internet, (2) information that we partner with third parties to access, and (3) information that our users, human trainers, and researchers provide or generate."

Three sources, from the company that builds the model. Hold onto that list. It maps onto this article's four, and it is missing one that now matters most.

Can you influence it? Indirectly, slowly, and only through what other people publish about you. The cutoff is already set. The one lever you control is a crawler instruction: disallowing the training crawler in robots.txt. That is a visibility trade-off, not a win.

2. Retrieval and live search: the layer that answers today

This is the layer that answers the question you asked five minutes ago. When an assistant retrieves, it searches live, pulls documents, and writes the answer on top of them.

Ahrefs explains the mechanic better than anything else in the top ten: "A RAG-enabled model can look things up first, then answer." Their closed-book versus open-book exam analogy is the clearest mental model available. Same page, on grounding: "When an AI answer is grounded, it is tethered to specific retrieved sources, which dramatically reduces the hallucination risk."

This is where Google AI Overviews live, and it is why Google AI Mode behaves so differently from a plain search result. The same applies to Perplexity and to any grounded chat answer.

Can you influence it? Yes, directly, and faster than anything else on this list. This is the layer your content team already has the skills for.

3. Licensed and proprietary data deals: the layer money buys

Some of the information an AI system knows was not crawled. It was purchased.

Google's reported Reddit licensing deal, signed in February 2024, is quoted at roughly 60 million US dollars per year, structured as a training and grounding contract for Search and Gemini. OpenAI's reported Reddit deal, signed in May 2024, is quoted at over 70 million dollars and grants ChatGPT retrieval access to the corpus.

Read those two numbers again and the strategic picture simplifies. One of your four sources is a procurement decision made at a level your marketing department cannot reach. No content calendar competes with a contract like that.

Content provenance now has a regulatory layer too. EU AI Act Article 50 transparency provisions became applicable on 2 August 2026, and Anthropic began marking output from Claude models launched on or after that date, with files carrying signed C2PA provenance metadata. Useful context if you work with machine-generated assets. It is not a ranking opportunity, and I would not build a strategy around it.

Can you influence it? Through a contract, if you have one to sell. For most marketing teams, the honest answer is no.

4. Live tool access and agent connectors: the layer being built now

The fourth source is not a corpus at all. It is a live connection between the assistant and a system that holds an answer: an API, a database, a connector, an MCP server.

The timeline is short. OpenAI adopted the Model Context Protocol at DevDay in October 2025. In December 2025 it renamed connectors to apps. In July 2026 the directory merge turned apps into plugins. Three names for one mechanism in eighteen months. Custom remote connectors launched for Claude on Pro, Max and Enterprise on 13 March 2026, and write-capable connectors on ChatGPT are gated to Business, Enterprise and Edu workspaces, with Plus and Pro users limited to read and fetch. More than 10,000 MCP servers now exist. Research Solutions shipped Scite MCP on 26 February 2026, wiring ChatGPT, Claude, Copilot, Cursor and Claude Code into 250 million scientific articles. That is a content layer that answers questions with no web page retrieved at all.

Clay's roadmap post is the clearest signal that this is being built right now. It is dated 24 July 2026, and its summer release programme ran through late September 2026. The line that matters: "Everything described, from the data up through execution, will be accessible to agents and through the CLI." The same series states: "you shouldn't have to log into Clay to use Clay."

This layer is also visibly early. Documented 2026 problems include custom apps intermittently vanishing from the directory, OAuth flows completing without the connector appearing in chat, and tool definitions freezing at approval time so a server-side change needs a manual admin refresh.

Can you influence it? Only if your product ships a connector. For a marketing team, this is someone else's roadmap.

Why the Answer Changes Depending on Which Assistant You Ask

Three variables explain almost all of the variation you see between assistants.

Training cutoff. Each model froze at a different point. Ask about something that happened after a model's cutoff and you will get a confident answer built from older retrieval or older training.

Retrieval index. Google's AI Overviews and AI Mode retrieve from Google's index. ChatGPT retrieves from its own. Perplexity retrieves from its own. Different indexes, different sources, different winners on the same prompt.

Connector availability. If the person asking has a connector installed that your competitor's data sits inside, the answer does not come from the web at all.

The overlap numbers show how little these three variables have in common, and they are the reason AI brand visibility has to be measured per assistant rather than in aggregate. Per 5WPR's State of AI Citations 2026, only about 11 percent of domains are cited by both ChatGPT and Perplexity. Wikipedia accounts for nearly half of ChatGPT's top-10 source share, while Reddit accounts for roughly 46.7 percent of Perplexity's. On Google AI Mode, Wikipedia shows up in only around 2 percent of responses, and LinkedIn appears in nearly 15 percent.

One more number, with its method attached because the methods disagree: Ahrefs Brand Radar, measuring 3 million US queries in June 2026, found YouTube the most-cited domain in AI Overviews at 20.9 percent, up 34 percent in six months. Other studies using different datasets produce different shares. Treat all of these as directional. What they agree on is the shape, not the decimal.

The Three-Layer Model, and Why It Matters More Than the Four-Source List

Ahrefs published the strongest structural answer on this SERP, and it deserves credit by name: "AI gets its knowledge from three distinct layers: training data, retrieval systems, and live tool access like APIs and MCPs."

That framework is better than a flat list because it tells you what kind of memory is answering. Its weak point is that it never gives you an order of operations. You learn what the layers are. You do not learn which one to work on first.

The four-source list and the three-layer model do different jobs. Here is how they line up, with the verdict each one needs and neither one provides.

Source

Ahrefs layer

Can you move it?

How fast

Training data

Training data

Indirectly, via third parties

12 months or more

Retrieval and live search

Retrieval systems

Yes, directly

Weeks

Licensed and proprietary data

Training data

No, by contract only

Not your call

Live tool access and connectors

Live tool access

No, unless you ship a connector

A product decision

OpenAI's own three-source list maps onto this cleanly. Its source 1 covers both training data and retrieval. Its source 2 is the licensed and proprietary layer. Its source 3, information from users and human trainers, is what most people call human-labelled data. Its page does not mention live tool access at all, which is the newest layer and the one that changed the most in 2026. The vendor documentation is one layer behind its own product surface.

So: four sources, mapped onto three layers, cross-checked against the first-party list, and then ranked by how fast a marketing team can move each one. That rank order is the part nobody has published, and it is what the rest of this page does.

What a Marketer Can Actually Influence This Quarter

Two of the four sources respond to work your team already knows how to do. Here is each one, with an owner and a metric attached, so it does not turn into another strategy deck.

Dark execution dashboard with two active influence levers lit beside two locked ones, joined to a rising measurement gauge in orange and blue

Retrieval: structure and factual density

When an assistant retrieves, the unit of competition is not the page. It is the extractable passage inside it. A model does not cite your article. It cites the paragraph that answers the question.

Citation studies keep landing on the same three formatting patterns: the direct answer sits in the top third of the page, self-contained answer blocks of roughly 40 to 60 words get quoted more often than answers buried in long paragraphs, and roughly one linked statistic per 150 to 200 words outperforms lower factual density. The exact multiples vary by study and I would not bet a budget on any single one. The direction is consistent, and none of it costs anything to follow.

On schema, be careful, including with advice you read on this site. There is a widely shared claim that structured data multiplies AI citations. Our own controlled testing found no such effect, and you can read what we measured in schema and AI visibility. Treat structured data as hygiene: it makes your entities unambiguous and your facts machine-readable. That is entity optimization, and it is worth doing for the clarity, not for a multiplier that does not reproduce.

Owner: whoever owns the content. Metric: generative AI impressions per page in Search Console, which we get to in a moment.

Citation: what makes a source quotable

Being retrieved and being quoted are different events. A model can read your page, use it to build an answer, and never name you. Quotability is a property of the passage, and it is why LLM optimization is mostly about how you write a paragraph rather than how you write a page.

Two findings should encourage you. Ranking inside the top 10 is not the gate: only a minority of AI Overview citations come from top-10 ranking pages. And brand-owned websites are gaining: Presenc AI's April 2026 analysis of 84,000 queries found the citation share going to brands' own websites rose from 26 to 31 percent in twelve months. If someone told you that only third-party sites get quoted, that was true once and is less true now.

What makes a passage quotable is unglamorous. One claim per paragraph. The claim stated in the first sentence. Numbers with a date. No pronouns pointing at sentences the model cannot see. If you want the detailed version of this, how to rank on ChatGPT covers the retrieval-side mechanics.

Owner: the content lead. Metric: how often your brand is cited across a fixed set of prompts, tracked on a schedule, because citation sources move even when the answer does not.

Presence in third-party sources

Ahrefs has the strongest execution sentence on the whole SERP, and it is worth repeating: "A brand that exists only on its own domain is largely invisible to the model's training data." Wikipedia entries, forum discussions, third-party reviews and press coverage sit in both the training layer and the retrieval layer.

Reddit is the case everyone asks about. The figure that circulates most is 40.1 percent for Reddit's share of AI citations, and a thread on r/it at position 6 is where most people meet it: "Reddit tops the list at 40.1% because it's a goldmine of user-generated discussions, but that doesn't mean AI blindly parrots it." Treat it as a claim made in that thread, not as a measured fact. The number traces back to a Semrush study from June 2025 that analyzed 150,000 citations across 5,000 keywords, and what it measured is citation frequency: which links AI responses pointed to. It is not a measurement of training data, and the two get conflated constantly. Independent datasets disagree with it by methodology, putting Reddit anywhere from around 5.5 percent to 21 percent of AI Overviews.

Then there is the volatility. A Semrush three-month study found ChatGPT's Reddit citations collapsing from roughly 60 percent of responses toward 10 percent in late 2025 before recovering, with the citation share shifting toward PR Newswire, Forbes and Medium. A single-source dependency is a strategic vulnerability, and Reddit is the documented example. If you want to work this channel properly, Reddit SEO covers the mechanics without the astroturfing.

The measurable part is brand search volume. 5WPR's 2026 analysis names it the strongest known predictor of AI citation at a correlation of 0.334, materially stronger than backlinks. You can read it in Search Console on Monday morning, which makes it the most useful number in this section. It also explains the concentration: the top 15 domains capture about 68 percent of all AI citations, and on AI Overviews specifically the top 1 percent of domains take 47 percent.

Owner: brand or PR, not the SEO team. Metric: brand search volume, plus third-party mentions on the domains that dominate your category.

What You Cannot Influence, and Why to Stop Trying

I have seen more wasted budget in this section than in the other two combined, and it usually comes from a good instinct: the team learns that training data matters and tries to influence it.

Training data is a lagging indicator. Your cutoff was set before you read this page. Public internet content is one of three stated inputs to OpenAI's models, and site owners can disallow the training crawler in robots.txt. That removes you from future training. It is a trade-off, not a win.

Licensed data is a procurement decision. The reported Reddit deals, around 60 million US dollars per year for Google and over 70 million for OpenAI, are not a marketing budget line. You cannot out-publish a licensing contract.

Live tool access belongs to product teams. Unless your company ships a connector and gets it into a directory, this layer is out of scope. The Clay timeline shows how fast it moves and the connector reliability problems show how early it still is.

Naming what you cannot influence is what makes the other two sections worth reading. A guide that promises influence over all four sources is selling you something.

How to Check Your Own Footprint in 30 Minutes

Four checks, cheapest first, so you have a verdict before you have spent your morning.

1. Ask the assistants your buyers actually use. Four or five prompts that describe your category, run in ChatGPT, Claude, Perplexity and Google. Log whether you appear and who does instead. Ten minutes, and usually the most uncomfortable part of the exercise.

2. Check which of your pages earn generative AI impressions. If your property has the Search Console generative AI report, sort by page. This is the only first-party version of this number that exists. Five minutes.

3. Check your brand search volume trend. The strongest known citation predictor, and it is already in your Search Console. Five minutes.

4. Check third-party presence on the domains that dominate your category. Start with the concentration data from the section above, then the subreddits and review sites your buyers read. Ten minutes.

The Measurement Problem

Here is the sentence that reframes everything above. You cannot manage any of this without tracking it, and until recently almost nothing tracked it.

That changed on 3 June 2026, when Google announced a dedicated generative AI performance report in Search Console, with separate views for Search and Discover. Global rollout was confirmed as of 31 August 2026. Google's note reads: "As of August 31, 2026, we've rolled out these insights to all websites worldwide." It reports impressions, pages, countries, devices and dates, with hourly to monthly granularity.

It does not report clicks, CTR, position or queries. There is also a companion opt-out toggle that removes a site from AI Overviews, AI Mode and AI Overviews in Discover without affecting indexing or ranking.

Report impressions without clicks and you have the proof of a structural change: a citation is now a separate event from a visit. Your brand can be named in an AI answer that produces zero sessions in your analytics, and that is the measurement problem in one line. If you want the tooling landscape for it, AI citation tracking tools and AI search monitoring cover the options, and LLM visibility covers what the metric actually means once you have it. We built AI visibility tracking into our own platform for the same reason: the manual version of this audit does not survive a busy quarter.

None of this replaces answer engine optimization as a discipline, and it does not replace SEO either. It adds a measurement layer that tells you whether the answers your buyers read are about you or about someone else. If you are starting that discipline from scratch, the AEO guide is the sensible entry point. Track this weekly, or you will be reading about a citation shift three months after it happened, the same way the 2025 Reddit rebalancing caught most teams off guard. LLM citations explains why answers ignore perfectly good content, which is the next question after this one.

The gap between the four sources is closing in one direction only: retrieval and connectors keep growing and training data keeps receding into the background. Which means the two sources you cannot touch are becoming less relevant every year, while the two that respond to your work are becoming the ones that decide the answer. If your team has never run the 30-minute audit above, this is a much better week for it than the week after your next competitor gets named in an AI Overview you should have owned.

Frequently Asked Questions

Where does AI get its information?

Four sources: training data the model absorbed before release, retrieval from a live search index at the moment you ask, licensed and proprietary datasets the vendor bought, and live tool access through connectors and APIs. Only retrieval and citation are reliably influenced by marketing work, and retrieval responds in weeks rather than years.

Does AI learn from the internet in real time?

No. Training happens in batches and freezes at a cutoff. What feels like real-time learning is retrieval: the model searches when you ask and writes an answer from what it found. That distinction matters because retrieval is the one layer your team can move quickly, and it is the layer that decides most answers today.

Does Reddit really feed AI answers?

Reddit is heavily cited, but be careful with the numbers. The 40.1 percent figure you see most often comes from a Reddit thread reposting a Semrush study from June 2025 that measured 150,000 citations across 5,000 keywords. That measures citation frequency, not training data. Other datasets put Reddit between roughly 5.5 and 21 percent of AI Overviews, and ChatGPT's reliance on it swung hard in both directions during late 2025.

Can I remove my content from AI training?

Not completely. You can block a training crawler in robots.txt, which prevents future collection from your site, and you can opt out of AI Overviews and AI Mode through the Search Console toggle without affecting your normal indexing or ranking. Neither removes information that third parties publish about you. Licensed agreements are not something a site owner can opt out of.

What is the difference between training data and retrieval?

Training data is what the model already knows and cannot update: it is a snapshot. Retrieval is what the model fetches right now to answer this specific question. Ahrefs' phrasing is still the clearest: training data, retrieval systems, and live tool access. If an assistant knows about last week's news, you are seeing retrieval. If it confidently states something that stopped being true two years ago, you are seeing training data.

Your competitors are already using AllAble. Are you?

The marketers pulling ahead aren't working harder. They're just working with one tool that does everything — that tool is AllAble. Try it yourself!