AI NewsGo-to-marketReported
Chris Green mapped the data sources feeding AI answers, and tagged each one by how strong the evidence is
Chris Green published a tiered map of the data sources feeding AI answers on Search Engine Journal on 22 September, marking each as confirmed and live, confirmed and historical, or strong evidence with no confirmed deal.
Why it mattersA GEO plan that spends the same effort on Wikipedia, Yelp and a package registry is spending on three sources with different confidence levels, so the tiering tells a team where a placement is likely to produce citations and where it is a guess.
Getting your brand named in an AI answer starts with knowing who is feeding the model. Chris Green, senior consultant at Torque Partnership, published a tiered map of the data sources on Search Engine Journal on 22 September 2026, first posted on his own Substack the same day. Each source carries an evidence label: tier 1 for confirmed retrieval or grounding that is live today, tier 2 for confirmed training or licensing, tier 3 for confirmed historical pretraining, and tier 4 for strong evidence with no confirmed deal.
What the tiers actually say
Green places Google Search, Bing Search, Google Maps, Google Business Profile, Yelp, Wikipedia, Wikimedia, the Google Merchant Center feed, Google Hotel Center feeds and live publisher pages in tier 1. The Google grounding docs, Yelp's own 10-Q filing and OpenAI's help centre are cited for each. Tier 2 covers training-and-licensing deals: OpenAI's partnerships with the Financial Times, Axel Springer, the Associated Press and News Corp; Google's reported 60-million-dollar-a-year deal with Reddit; GitHub and Stack Overflow as named licensable data platforms. Tier 3 covers historical pretraining corpora: Common Crawl, cleaned derivatives such as C4, and archived news collections.
The named numbers
Green cites GPT-3's disclosed sampling mixture at 60 percent Common Crawl, LLaMA 1 at 67 percent, and LLaMA's C4 share at 15 percent, from the original papers. Merchant feeds for OpenAI shopping refresh as often as every 15 minutes, from OpenAI's own developer documentation. Google added hotel booking inside AI Mode with Google Pay in August 2026, from TechCrunch. He notes that Reddit is reportedly weighing whether to renew the Google deal, from Fortune and Neowin, and marks the training side of the Reddit entry as unstable for that reason.
Where Green flags weak evidence and one entry to cut
Tier 4 is the interesting one, because it names sources Green thinks are likely feeding the models but where he could not find a public deal: Foursquare, Tripadvisor, OpenStreetMap, Microsoft Merchant Center, marketplace feeds beyond Shopify, and reservation and inventory APIs. He labels npm and PyPI as the weakest entry in the whole table, with no disclosed agreement or documented retrieval use, and writes that the entry is a candidate to cut. He also removed Google Shopping from an earlier draft, on the grounds that AI writes to that surface and reads its underlying data from the Merchant Center feed instead.
Two smaller signals worth naming
The Yelp entry in tier 1 covers actions as well as citations. Green cites Yelp's own blog and 10-Q for a live product where ChatGPT users can book a Yelp table, join a waitlist, or send a Request-a-Quote to a provider. Wikipedia's licence sits differently from every other tier 1 source: it is principally CC BY-SA, which carries attribution and share-alike obligations on any downstream use.
Green states one caveat plainly. The table is a snapshot on 22 September 2026, and AI search deals are moving fast enough that specific entries will date. His advice is to work backwards from real customer questions, generate answers on ChatGPT, Google AI Mode, Perplexity and Copilot, and check which cited sources the brand is already inside and which it is not. That audit says where a placement is likely to produce citations, and where a tier 4 source is still a guess.
Source
Which Data Sources Should You Care About For AI Search?, by Chris Green, Search Engine Journal, 22 September 2026. Original: Which Data Sources Should You Care About For AI Search? on Chris Green SEO.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.



