When a grounded AI model answers a question about a brand, it cites its sources. We collected those citations at scale: 167,551 URL-grounded citations (189,974 attribution rows) across 128 brands, 13 languages, and 12 home markets, from GPT, Gemini, and Perplexity with live web search on.
The dataset merges three Rankfor studies into one citation corpus, resolves Google's grounding redirectors back to the real publishers, and tags every citation with its domain, source type, language, and model. It is open under CC-BY-4.0: doi.org/10.5281/zenodo.20829524.
Here is what it shows.
85.7% of your AI reputation is written by someone else
Across the full corpus, AI grounds brand answers in third-party sources 85.7% of the time. The brand's own website carries 14.3%.
Read that against how marketing budgets are split. Most brand teams spend most of their content money on the 14.3% they control and treat the 85.7% as PR's problem. The models treat it the other way around: when they need to describe you, they overwhelmingly quote pages you never wrote.
18% of domains carry 80% of the citations
Citation volume follows a power law: about 18% of domains account for 80% of all citations (Zipf alpha 0.86, R² 0.983). The grounding web is a small club. A brand that gets itself described correctly on a few dozen high-frequency domains has covered most of what AI will ever quote about it.
Which domains? Wikipedia is the most-cited domain in 11 of the 12 languages. The exception is Lithuanian, where the business daily vz.lt outranks it. If your Wikipedia presence is thin or wrong, it is thin or wrong in almost every language at once.
Poland runs on different rails
Poland is the corpus's outlier market. The top-cited domain for Polish brand answers is YouTube, and four HR and careers portals together out-cite Polish Wikipedia roughly two to one. AI describing a Polish brand quotes its job ads and employer-review pages more readily than its encyclopedia entry. We unpack that finding in the job-boards study.
The lesson generalizes: grounding profiles are market-specific. The domain list that controls your reputation in Germany is a different list in Poland, and nobody audits the second one.
The models cite differently too
Perplexity is the highest-volume citer in the corpus, quoting more sources per answer than GPT or Gemini. That makes it the easiest model to influence through breadth (be present on many domains) while the sparser citers reward depth (be the best source on the few domains they trust).
What to do with this
Three moves follow directly from the data:
- Audit the 85.7%. List the domains AI actually cites about your category (the dataset includes the per-language domain tables) and check what they say about you. Our Index brand pages show each measured brand's top citation domains.
- Fix Wikipedia and the local exception. One correct, well-sourced Wikipedia article covers 11 languages of grounding at once. Then find your market's vz.lt.
- Treat high-frequency domains as distribution partners. With 80% of citations flowing through 18% of domains, two or three placements on the right ones move more AI answers than fifty blog posts on your own site.
The full dataset, the analysis ledger, and a reproduction script are in the Zenodo record: 10.5281/zenodo.20829524. For the companion question, what AI knows about brands that it cannot cite at all, see our sourcing opacity study, where 70.2% of brand knowledge traced to no source.
