research

Where AI Gets Its Answers: 167,551 Citations Behind Brand Reputation

July 13, 2026
7 min read
n=167,551
Dmitrij Żatuchin
AI VisibilitySource AttributionCitationsGroundingWikipediaOpen DatasetGEO

When a grounded AI model answers a question about a brand, it cites its sources. We collected those citations at scale: 167,551 URL-grounded citations (189,974 attribution rows) across 128 brands, 13 languages, and 12 home markets, from GPT, Gemini, and Perplexity with live web search on.

The dataset merges three Rankfor studies into one citation corpus, resolves Google's grounding redirectors back to the real publishers, and tags every citation with its domain, source type, language, and model. It is open under CC-BY-4.0: doi.org/10.5281/zenodo.20829524.

Here is what it shows.

85.7% of your AI reputation is written by someone else

Across the full corpus, AI grounds brand answers in third-party sources 85.7% of the time. The brand's own website carries 14.3%.

Read that against how marketing budgets are split. Most brand teams spend most of their content money on the 14.3% they control and treat the 85.7% as PR's problem. The models treat it the other way around: when they need to describe you, they overwhelmingly quote pages you never wrote.

18% of domains carry 80% of the citations

Citation volume follows a power law: about 18% of domains account for 80% of all citations (Zipf alpha 0.86, R² 0.983). The grounding web is a small club. A brand that gets itself described correctly on a few dozen high-frequency domains has covered most of what AI will ever quote about it.

Which domains? Wikipedia is the most-cited domain in 11 of the 12 languages. The exception is Lithuanian, where the business daily vz.lt outranks it. If your Wikipedia presence is thin or wrong, it is thin or wrong in almost every language at once.

Poland runs on different rails

Poland is the corpus's outlier market. The top-cited domain for Polish brand answers is YouTube, and four HR and careers portals together out-cite Polish Wikipedia roughly two to one. AI describing a Polish brand quotes its job ads and employer-review pages more readily than its encyclopedia entry. We unpack that finding in the job-boards study.

The lesson generalizes: grounding profiles are market-specific. The domain list that controls your reputation in Germany is a different list in Poland, and nobody audits the second one.

The models cite differently too

Perplexity is the highest-volume citer in the corpus, quoting more sources per answer than GPT or Gemini. That makes it the easiest model to influence through breadth (be present on many domains) while the sparser citers reward depth (be the best source on the few domains they trust).

What to do with this

Three moves follow directly from the data:

  1. Audit the 85.7%. List the domains AI actually cites about your category (the dataset includes the per-language domain tables) and check what they say about you. Our Index brand pages show each measured brand's top citation domains.
  2. Fix Wikipedia and the local exception. One correct, well-sourced Wikipedia article covers 11 languages of grounding at once. Then find your market's vz.lt.
  3. Treat high-frequency domains as distribution partners. With 80% of citations flowing through 18% of domains, two or three placements on the right ones move more AI answers than fifty blog posts on your own site.

The full dataset, the analysis ledger, and a reproduction script are in the Zenodo record: 10.5281/zenodo.20829524. For the companion question, what AI knows about brands that it cannot cite at all, see our sourcing opacity study, where 70.2% of brand knowledge traced to no source.

Continue Reading

research

We Asked 3 AI Platforms About 24 Brands. They Could Not Agree on Anything.

The largest multi-industry study of AI reputation sourcing to date. 24 companies, 8 industries, 1,311 responses, 3,041 coded source attributions across GPT-5.2 and Gemini 3 Flash. Key finding: 70.2% of what AI "knows" about brands cannot be traced to any source. Models disagree on sentiment (0.83 vs 0.32), consistency averages just 0.54, and being publicly traded gives zero advantage. Published in Springer Discover Artificial Intelligence.

article

One in Every Nine Sources AI Uses to Describe Polish Brands Is a Job Board No Marketer Owns

We pulled 17,891 cited sources that AI uses to describe 46 Polish consumer brands across three frontier models with grounding on. After stripping brand-owned domains, four HR portals (livecareer.pl, pl.indeed.com, interviewme.pl, randstad.pl) carry 460 citations between them, 10.6% of the TOP 20 cross-cutting domains. That makes employer-reputation the third largest source category AI relies on for Polish brands, ahead of Wikipedia, Business Insider PL, and Forbes PL. None of these pages is owned by marketing or PR.

research

How We Measure AI Brand Reputation: The Method Behind the Rankfor Index 2026

The full method behind the Rankfor Index 2026: 66 brands, 12 languages, 3 AI models with live grounding, 35,640 responses, 131,667 citations. Five weighted components (sentiment, recommendation, source quality, consistency, stability) produce the 0-100 AI Visibility Score published on every brand page, plus the metrics only multilingual data can expose: bilingual penalty, citation concentration, and blind languages.

article

The "Best Of" Trap: Why Self-Promotional Listicles Are Losing Google and Winning ChatGPT

Three signals in one week: Google penalized self-promotional listicles (SaaS brands losing 34-49% visibility), reduced crawl limits by 86.7%, while Ahrefs found 44% of ChatGPT citations come from those exact listicles. Google is shrinking, AI is expanding. Includes a 5-minute brand audit playbook and split strategy for navigating both discovery systems.

BeVisible Club

Get research like this before we publish it.

New AI visibility studies, ranking patterns, and source-stack data delivered to your inbox a week before they go public.

Join BeVisible Club

Want to Know How AI Sees Your Brand?

Rankfor.AI measures your brand's AI visibility across all major platforms and provides actionable recommendations.

About the Author

Dmitrij Żatuchin

Founder

Dmitrij Żatuchin is the founder of Rankfor.AI. A computer scientist with a PhD in semantic web technologies, he bridges the gap between how AI reasons about brands and how brands want to be understood. With over two decades of software architecture experience and academic roles at Estonian Business School and EUAS, Dmitrij builds the measurement infrastructure brands need to transition from optimizing for search engines to becoming visible for reasoning engines.

© 2025-2026 Rankfor.AI™ All rights reserved.

Ask AI about Rankfor.AI

Rankfor.AI sp. z o.o., Skarbowcow 23B, 53-025 Wroclaw, Poland

KRS: 0001190083 | NIP: 8993033605

We use cookies

We use essential cookies to make our site work. With your consent, we may also use analytics cookies to understand how you use our tools (like the Dice Roller) so we can improve them. Learn more