Disclosure: This guide is published by Rankfor.AI. Founder Dmitrij Żatuchin is included, and Rankfor.AI sells software related to his research. His 2026 manuscripts link to their public repository records so readers can inspect them directly.
Last checked: 7 September 2026. Roles and publication records can change after this date.
The short answer
- Jason Barnard is the only person in this group with verified, relevant, peer-reviewed journal articles.
- Dmitrij Żatuchin has ten distinct current preprints. Malte Landwehr and Tomek Rudzki share two SSRN working papers.
- A citation, brand mention, recommendation, and visit are different outcomes. Combining them into one visibility score hides useful information.
- Repeated prompts, model choice, language, location, and retrieval mode can materially change the measured result.
- Third-party pages supply most cited brand evidence in several large studies, while most public GEO evidence remains proprietary, observational, or commercially connected.
Why these 15?
Inclusion required a current AI-search role plus at least one identifiable public study, technical investigation, framework, scholarly paper, or operating product. Coverage is limited to these 15 and leaves many academic GEO authors outside its scope.
Each person was assessed on four questions:
- Do they publish original measurements or system observations?
- Is the method visible enough to judge the scope?
- Are the data and limits stated?
- Is the work an industry study, a public preprint, or a peer-reviewed article?
Żatuchin receives a paper-by-paper section because his complete 2026 series is the primary case study in this review. The other profiles use representative public work. Commercial affiliation matters because almost everyone here sells connected software, consulting, or agency services. Open methods, open data, replication, and careful attribution help readers judge those conflicts.
Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande introduced Generative Engine Optimization as a formal research framework in November 2023. Their GEO paper was later published at ACM KDD 2024.
AI visibility researchers and practitioners at a glance
| Person | Follow for | Strongest public work type | Relevant scholarly status by 7 September 2026 |
|---|---|---|---|
| Jason Barnard | Entity resolution and AI brand representation | Applied journal frameworks | Two peer-reviewed journal articles |
| Dmitrij Żatuchin | Sampling, stability, multilingual measurement, and sources | Large empirical preprints | Nine arXiv preprints and one separate Research Square preprint |
| Malte Landwehr | Prompt panels and cited-listicle exposure | Empirical working papers and vendor data | Two shared SSRN working papers; no peer-reviewed version located |
| Tomek Rudzki | Query fan-out and source selection | Empirical working papers and vendor data | Two shared SSRN working papers; no peer-reviewed version located |
| Metehan Yeşilyurt | Retrieval, reranking, and citation traces | Technical investigations and vendor data | No attributable scholarly paper located |
| David Konitzny | ChatGPT retrieval and page-reading traces | Product-specific technical investigation | No attributable scholarly paper located |
| Lily Ray | Search quality, trust, citations, and recommendations | Practitioner experiments | No attributable scholarly GEO paper located |
| Kevin Indig | AI traffic, citations, and business impact | Multi-study industry synthesis | No attributable scholarly GEO paper located |
| Aleyda Solís | International and multilingual AI-search measurement | Modeled traffic and citation studies | No GEO paper located; editor-reviewed 2020 Web Almanac chapter |
| Cyrus Shepard | Citation-factor evidence synthesis | Expert narrative synthesis | No attributable scholarly GEO paper located |
| Wil Reynolds | AI-assisted buyer research and business measurement | Small user study and agency research | No attributable scholarly GEO paper located |
| Ross Simmonds | Third-party distribution and source ecosystems | Large company study | No attributable scholarly GEO paper located |
| Mike King | Relevance Engineering and agentic retrieval | Practitioner framework and lab research | No attributable scholarly GEO paper located |
| Evan Bailyn | Commercial GEO strategy and buying-intent research | Proprietary agency studies | No attributable scholarly GEO paper located |
| Ipek Isler (İpek İşler) | AI visibility product operations | Operating product | No attributable scholarly paper or transparent primary study located |
Evidence labels used in this guide
- Peer-reviewed: reviewed through the journal or conference's stated process.
- Preprint or working paper: public scholarly manuscript without verified peer review.
- Industry study: empirical work published by a practitioner or company.
- Technical investigation: observed traces, logs, or code from a specific product version.
- Framework or tool: an operating model or product without a methods-led validation study.
A DOI is a persistent identifier; peer review requires separate confirmation from a journal or conference. arXiv, Research Square, Zenodo, and SSRN host public research records under different processes.
Featured research program: Dmitrij Żatuchin's ten 2026 studies
Current role: Founder of Rankfor.AI and researcher affiliated with the Estonian Entrepreneurship University of Applied Sciences.
Best for: Sampling design, reliability, multilingual measurement, and source coverage in AI brand research.
Strongest public evidence: Ten linked 2026 manuscripts spanning brand recommendations, sourcing, multilingual measurement, sampling, reliability, and discovery.
Żatuchin's 2026 research focuses on a measurement problem: an "AI visibility score" changes when the model, language, location, prompt, retrieval mode, repetition count, or brand-extraction rule changes.
That is a wider program than reverse-engineering one answer engine. It covers competitive ownership, citation sources, multilingual measurement, reliability, individual experts, buyer-prompt sampling, and contextual bias.
1. Gender framing changes the category and brand mix
Gender Bias in Large Language Model Brand Recommendations tested 1,279 public-version queries across Christmas and Valentine buying contexts. Female-framed Christmas prompts returned fewer brands, while the direction changed in the Valentine study on Gemini. Cross-model agreement on the actual brands ranged from zero to 0.68.
Plain language: answers to "for him" and "for her" questions concentrated in different product categories and brand pools. The pattern changes by occasion and model, so a single universal gender-bias rule would be misleading. This is a Research Square preprint, and no accepted peer-reviewed version was verified.
2. Most AI recommendation categories remain contested
Who Owns the AI Recommendation? analyzed 3,750 answers covering 50 brands, five industries, 250 brand-free buying questions, three models, and five repetitions. Recommendation concentration was moderate, only 8% of queries were competitive vacuums, and the three models agreed on the top brand 41.6% of the time.
Plain language: most categories have several plausible winners. A brand that leads in one engine may lose in another, and the strongest opportunities often sit in questions where no stable leader exists.
3. English-only monitoring misses local-market visibility
The Language Blind Spot collected 35,640 grounded answers for 66 brands across 11 markets, 12 languages, and three models. Local-language questions raised the recommendation share of local champions much more than that of global multinationals. Model choice still had a larger effect on overall stability.
Plain language: an English dashboard can make a local brand look weaker than it is. Market-language prompts reveal competitors and reputation patterns that English prompts leave out.
4. Other websites write most of the AI story about a brand
How Large Language Models Source Brand Reputation Across Languages and Markets examined 167,551 URL-grounded citations across 128 brands. The merged corpus contained 13 languages; the main comparison covered 12, and Wikipedia led in 11 of those 12. Third-party pages supplied 85.7% of citations, while brand-owned pages supplied 14.3%. About 18% of cited domains supplied 80% of all citations. Poland differed, with YouTube and careers sources playing a larger role.
Plain language: in these grounded answers, company-owned pages supplied a small share of the observed citations. A focused set of outside publishers, directories, communities, and local sources appeared much more often.
5. More prompt repeats eventually buy very little reliability
Where Does the Noise Come From? decomposed variation across 12,933 responses for 20 Central and Eastern European brands, eight languages, and three models. Query language explained 26.5% of the variance in a single response, compared with 1.5% for brand identity. Adding a block of five repetitions, from five to ten, reduced modeled relative-error variance by only 0.00030.
Plain language: teams gain more information by adding languages and models than by asking one identical question many more times. The measured outcome was multilingual sentiment polarity, so the exact variance shares are outcome-specific and provisional.
6. Fixed expert lists miss almost everyone an AI names
Who Gets Named used 2,400 grounded calls across 120 buyer questions, four models, four markets, five languages, and five repetitions. An individual professional appeared in 25.8% of answers. This is a lower-bound estimate because the high-precision detector recovered 61.7% of names in its validation sample. Category and model changed the rate sharply. A fixed list of 939 people matched only 0.47% of 27,293 name-shaped mentions, and only 26 people from that list appeared at all.
Plain language: a tracker built from a company roster sees a tiny slice of the people that AI may recommend. Open discovery is necessary when the research question is "who gets named?"
7. Buyer-prompt research needs a demand sample
Demand-Side Measurement for GEO, coauthored with Daniil Dzemesjuk, introduces PersonaGen-1M: 1,031,732 synthetic buyer personas, 511 industry labels, four market contexts, 19.4 million structured attributes, and 5.16 million attached queries. Each persona carries an intent label and preferred source types.
Plain language: analysts often invent the questions used in an AI visibility audit. A large, structured prompt corpus provides a repeatable sampling frame. The personas are synthetic, and the dataset enables later tests; real buyer behavior remains untested.
8. Language opens the local shortlist; location selects the local market
The Language of the Question Selects the Market ran 234 controlled tests through the logged-out ChatGPT interface and the OpenAI API across four exit countries and six languages. When language matched country, a global brand led only one of 24 runs. English questions on the same connections produced no local brands in Estonia or Türkiye. Holding language constant while changing the exit IP changed which country's suppliers appeared.
Plain language: question language helps decide whether local firms enter the answer. Location helps decide which local market the system searches. The study covers one service, a short collection window, and a small category set, so the effect should be retested elsewhere.
9. The right repeat count must be estimated
The Dice Roll Method reanalyzed about 190,000 observations across more than 270 brands. Across three independent corpora, 37 of 39 preregistered D-study prediction cells replicated, two replicated partially, and none failed. In the original data, the author labels five repeats exploratory and ten confirmatory, although the corresponding reliability coefficient at ten repeats remained below 0.80. Fifteen crossed the higher reliability target. The fixed tiers did not transfer cleanly to every outside dataset.
Plain language: one answer is one draw from a changing distribution. Run a small sample, estimate the variation, and calculate how many repeats the decision requires. Five, ten, and fifteen are context-specific planning anchors.
10. Retrieval-enabled brand lists narrowed sooner while sources stayed open
Repeated Queries Exhaust an LLM's Brand Recommendations but Not Its Sources collected 4,500 responses across 50 questions, six engines, and 15 runs per question-engine cell, then used open extraction to identify 1,470 organizations. Five engines without web search were still adding unseen organizations in 86% to 92% of cells at run 15. The sole retrieval-enabled engine produced a smaller brand repertoire and reached the study's estimated completeness threshold sooner, although 64% of its cells still added a new organization on run 15. Cited domains continued to accumulate.
Plain language: in this sample, the retrieval-enabled engine narrowed its brand list sooner while its source list kept growing. Estimated completeness used a Chao2 lower bound, leaving absolute exhaustion unproven. Because only one engine in the main comparison used retrieval, engine identity and retrieval mode remain confounded. A fixed roster can create a false plateau by ignoring every new name outside the list.
What the ten studies say together
Practical implication: Żatuchin's recurring result is simple: AI visibility is a distribution of possible answers. Measurement needs multiple engines, languages, relevant locations, repeated runs, clear outcome definitions, and open extraction where the candidate set is unknown.
Limit or commercial interest: Rankfor.AI sells measurement software related to this research. The manuscripts are preprints, and several use proprietary collection systems or synthetic data.
Scholarly record: these ten distinct current manuscripts are preprints. No publisher version-of-record or journal DOI for the current AI-brand manuscripts was verified. An older Research Square version of the Dice Roll Method is an additional repository record within the same paper lineage and counts within that single work. The current arXiv record appears to cite the DOI of the separate gender-bias paper in its supersession note; the earlier Dice Roll DOI is 10.21203/rs.3.rs-8980233/v1. Żatuchin has peer-reviewed publications outside this topic, including a 2024 Discover Education article.
Peer-reviewed applied research
Jason Barnard
Current role: Founder of Kalicube.
Best for: Entity resolution, knowledge graphs, and the way AI systems represent people and companies.
Barnard's core practice framework separates understandability, credibility, and deliverability.
Strongest public evidence: Barnard has two verified peer-reviewed applied papers:
- With Matt Artz, Search Marketing in the Age of AI, Journal of Digital & Social Media Marketing, 11(3), 244 to 260, 2023. It proposes a marketing framework for search, assistive, and answer engines. The publisher says every article is peer-reviewed, and the Aalborg University record classifies it as a peer-reviewed journal article.
- Engineering the AI Résumé, Journal of AI, Robotics & Workplace Automation, 4(4), 338 to 355, 2026. It defines a framework for evaluating how AI represents an entity, includes falsifiable predictions, and gives a minimal replication protocol. The official journal volume lists the article, and the journal describes its anonymous peer-review process.
Practical implication: clear identity, corroborated claims, and accessible evidence make it easier for machines to resolve who a company is and what claims belong to it.
Limit or commercial interest: these are applied, practitioner-derived framework papers. Controlled retrieval experiments and independent validation of Kalicube's commercial method remain outside their scope.
Scholarly record: two verified peer-reviewed applied journal articles. Barnard also maintains a larger register of Zenodo and SSRN working papers; repository records with DOIs still require a separate peer-review label.
Follow: Publications · Kalicube
Researchers with scholarly GEO working papers
Malte Landwehr
Current role: CPO and CMO at Peec AI.
Best for: Connecting product-scale answer-engine data with prompt panels and marketing questions that teams can test.
Strongest public evidence: Landwehr coauthored two 2026 SSRN papers with Jan Ehrlinspiel and Tomek Rudzki. Prompt Tracking Works, But Not as a One-Prompt Measurement System combines 288 human-written prompts with 1,466 controlled prompts: 54 base prompts and 1,412 variants. It finds that prompt format, meaning, funnel stage, and engine shift brand-visibility baselines.
Cited-Listicle Rank-Tier Exposure, Author Type, and LLM Brand Visibility uses about 5.7 million brand-level observations across B2B software, emerging marketing technology, and US finance. Higher placement in frequently cited third-party listicles was associated with a higher chance of being named and earlier placement in the answer.
Practical implication: track a portfolio of prompts across distinct buyer needs, and treat frequently cited third-party listicles as measurable distribution surfaces.
Limit or commercial interest: the listicle study is observational, leaving the causal effect of moving a brand up one list untested. All three authors declare Peec AI commercial interests.
Scholarly record: two relevant SSRN working papers. SSRN's publication policy provides no platform peer review.
Tomek Rudzki
Current role: GEO researcher at Peec AI with a technical SEO and research-and-development background.
Best for: Large-scale query fan-out and engine-specific source analysis.
Rudzki coauthored the same two SSRN working papers with Landwehr and Ehrlinspiel.
Strongest public evidence: Rudzki's industry studies map retrieval at much larger scale. A five-million fan-out analysis found different research breadth by engine: Perplexity averaged 1.4 fan-outs, ChatGPT 2.1, and Grok 6.8 in the observed April 2026 sample. Commercial modifiers such as "best," comparisons, reviews, tools, and features appeared frequently. A separate 30-million-source study showed that source preferences differ across engines.
Practical implication: the user's visible question is often the start of a larger hidden search process. Brands need coverage across the related questions an engine may generate, and each engine deserves its own measurement.
Limit or commercial interest: the large retrieval datasets come from Peec AI, where Rudzki works, and source behavior can change with product updates.
Scholarly record: two relevant shared SSRN working papers; no peer-reviewed GEO version was located.
Technical and industry researchers
Metehan Yeşilyurt
Current role: GEO researcher at Peec AI and cofounder of AEO Vision.
Best for: Reverse-engineering fan-out, reranking, retrieval, and citation selection from product traces and large vendor datasets.
Strongest public evidence: Yeşilyurt's largest public study analyzed 64.77 million Reddit citations across 20 countries from ChatGPT, Google AI Overviews, Google AI Mode, and Gemini. These systems often cited machine-translated English Reddit pages in non-English markets, while ChatGPT's share changed sharply during the collection period. In a separate Perplexity check, no ?tl= citation was found. The result shows engine-specific behavior and fast drift.
His Reciprocal Rank Fusion investigation found RRF-related code and parameters in client traces. It supports a useful hypothesis: consistent appearance across several fan-out searches may beat one isolated first-place result.
Practical implication: measure whether a brand appears across related retrieval queries and engines. One isolated rank reveals little about the full answer process.
Limit or commercial interest: the largest dataset comes from Peec AI, where Yeşilyurt works. The code investigation is a product-version snapshot with no visibility into a permanent ranking formula.
Scholarly record: no attributable peer-reviewed paper, arXiv paper, or SSRN preprint was located. His public work is technical industry research.
Follow: metehan.ai · LinkedIn
David Konitzny
Current role: GEO researcher at Peec AI, following earlier technical and in-house SEO work.
Best for: Observing what ChatGPT Deep Research searches, reads, and rereads during a task.
Strongest public evidence: How ChatGPT Deep Research Reads Your Site recorded WebSocket activity from more than ten free accounts. The study observed a search, open, and find sequence; an initial reading window around 5,700 characters; keyword-targeted rereading; and much less main content in the initial window on navigation-heavy pages. It also found that being read and being cited are separate events.
Practical implication: page order and clutter can affect what this research agent sees before it chooses sources.
Limit or commercial interest: this is a June 2026 snapshot from more than ten free accounts and about 20 requests per account. Its scope covers ChatGPT Deep Research, excluding standard ChatGPT Search, Agent mode, Gemini, and Perplexity. Konitzny works for Peec AI.
Scholarly record: no attributable relevant scholarly paper or preprint was located.
Follow: LinkedIn · Peec AI research
Lily Ray
Current role: Founder of Algorythmic and also listed as VP of SEO and AI Search at Amsive.
Best for: Search quality, trust signals, organic visibility, and the gap between an AI citation and a brand recommendation.
Strongest public evidence: In Why Calling Yourself the "Best" Could Be Helping Your Competitors Win in AI Search, Ray tested 100 B2B "best category" questions on three dates. Eighty questions produced a Google AI Overview. Among 323 instances where an Overview cited a company's self-promotional listicle, the publisher's brand was absent from the recommendation in 224 cases, or 69%.
Her core-update and AI-citation study followed 11 selected subfolders and found that losses in Google visibility coincided with an average 22.5% fall in AI citations.
Practical implication: a company can supply the page that an AI cites while the answer recommends its competitors. Citation share and recommendation share measure different jobs.
Limit or commercial interest: the 11-subfolder sample is small and selected, so it supports a shared-signal hypothesis without proving a causal link. Ray sells consulting connected with the subject.
Scholarly record: no attributable relevant scholarly paper or preprint was located.
Follow: Lily Ray's newsletter · LinkedIn
Kevin Indig
Current role: Independent growth adviser and publisher of Growth Memo.
Best for: Connecting AI citations, search strength, traffic patterns, and business impact.
Strongest public evidence: Indig's Winners and Losers of AI Search synthesis draws on 20 months of work, more than one million AI answers, and over 100,000 citations. It reports that the first organic result received about four times as many AI citations as result ten, while repeating one prompt produced very low citation persistence.
His Ghost Citations study with Semrush examined 3,981 domain appearances from 115 prompts across 14 countries. It found that 61.7% were citations without a brand mention and 25.1% were brand mentions without a citation.
Practical implication: search strength still helps, while a citation, a mention, a recommendation, and a click remain four different outcomes. Reporting them as one score hides useful information.
Limit or commercial interest: the million-answer talk combines proprietary and partner studies. The Ghost Citations percentages come from 115 prompts and apply to that sample.
Scholarly record: no attributable peer-reviewed GEO paper or scholarly preprint was located.
Follow: Kevin Indig · LinkedIn
Cyrus Shepard
Current role: Founder of Zyppy and former Moz SEO lead.
Best for: Comparing many public citation studies and identifying claims that have support across several sources.
Strongest public evidence: AI Citation Ranking Factors Analysis reviews 54 experiments, studies, patents, and case studies, then manually scores 23 suspected citation correlates. The highest-rated signals include URL accessibility, ordinary search rank, fan-out rank, preview permissions, direct query-answer fit, appropriate content format, and placing the answer near the start of the page. The analysis finds little credible evidence for [llms.txt](/glossary/term/llms-txt) as a citation lever.
Practical implication: pages that can be crawled, found, understood, and quoted cleanly tend to appear more often in AI citations. Shepard explicitly calls the items correlates and rejects the label of confirmed ranking factors.
Limit or commercial interest: the manual scoring is an expert narrative synthesis, outside the category of a registered systematic review or meta-analysis. Shepard sells SEO products and consulting.
Scholarly record: no attributable scholarly GEO paper or preprint was located.
Follow: Zyppy Signal · Zyppy
Aleyda Solís
Current role: International SEO and AI search consultant, founder of Orainti, and cofounder of Finchling.
Best for: Cross-market measurement, local search platforms, and turning AI presence into business reporting.
Strongest public evidence: Solís's global AI search study uses Similarweb estimates for 87.6 million click-producing AI referrals and 57,696 domain-market entries across ecommerce, finance, and travel in ten markets. It shows that local platforms can outperform global defaults in their home countries.
An observational 15-ecosystem SaaS study combines Semrush and Similarweb estimates. It found that third-party pages supplied about 84% to 93% of citation weight in ChatGPT and AI Mode. The pages used as evidence were often different from the pages that received the visit.
Practical implication: international GEO needs local prompts and local source maps. The page that supplies evidence and the page that receives the customer can be different assets.
Limit or commercial interest: the traffic study relies on modeled Similarweb visits, and the SaaS study combines third-party estimates. Both are observational. Solís sells consulting connected with the findings.
Scholarly record: no attributable scholarly LLM-search or GEO paper was located. Solís coauthored the data-based 2020 Web Almanac SEO chapter, an editor-reviewed community research report outside conventional journal peer review.
Follow: Aleyda Solís · LinkedIn
Wil Reynolds
Current role: Founder and CEO of Seer Interactive.
Best for: Connecting answer-engine measurements with customer behavior and business results.
Strongest public evidence: Reynolds and Andrea Haley observed 28 people completing 84 AI-assisted research sessions. The report found large differences between pre-task consideration and post-task intended choice. The small convenience sample shows why customer research needs outcomes beyond citation counts.
Seer also publishes larger studies through named research staff. Nick Haigler's query fan-out study ran 100 prompts 13 times and collected 11,029 unique fan-out queries. Only eight appeared in every run. Credit for that work belongs to Haigler and Seer.
Practical implication: track recurring needs and themes, then connect them to consideration, direct visits, signups, or sales. Exact hidden queries change too often to act like a keyword rank report.
Limit or commercial interest: the user study uses a small convenience sample and non-equivalent pre-task and post-task measures. The larger fan-out work belongs to named Seer researchers. Reynolds owns an agency that sells related services.
Scholarly record: no attributable peer-reviewed GEO paper or scholarly preprint was located.
Follow: LinkedIn · Seer research
Ross Simmonds
Current role: Founder of Foundation Marketing and creator of Distribution.ai.
Best for: Distribution across the publishers, communities, videos, and professional networks that answer engines use.
Strongest public evidence: The Foundation and AirOps team reports in its Hidden Selection Phase study that it analyzed 5.1 million responses and 57.2 million citations for 50 B2B brands across seven verticals and five AI platforms. Brand-owned pages supplied 10.15% of citations overall and 2.2% for unbranded discovery prompts. Reddit supplied about 20.8% of outside citations, YouTube 13%, and LinkedIn 11%.
Practical implication: AI systems usually meet a brand through somebody else's content. Distribution, public proof, and third-party description belong inside an AI visibility program.
Limit or commercial interest: the report is company research by the Foundation and AirOps teams. The raw data and full processing method remain closed, and both organizations have commercial interests related to the findings.
Scholarly record: no attributable scholarly GEO paper or preprint was located.
Follow: Foundation Lab · Ross Simmonds
Framework authors and product builders
Mike King
Current role: Founder and CEO of iPullRank and principal advocate of Relevance Engineering.
Best for: Understanding AI search through information retrieval, technical access, entity resolution, and passage selection.
King's AI Search Manual connects information retrieval, technical access, entity understanding, content, digital PR, and measurement into one operating model.
His Agentic RAG work explains why many answer systems plan, search, inspect, and revise through several steps.
Strongest public evidence: iPullRank Lab analyzed 79,000 URL-query pairs across four platforms. It found correlations between citations and search rank, topical match, entity richness, and concise coverage. Francine Monahan authored the published analysis, so the study should be credited to Monahan and iPullRank Lab within King's broader framework.
Practical implication: diagnose discovery, retrieval, passage selection, and synthesis separately. A page can fail at any one of those stages before the user sees an answer.
Limit or commercial interest: the 79,000-pair analysis is correlational company research authored by Francine Monahan. King and iPullRank sell related agency services.
Scholarly record: no attributable peer-reviewed GEO paper or scholarly preprint was located. The manual, essays, presentations, code, and lab work are practitioner publications.
Follow: AI Search Manual · LinkedIn
Evan Bailyn
Current role: Founder of First Page Sage.
Best for: Commercial GEO frameworks, buying-intent questions, and third-party evidence associated with recommendations.
Bailyn helped package GEO as an agency practice and publishes proprietary studies about commercial recommendations, citations, and traffic.
Bailyn anticipated a related agency service in May 2023, using the names Generative AI Optimization and AI Optimization. Aggarwal and colleagues formalized GEO as a research framework in November 2023. First Page Sage's GEO-titled page states an original publication date of 13 March 2024.
Strongest public evidence: First Page Sage says its GEO recommendation-factor study covers 11,128 commercial questions across four chatbots. A separate buying-intent study reports 36,140 questions across four models and 17 industries. First Page Sage's proprietary framework associates authoritative third-party lists, review platforms, accreditations, customer evidence, and public sentiment with recommendations.
Practical implication: Bailyn treats GEO as a commercial demand channel and asks which outside evidence is associated with entry into a buyer's AI-generated shortlist.
Limit or commercial interest: the raw data, coding protocol, and full statistical model remain closed. First Page Sage also publishes rankings of GEO experts that include its own founder. We excluded those marketing pages from evidence of research standing.
Scholarly record: no attributable peer-reviewed GEO paper or scholarly preprint was located.
Follow: First Page Sage research · GEO study
Ipek Isler (İpek İşler)
Current role: Cofounder and CEO of AEO Vision. Her name appears as Ipek Isler in many English-language profiles.
Best for: Operating an AI visibility product that tracks prompt-level brand mentions, citations, sentiment, competitors, and changes over time.
Strongest public evidence: AEO Vision is an operating product that turns answer-engine observations into a visibility intelligence layer for marketing teams.
Practical implication: compare mentions, citations, sentiment, and competitors as separate prompt-level measures over time.
Limit or commercial interest: AEO Vision's own "best experts" page ranks Yeşilyurt first and İşler second. That is circular promotional evidence and was excluded from the evidence used to establish research standing. The company's technical research is more directly attributable to Yeşilyurt.
Scholarly record: no attributable peer-reviewed paper, scholarly preprint, or transparent methods-led primary study was located. "AI visibility product builder and operator" is the accurate current label.
Follow: AEO Vision
What this group has learned in plain language
One answer is too little for a market claim
Generated answers change from run to run. Five repetitions can support exploration in some settings, yet the required sample depends on the model, prompt, outcome, and size of the difference a team needs to detect.
Citation, mention, recommendation, and visit are separate events
A model can cite a page without naming its publisher. It can name a brand without linking to it. It can use one page as evidence and send the user to another. Each event needs its own metric.
Every engine is a separate channel
ChatGPT, Gemini, Perplexity, Grok, Claude, and Google AI surfaces use different search breadth, source mixes, and answer styles. An average across engines can hide the exact channel where a brand wins or disappears.
Language and location change the competitive set
Translated English prompts produce different results from local-language buyer questions. Query in the market language, capture the location, and keep those dimensions visible in reporting.
Third-party sources carry most of the brand story
Several independent and company datasets converge on the same broad direction: answer engines often rely far more on outside pages than on the brand's own domain. The useful source map includes publishers, review sites, communities, video platforms, directories, and local-market authorities.
Hidden retrieval widens the real query set
Answer engines often generate related searches behind the visible prompt. Strong coverage across the buyer's topic and adjacent questions improves the chance that at least one useful page enters retrieval.
Visibility needs a business outcome
Share of voice can show movement. Marketing still needs to connect that movement to corrected brand facts, shortlist inclusion, branded demand, direct visits, qualified leads, or revenue. A visibility chart without a decision attached will eventually lose executive support.
How to use this list
Following researchers is the starting point. Extract a testable claim, check whether its evidence matches your engine and market, reproduce it on your own prompt sample, and track the result over time. For a defensible AI visibility program:
- Sample real buyer needs across the funnel and market. Prompts chosen in a conference room are too narrow.
- Run each prompt several times on each relevant engine.
- Keep language, location, model, date, retrieval mode, and account state in the record.
- Report mentions, recommendations, citations, cited domains, sentiment, and visits as separate measures.
- Use open brand extraction when the full competitor set is unknown.
- Audit the outside sources that answer engines repeatedly use for the category.
- Link visibility changes to a customer or commercial outcome.
Frequently asked questions
Which AI visibility researchers here have peer-reviewed papers?
Jason Barnard is the only person in this 15-person review with verified, directly relevant, peer-reviewed journal articles. His two papers are applied practitioner frameworks. They are separate from controlled information-retrieval experiments.
Which GEO researchers publish preprints or working papers?
Dmitrij Żatuchin has nine current arXiv preprints and one separate Research Square preprint. Malte Landwehr and Tomek Rudzki coauthored two SSRN working papers with Jan Ehrlinspiel. No peer-reviewed version of those manuscripts was verified by the cutoff date.
What is the difference between a peer-reviewed paper, a preprint, and a vendor study?
A peer-reviewed paper has passed the review process stated by a journal or conference. A preprint is a public manuscript outside a verified journal or conference review record. A vendor study can contain strong empirical work, although access to its raw data, sampling frame, or full method is often limited. A DOI is a persistent identifier; peer-review status requires separate confirmation.
Who researches prompt stability and repeated sampling?
Żatuchin's Dice Roll and saturation studies focus directly on repetition counts, reliability, and unseen brands. Landwehr, Rudzki, and Ehrlinspiel study prompt portfolios and the effect of wording, meaning, funnel stage, and engine choice on visibility baselines.
Who studies query fan-out and answer-engine retrieval?
Rudzki publishes large query fan-out datasets. Yeşilyurt investigates reranking and retrieval traces. Konitzny observes ChatGPT's search and page-reading behavior. King explains the wider retrieval pipeline through Relevance Engineering and Agentic RAG.
Who studies multilingual and international AI visibility?
Żatuchin studies language and location through controlled and multi-market datasets. Solís focuses on international traffic, local AI platforms, and cross-market source behavior. Both show that English-only monitoring misses important local variation.
Are AEO and GEO the same thing?
They are overlapping industry labels without a standardized boundary. AEO often refers to appearing as a direct answer. The original GEO paper defines methods for improving visibility in generative-engine responses. This guide uses AI visibility as the broader measurement category.
How were the people in this guide selected?
The guide uses the inclusion criteria in "Why these 15?" and groups people by evidence type. Many academic authors working on generative retrieval remain outside its scope.
Search method and limits
We checked primary author pages, original study pages, arXiv, Research Square, SSRN, DOI records, publisher pages, university repositories, and bibliographic indexes. Searches used name variants and combinations of each name with GEO, generative engine optimization, AI search, answer engine, LLM visibility, retrieval, and citation.
The field moves faster than journal publishing. Product behavior observed in June may change by September. Most large datasets are proprietary, full reproduction is unavailable for many vendor studies, and publication databases can miss obscure or newly indexed work.
"No scholarly paper located" is the result of a documented search through the cutoff date. An obscure or unindexed record may still exist. New journal versions should be added only after the publisher page or proceedings record confirms publication and peer-review status.
The strongest reading habit in GEO is simple: check who collected the data, what exactly was measured, whether the result is causal or correlational, and which evidence label belongs beside the claim.
Test the findings on your own brand
Several studies in this guide reach the same measurement lesson: one prompt is one observation. Use Answer Trail to compare AI answers to the same buyer question, the brands they name, and the sources they use.
Answer Trail is built by Rankfor.AI, the publisher of this guide.
BeVisible Club
Get research like this before we publish it.
New AI visibility studies, ranking patterns, and source-stack data delivered to your inbox a week before they go public.
Join BeVisible ClubWant to Know How AI Sees Your Brand?
Rankfor.AI measures your brand's AI visibility across all major platforms and provides actionable recommendations.
About the Author

Founder & CEO
Dmitrij Żatuchin is the founder of Rankfor.AI. A computer scientist with a PhD in semantic web technologies, he bridges the gap between how AI reasons about brands and how brands want to be understood. With over two decades of software architecture experience and academic roles at Estonian Business School, Dmitrij builds the measurement infrastructure brands need to transition from optimizing for search engines to becoming visible for reasoning engines.
