15,000 buyer personas published on HuggingFace to validate what AI thinks your customers want—and why we didn't publish all 593K.
Why We Published This Dataset
Most buyer persona datasets are either synthetic noise or locked behind enterprise paywalls. We published PersonaGen-15K on HuggingFace to prove a different point: AI-generated personas aren't hallucinations—they're a reflection of what AI models learn from the web.
This is a 15,000-persona subset of our full 593,181-persona dataset, a historical 10% sample from the earlier 149K version, published in Springer's Discover Artificial Intelligence journal. It's stratified by industry, anonymized, and compressed into a single 8.9 MB Parquet file. The full dataset is the foundation of our commercial AI visibility intelligence platform. This subset is the proof.
If you're a researcher, data scientist, or marketing technologist asking "do AI-generated personas actually map to real buyer behavior?"—this is your validation set. If you're evaluating whether to license the full dataset for your product or research, this is your quality check.
You can download it here: PersonaGen-15K on HuggingFace.
What's Inside the Dataset
PersonaGen-15K contains 14,955 buyer personas across 25 industries, generated using Google's Gemini 2.0 Flash. Each persona includes:
- Demographics: Age range, education, job function, company size
- Goals: What they're trying to achieve (business objectives, personal outcomes)
- Pain points: Operational friction, budget constraints, capability gaps
- Search queries: The questions they ask AI assistants and search engines
- Uncovered needs: Implicit requirements not stated in searches but revealed through behavioral analysis
- Market context: Industry, competitive pressures, regulatory environment
- Purchase triggers: Events that shift them from research to evaluation
The dataset is stored in ZSTD-compressed Parquet format for efficient loading in pandas, Polars, or DuckDB. Industry distribution is preserved within 0.2% of the original dataset, meaning the stratified sample is statistically representative.
PersonaGen-15K: Dataset Composition
Anonymization: We hashed all persona names to Persona_{hash}, stripped narrative stories, and removed internal UUIDs. The dataset is research-grade, not PII.
Why 10%, Not 100%?
The full 593,181-persona dataset is our commercial product. It powers AI visibility intelligence for B2B marketing teams who need to understand how AI models recommend their brand—or don't.
We published 10% (15K personas) for three strategic reasons:
1. Validation, not substitution. 15,000 personas is enough to reproduce our research, validate our methodology, and build proof-of-concept tools. It's not enough to replace the commercial dataset. If you're training a recommendation model or building a persona search engine, you need the full 593K.
2. The free sample that proves the quality. Every researcher who downloads this dataset and cites it in their work is validating our methodology. Every data team that evaluates it for licensing is experiencing the depth of the schema. This is demand generation with zero customer acquisition cost.
3. Academic authority compounds. Publishing on HuggingFace with a citable DOI means other researchers reference this dataset. Those citations reinforce our position: if you're studying AI-generated personas, buyer intent modeling, or LLM recommendation behavior, this is the canonical dataset.
Full Dataset vs Open Subset
Compare that to a EUR 2,500 LinkedIn ad campaign. HuggingFace gives us higher-quality leads—researchers, data scientists, enterprise ML teams—at EUR 0 cost.
Who This Dataset Is For
Researchers studying AI behavior, persona generation, or buyer intent modeling. This is your training set for evaluating how well AI models capture real-world buyer archetypes.
Data scientists building recommendation systems, search relevance models, or content personalization engines. Use this to test whether your model can predict which personas care about which features.
Marketing technologists evaluating tools like HubSpot, Salesforce Marketing Cloud, or SEMrush. If your vendor claims "AI-powered buyer insights," you can now benchmark their outputs against a published dataset.
Product teams at martech or sales intelligence platforms. If you're considering licensing the full dataset for OEM integration (EUR 50K-200K/year), this is your proof-of-value.
Academic labs researching fairness, bias, or persona diversity in AI-generated datasets. The full paper includes bias analysis across gender, age, and industry representation.
What You Can Build With It
We built three open tools using this dataset. You can too.
1. Persona Matcher Upload your website content or product descriptions. The tool matches you to the most resonant personas in the dataset, ranked by semantic similarity. Think of it as "find my audience" based on what AI thinks your content signals.
2. Benchmark Calculator Compare your persona coverage against industry peers. If you're a B2B SaaS company, how does your audience diversity compare to the median? Are you over-indexing on technical buyers and missing economic buyers?
3. AI Perception Readiness Index Score your brand's visibility across AI models. The dataset powers the baseline: which personas are underserved by current AI recommendations, and where are the strategic gaps?
You can replicate these tools, improve them, or build something entirely different. The dataset is Apache 2.0 licensed. Attribution required, commercial use allowed.
What the Full Dataset Offers
The full 593,181-persona dataset includes everything in the 15K subset, plus:
- 40x more coverage: 593,181 personas across 339 industries (25 primary verticals) with deeper sub-segment stratification
- Longitudinal tracking: Persona evolution over time as we regenerate the dataset quarterly
- Narrative stories: Full buyer journey narratives (removed from the public subset for anonymization)
- Competitive context: Which personas prefer competitors, and why
- API access: Query personas by industry, pain point, or search intent via REST API
- OEM licensing: White-label integration into your product, with SLA and support
The full dataset is available through enterprise licensing (EUR 25K-120K/year) or API subscriptions (starting at EUR 2,500/month). Contact us at contact@rankfor.ai for pricing and evaluation access.
How to Use This Dataset
Download it from HuggingFace: https://huggingface.co/datasets/Rankfor/PersonaGen-15K
Load it in Python:
import pandas as pd
df = pd.read_parquet('hf://datasets/Rankfor/PersonaGen-15K/personagen-15k.parquet')
print(df.head())
Cite the research: Żatuchin, D. (2025). PersonaGen-593K: A Large-Scale Dataset of AI-Generated Buyer Personas for Studying LLM Recommendation Behavior. Discover Artificial Intelligence. [Link to paper]
Explore the tools: Try the Persona Matcher, Benchmark Calculator, or AI Perception Readiness Index to see what you can build with persona data.
Why This Matters for Your Brand
If AI assistants are recommending solutions to your potential customers, you need to know:
- Which personas are asking questions AI can't answer about your category (opportunity gaps)
- Which personas AI recommends to your competitors instead of you (competitive vulnerability)
- Which personas you're over-investing in with content that AI already favors (wasted effort)
This dataset is the foundation of that analysis. The 15K subset proves the methodology. The full 593K dataset powers the intelligence layer that answers those questions for your brand.
Download the dataset. Build with it. If you need the full commercial dataset, we're one email away.
About Rankfor.AI: We help B2B marketing teams understand their AI visibility—what AI models know about their brand, which personas they reach, and where competitors dominate the conversation. Our platform is built on the largest published dataset of AI-generated buyer personas. Learn more at rankfor.ai.
Download the dataset: HuggingFace Explore the tools: Persona Matcher | Benchmark Calculator | Readiness Score License the full dataset: contact@rankfor.ai
