research

When AI Changes Its Mind About Gender: Three Studies Prove Bias Is Real, Systematic, and Seasonal

February 15, 2026
20 min read
n=929
Dmitrij Żatuchin
Gender BiasAI ResearchLLM RecommendationsSeasonal BiasCategory GatekeepingPASORGeminiGPT

We ran 1,279 queries across three AI platforms and two gift-giving seasons. The result: definitive statistical proof that AI treats brands differently based on gender - and that bias changes direction depending on whether you ask during Christmas or Valentine's Day.


The Discovery That Ended the Debate

For a year, we have been documenting gender bias in AI brand recommendations. First, we found the Kindle phenomenon: 100% visibility for wife queries, 0% for husband queries. Then we expanded to child recipients and found the bias persisted: 28% fewer brands for daughters than sons.

Some argued these were isolated findings. Platform quirks. Random variance. Statistical noise.

This research ends that argument.

1,279 queries. Three complementary studies. Two seasonal contexts. One conclusion: gender bias in AI recommendations is systematic, statistically proven, and context-dependent.

The paper, submitted to Springer's Human-Centric Intelligent Systems journal, provides formal academic validation of what we first observed in real-world testing. The statistical evidence is no longer debatable.


The Three-Study Design: Why This Research Is Definitive

Most bias research tests one scenario and calls it done. We designed three interlocking studies to eliminate every alternative explanation:

Study 1: Christmas Gifts for Adults (December 2025, n=299)

Prompts tested: husband/wife/partner/2025 gift Key finding: 41% fewer brands for wife queries on Gemini Statistical proof: χ² = 137.32, p < 0.001

Study 2: Christmas Gifts for Children (January 2026, n=480)

Prompts tested: son/daughter/child/2026 gift Key finding: 28% fewer brands for daughter queries across all platforms Statistical proof: χ² = 524.32, p < 0.001 (nearly 4x stronger than Study 1)

Study 3: Valentine's Day Gifts for Partners (February 2026, n=500)

Prompts tested: husband/wife/boyfriend/girlfriend/partner Key finding: Seasonal REVERSAL - female-framed prompts receive 8.8% MORE brands on Gemini Statistical proof: χ² = 89.45, p < 0.001

Combined analysis: 1,279 independent observations across 12 prompt conditions and 3 AI platforms.

Three Studies: Sample Size & Gender Gap by Study

299480500-41-288.8Study 1: Christmas AdultsStudy 2: Christmas ChildrenStudy 3: Valentine's Day0100200300400500
Sample SizeFemale Brand Gap (%)

The statistical validation is overwhelming. Cramér's V values of 0.23-0.38 indicate medium to large effect sizes. Bootstrap confidence intervals confirm the effects are strong. This is not random variance. This is systematic bias.


The Reversal That Changed Everything

Here is the finding that transforms how we understand AI bias:

Christmas framing: AI recommends 28-61% fewer brands for female-targeted queries Valentine's Day framing: AI recommends 8.8% MORE brands for female-targeted queries on Gemini

Same AI. Same brand universe. Different season. Opposite bias direction.

PlatformChristmas Wife GapValentine's Wife GapSeasonal Shift
Gemini-41% brands+8.8% brands49.8% reversal
GPT-15% brands+12.1% brands27.1% reversal
Grok-35% brands-2.3% brands32.7% reversal

Seasonal Bias Reversal: Christmas vs Valentine's Day

-41-15-358.812.1-2.3GeminiGPTGrok−40−30−20−10010
Christmas Wife Gap (%)Valentine's Wife Gap (%)

This discovery proves three critical points:

  1. Bias is context-dependent. The same query structure produces opposite results in different seasonal contexts.
  2. AI has learned cultural patterns. Valentine's Day is coded as female-oriented in training data, reversing the default male bias.
  3. Single-point measurements are meaningless. Your brand might dominate one seasonal context while being invisible in another.

The academic term for this is "context-dependent bias modulation." The practical term is: your AI visibility strategy must account for seasonality, or you are measuring half the picture.


69 Gender-Locked Brands: The Full List

Across all three studies, we identified 69 brands that appear exclusively in one gender framing:

Male-Locked Brands (48 total):

  • Tech/Gadgets: Garmin, GoPro, Oculus, Philips Norelco, Manscaped
  • Tools: DeWalt, Milwaukee, Leatherman, Craftsman, Benchmade
  • Outdoor: Traeger, Weber, Osprey, North Face, Patagonia (partial)
  • Grooming: Dollar Shave Club, Manscaped
  • Watches: Rolex, Omega, Seiko, Tissot
  • Sports: Titleist, Under Armour, Theragun (Christmas only)

Female-Locked Brands (17 total):

  • Home: Stanley, KitchenAid, Vitamix, Breville, Our Place
  • Beauty: Glossier, Kate Spade
  • Wellness: Oura (Christmas only), Peloton (Christmas only)
  • Reading: Kindle (Christmas-dominant, Valentine's neutral)
  • Jewelry: Tiffany (Valentine's only)

Platform-Specific Locks (4 brands):

  • Ray-Ban, LG, Google, Samsung appear only in temporal/neutral queries

Study Note: Some brands shifted between studies. Theragun was male-locked in Study 1 but neutral in Study 3. Kindle went from 100% wife-locked to seasonal-dependent. AI's gender categorization is not static.

If your brand is on this list, you are invisible to half your potential market depending on how the question is framed. No amount of content optimization will fix that - the bias exists in AI's category model, not in your messaging.


The Mechanism: Category Gatekeeping

We tested two competing hypotheses for how AI creates gender bias:

Hypothesis 1: Stereotypical Concentration AI recommends the same few stereotypical brands repeatedly for female queries (low diversity).

Hypothesis 2: Category Gatekeeping AI excludes entire product categories from female queries but distributes recommendations evenly within the allowed categories (high diversity, but smaller pool).

Shannon entropy analysis proved Hypothesis 2 is correct.

Female-targeted queries score 0.86-0.88 on Shannon entropy (nearly identical to male queries). Gini coefficients are actually LOWER for female queries (0.30-0.48) than male queries (0.41-0.57), meaning more even distribution among recommended brands.

AI is not being stereotypical. AI is being exclusionary.

Category gatekeeping means AI decides which product categories are "appropriate" for each gender, then excludes entire brand sets from consideration. A brand in a "male-coded" category never even enters the recommendation pool for female queries.

This is worse than stereotyping. Stereotyping would at least keep your brand visible (even if positioned incorrectly). Gatekeeping makes your brand invisible to entire query patterns.

How Gatekeeping Works in Practice

When someone asks "What is the best gift for my wife?", AI does not search all brands and pick stereotypical ones. Instead:

  1. AI identifies "female-appropriate" product categories: jewelry, home goods, wellness, beauty, reading
  2. AI searches ONLY those categories for brand recommendations
  3. Brands in "male-coded" categories (tools, tech gadgets, outdoor gear, watches) are excluded from consideration
  4. Recommendations are distributed evenly among the allowed categories

The same process happens in reverse for husband queries, just with a larger allowed category set (hence the higher brand counts).

This is why Kindle appears for wife queries and not husband queries. Amazon's e-reader belongs to the "reading/home" category cluster, which AI has coded as female-appropriate. The product is gender-neutral. The category assignment is not.


Cross-Model Agreement: Still Terrible

Despite analyzing 1,279 queries, AI platforms still cannot agree on which brands to recommend:

StudyGemini-GPT OverlapGemini-Grok OverlapGPT-Grok Overlap
Study 1 (Christmas adults)14% (Jaccard 0.08)45% (Jaccard 0.38)28% (Jaccard 0.17)
Study 2 (Christmas children)10% (Jaccard 0.10)42% (Jaccard 0.35)25% (Jaccard 0.20)
Study 3 (Valentine's Day)18% (Jaccard 0.12)51% (Jaccard 0.45)33% (Jaccard 0.25)

Cross-Model Agreement (Jaccard Index) Across Studies

0.080.10.120.380.350.450.170.20.25Study 1Study 2Study 300.10.20.30.4
Gemini-GPTGemini-GrokGPT-Grok

Average cross-model agreement: 20-35%.

Translation: If you measure your AI visibility on only one platform, you are seeing 20-35% of the reality. The other 65-80% of brand recommendations are completely different on competing platforms.

Platform Personalities Remain Consistent

The three platforms maintain distinct "personalities" across all studies:

GPT-5.2: The Brand Minimalist

  • Averages 0.1-4.7 brands per response (lowest across all studies)
  • Often recommends zero specific brands, preferring categories ("consider a smartwatch")
  • Deliberate brand abstinence is a strategic choice, not a capability limitation

Gemini: The Brand Maximalist

  • Averages 3.7-11.9 brands per response (highest across all studies)
  • Consistently names 4-12 specific brands per answer
  • Most susceptible to gender bias effects (largest gaps between male/female queries)

Grok: The Efficient Middle

  • Averages 2.5-10.0 brands per response
  • Highest brand density per character (most product-focused)
  • Most stable across gender framings (smallest gaps)

If you are optimizing for AI visibility without testing all three platforms, you are optimizing for one personality while ignoring two others.


The PASOR Metric: Measuring What Actually Matters

Traditional metrics like "impressions" or "rankings" do not apply to AI recommendations. You are either included in the answer or you are invisible. There is no position 1 vs position 10.

We introduce the Prompt-Adjusted Share of Recommendation (PASOR) metric:

PASOR = (Your brand's appearance count / Total relevant queries) × 100

Example: Kindle in Study 1 (Christmas adults)

  • Wife queries: 100% PASOR (39 appearances / 39 wife queries)
  • Partner queries: 100% PASOR (30 appearances / 30 partner queries)
  • Husband queries: 0% PASOR (0 appearances / 40 husband queries)
  • Overall PASOR: 63.4% (69 appearances / 109 total queries)

But here is the problem: overall PASOR masks gender gaps. If you measure only your aggregate score, you will miss that you dominate one demographic while being invisible to another.

Gender-Segmented PASOR reveals the truth:

  • Male-framed PASOR: 0%
  • Female-framed PASOR: 100%
  • Gap: 100 percentage points

For brands seeking balanced visibility, the gap matters more than the average.

PASOR in Action: Valentine's Day Reversal

Study 3 PASOR for select brands:

BrandHusband PASORWife PASORBoyfriend PASORGirlfriend PASORGap
Tiffany0%73%0%80%-73pp
Apple67%60%70%57%+7pp (balanced)
LEGO33%20%37%23%+13pp
Kindle27%40%23%43%-13pp (Valentine's neutral vs Christmas extreme)

Apple achieves near-universal PASOR (57-70% across all framings). Tiffany is completely gender-locked to female prompts. Kindle's gender gap narrowed from 100 percentage points in Study 1 to 13-17 points in Study 3, proving category assignments can shift seasonally.


The Statistical Evidence: Beyond Doubt

For readers questioning whether these effects are real or just measurement noise, here is the formal statistical validation:

Chi-Square Tests (Category-Gender Independence)

Studyχ² Valuep-valueCramér's VEffect SizeInterpretation
Study 1137.32<0.0010.23MediumReject independence
Study 2524.32<0.0010.38Medium-LargeReject independence
Study 389.45<0.0010.26MediumReject independence

Translation: There is less than a 0.1% chance these patterns are random. Product categories and gender framing are systematically linked.

Effect Sizes (Brand Count Gaps)

ComparisonCohen's d95% CIEffect SizePractical Meaning
Study 1: Husband vs Wife (Gemini)1.24[0.89, 1.59]Large41% fewer brands
Study 2: Son vs Daughter (All models)0.82[0.61, 1.03]Large28% fewer brands
Study 3: Valentine's Wife REVERSAL-0.31[-0.58, -0.04]Small8.8% MORE brands

Effect sizes are reported with bootstrapped confidence intervals (10,000 resamples). All effects are strong to sampling variation.

Cross-Validation

We validated findings using eight independent statistical tests:

  1. Chi-square tests (category-gender association)
  2. Cohen's d (brand count differences)
  3. Shannon entropy (diversity within gender groups)
  4. Gini coefficient (concentration/inequality)
  5. Jaccard index (cross-model overlap)
  6. Cramér's V (effect size for chi-square)
  7. Bootstrap confidence intervals (robustness check)
  8. Brand migration analysis (temporal stability)

All eight methods converge on the same conclusion: gender bias is systematic, platform-dependent, and context-dependent.

Effect Sizes Across Studies (Cram\u00e9r\u0027s V)

0.230.380.26Study 1: Christmas AdultsStudy 2: Christmas ChildrenStudy 3: Valentine's Day00.050.10.150.20.250.30.350.4

What Changed Between Studies: Brand Migration Patterns

Some brands maintained consistent gender categorization across all three studies. Others shifted:

Stable Gender Locks:

  • Male-locked across all studies: Garmin, DeWalt, Milwaukee, Rolex, Omega
  • Female-locked across all studies: Stanley, KitchenAid, Glossier

Shifted Categorization:

  • Theragun: Male-locked (Study 1) → Neutral (Study 3)
  • Peloton: Male-locked (Study 1) → Neutral (Study 3)
  • Kindle: 100% female-locked (Study 1) → Seasonal-dependent (Study 3, still 13-17pp gap)
  • Oura: Female-locked (Study 1) → Neutral (Study 3)

Seasonal-Only Appearance:

  • Tiffany, Cartier: Invisible in Christmas studies, female-locked in Valentine's study
  • Ray-Ban, LG: Temporal queries only, never gender-framed

Key insight: AI's gender categorization evolves. A brand that was gender-locked last quarter might have shifted. Regular testing is the only way to know your current status.


Why "Partner" and "Child" Are Not Gender-Neutral

We tested gender-neutral language in all three studies. The results prove neutrality is an illusion.

Study 1: "Partner" = Wife Pattern

  • Partner queries share 69.2% of brands with wife queries
  • Partner queries share 37.5% of brands with husband queries
  • Partner is not neutral. It defaults to female framing.

Study 2: "Child" = Daughter Pattern

  • Child queries share 56.2% of brands with daughter queries
  • Child queries share 42.6% of brands with son queries
  • Child is not neutral. It defaults to female framing.

Study 3: "Partner" = Wife Pattern (Again)

  • Partner queries share 64.8% of brands with wife queries
  • Partner queries share 41.3% of brands with husband queries
  • Pattern persists across seasonal contexts.

If your brand strategy relies on "inclusive" or "gender-neutral" positioning, you may be optimizing for female-framed visibility while missing male-framed recommendations entirely. AI interprets neutrality as female by default.


The Seasonal Context Dependency: Christmas vs Valentine's

This is the finding that fundamentally changes how we understand AI bias.

Hypothesis tested: Does the same gender-framing produce consistent bias across different seasonal contexts?

Answer: No. Bias direction reverses.

MetricChristmas (Studies 1&2)Valentine's Day (Study 3)Change
Gemini wife vs husband gap-41% brands+8.8% brands49.8pp reversal
Female-locked brands17 brands12 brands (but different set)-29%
Male-locked brands48 brands22 brands (but different set)-54%
Gender-neutral PASOR balanceFemale-biasedMore balancedImproved

Why this happens: AI training data encodes cultural patterns. Christmas gift-giving is traditionally male-framed in Western culture (husbands buy gifts for wives, fathers buy for children). Valentine's Day is female-framed (boyfriends buy for girlfriends, husbands for wives).

AI does not apply a universal gender bias. It applies context-specific learned patterns from training data.

This explains why single-point measurements are misleading. A brand that dominates Christmas recommendations might be invisible during Valentine's Day, or vice versa. Seasonal testing is mandatory for complete visibility assessment.


What This Means for Your Brand

If You Are in a "Male-Coded" Category

Examples: Tools, tech gadgets, outdoor gear, sports equipment, luxury watches

Your reality:

  • Strong visibility in husband/son/boyfriend queries
  • Weak or zero visibility in wife/daughter/girlfriend queries during Christmas
  • May improve visibility for female-framed queries during Valentine's Day (platform-dependent)

Your risk:

  • Missing 40-50% of potential market during high-volume gift seasons
  • Content optimization will not fix this - the bias is in AI's category model

Your strategy:

  1. Test PASOR across all gender framings for each seasonal context
  2. Identify which seasonal contexts favor your category
  3. Time major campaigns around favorable seasonal windows
  4. Consider whether your category assignment is accurate (appeal to AI teams if not)

If You Are in a "Female-Coded" Category

Examples: Jewelry, home goods, beauty, wellness, kitchen appliances

Your reality:

  • Strong visibility in wife/daughter/girlfriend queries during Christmas
  • May face reduced advantage during Valentine's Day (seasonal reversal)
  • Weak or zero visibility in husband/son/boyfriend queries

Your risk:

  • Leaving 40-50% of market to competitors who solved the category problem
  • Being invisible to male gift-givers who want to buy your product for themselves

Your strategy:

  1. Test whether your product is genuinely gender-specific or just gender-categorized
  2. If product is neutral but category-locked, build content that bridges categories
  3. Monitor seasonal shifts - your Valentine's advantage may not persist year-round

If You Are Platform-Specific

Examples: Ray-Ban, Samsung, Google (temporal-only appearance), Tiffany (Valentine's-only)

Your reality:

  • Invisible in standard gender-framed queries
  • Appear only in temporal or seasonal-specific contexts
  • May be excluded from LLM training data or fall below recommendation thresholds

Your strategy:

  1. Investigate why you are absent from standard recommendations
  2. Boost brand mentions in contexts AI uses for training
  3. Test whether different query structures unlock visibility

If You Achieve Balanced PASOR (Like Apple)

Balanced brands: Apple, Sony, Bose, LEGO, Patagonia, Nintendo

Your reality:

  • Appear consistently across gender framings (±10-15% variance)
  • Maintain visibility across seasonal contexts
  • Not locked to any specific demographic category

Your advantage:

  • Maximum recommendation reach across all query types
  • Lower risk of seasonal volatility
  • Better insulated from category gatekeeping

Your strategy:

  1. Maintain balanced positioning across all content
  2. Monitor for category drift that could lock you into one demographic
  3. Use your balanced status as a competitive advantage

The Academic Validation: Peer Review Matters

This research has been submitted to Springer's Human-Centric Intelligent Systems journal - a peer-reviewed academic publication in the field of AI and human-computer interaction.

Why this matters:

  1. Independent review: Three anonymous experts in AI bias research will validate methodology, statistical rigor, and conclusions before publication.

  2. Reproducibility: Full methodology, statistical tests, and effect sizes are documented. Other researchers can replicate the findings.

  3. Citation permanence: Once published, this research becomes part of the permanent academic record on AI bias, cited by future studies.

  4. Industry credibility: Academic validation carries weight with enterprise buyers, regulators, and media.

The submission package includes:

  • Full paper (30 pages, 1,279 observations analyzed)
  • Statistical appendix with all eight validation tests
  • Raw data availability statement
  • Conflict of interest disclosure (research funded by Rankfor.AI, author is founder)

Current status: Under review (submitted February 2026). Expected decision: April-May 2026.

If you are presenting this research to executives, regulators, or journalists, you can cite it as "under peer review at Springer Human-Centric Intelligent Systems." Once accepted, it becomes citable academic literature.


The Implications for AI Fairness Research

This research contributes to the broader AI fairness literature in three ways:

1. Context-Dependent Bias Modulation

Prior work on AI bias assumes bias is static - if a model is biased, it applies that bias consistently. Our findings prove this is wrong for LLM recommendations.

Key contribution: Gender bias in commercial AI is not a fixed parameter. It is modulated by seasonal context, recipient type, and relationship framing. Fairness audits that test only one context will miss reversals in other contexts.

2. Open-World Learning Challenges

LLMs operate in open-world environments where user queries span unbounded demographic and seasonal contexts. Traditional closed-world bias detection (testing on fixed datasets) cannot capture this complexity.

Key contribution: We demonstrate that LLM bias manifests differently across seasonal contexts that were likely underrepresented in training data. This aligns with open-world learning research on how models generalize beyond training distributions.

3. Category Gatekeeping as a Bias Mechanism

Most bias research focuses on stereotypical concentration (recommending the same few options repeatedly). We identify a different mechanism.

Key contribution: AI does not stereotype female-targeted queries with narrow brand sets. It excludes entire product categories from consideration, then distributes recommendations evenly within the allowed categories. This "category gatekeeping" is harder to detect because diversity metrics (Shannon entropy, Gini coefficient) appear normal.


What You Should Do Now

1. Run a Multi-Seasonal PASOR Audit

Test your brand across:

  • All gender framings (husband/wife/boyfriend/girlfriend/partner)
  • All seasonal contexts (Christmas, Valentine's Day, Mother's/Father's Day, birthdays)
  • All three major platforms (Gemini, GPT, Grok)

Measure PASOR for each combination. Identify where you are visible, where you are locked out, and whether you face seasonal reversals.

2. Check Your Category Assignment

If you are gender-locked but your product is genuinely gender-neutral, investigate:

  • Which product categories does AI associate your brand with?
  • Are those categories coded as male or female in AI responses?
  • Can you bridge categories through content that crosses demographic boundaries?

3. Monitor for Brand Migration

Our research proves gender categorization is not static. Brands shift between studies. Set up quarterly testing to detect when your status changes.

4. Segment Your Visibility Metrics

Overall PASOR hides gender gaps. Always report:

  • Male-framed PASOR (husband/son/boyfriend)
  • Female-framed PASOR (wife/daughter/girlfriend)
  • Neutral-framed PASOR (partner/child)
  • Gap between male and female (this matters more than the average)

5. Adjust Strategy by Season

If you discover seasonal reversals (like the Christmas-to-Valentine's shift), time your campaigns accordingly:

  • Push female-targeted messaging during seasons when female-framed queries receive more brands
  • Push male-targeted messaging during seasons when male-framed queries dominate
  • Test emerging seasonal contexts (Prime Day, Black Friday, back-to-school)

Check If Your Brand Is Gender-Locked

We built a free tool that runs your brand through the same analysis we used in this research.

What you will see:

  • PASOR scores across all gender framings
  • Which seasonal contexts favor your brand
  • How your visibility compares across Gemini, GPT, and Grok
  • Whether you are category-locked like Kindle or balanced like Apple

Three prompts. Three platforms. Three seasonal contexts. Instant PASOR analysis.

No email required. Takes 90 seconds.

Check Your Brand's Gender Lock Status


Methodology Summary

  • Total Queries: 1,279 (299 + 480 + 500 across three studies)
  • Platforms Tested: Google Gemini 3 Flash, OpenAI GPT-5.2, xAI Grok-4-1
  • Prompts per Study:
    • Study 1: 4 conditions (husband/wife/partner/2025)
    • Study 2: 4 conditions (son/daughter/child/2026)
    • Study 3: 5 conditions (husband/wife/boyfriend/girlfriend/partner)
  • Temperature: 0.7 (Studies 2-3), default (Study 1)
  • Data Collection: December 2025 - February 2026
  • Statistical Validation: Chi-square tests, Cohen's d, Cramér's V, Shannon entropy, Gini coefficient, Jaccard index, bootstrap confidence intervals
  • Effect Sizes: Medium to large (Cramér's V = 0.23-0.38, Cohen's d = 0.82-1.24)
  • Peer Review: Submitted to Springer Human-Centric Intelligent Systems, February 2026

Full statistical appendix and raw data available upon request. Contact research@rankfor.ai for academic collaboration inquiries.


Kindle: 100% of Wife Queries. 0% of Husband Queries. Same AI.

Our original discovery (December 2025, n=299) that first identified the Kindle phenomenon and established the statistical foundation for gender bias in LLM recommendations. This research became Study 1 in the three-study analysis.

Key Finding: 41% fewer brands for wife queries. Chi-square 137.32, p<0.001.

Only 2 Brands Are Visible Across All AI Platforms. Is Yours One of Them?

The child recipient study (January 2026, n=480) that proved gender bias extends beyond adult contexts and identified universal visibility as nearly impossible to achieve. This research became Study 2 in the three-study analysis.

Key Finding: Only LEGO and Disney achieve universal cross-platform visibility. 28% fewer brands for daughter queries.


Research conducted by Dmitrij Żatuchin, Estonian Entrepreneurship University of Applied Sciences (EUAS). Submitted to Springer Human-Centric Intelligent Systems, February 2026. Full paper, statistical appendix, and raw data available upon request.


Key Terms

PASOR (Prompt-Adjusted Share of Recommendation): The percentage of relevant queries where your brand appears in AI responses. Unlike traditional rankings, PASOR measures inclusion vs exclusion, not position.

Category Gatekeeping: When AI excludes entire product categories from gender-specific queries rather than stereotyping within an allowed set. Identified as the primary bias mechanism through Shannon entropy analysis.

Gender-Locked Brand: A brand that appears exclusively in one gender framing (e.g., only husband queries) and never in others. Our research identified 69 such brands across three studies.

Seasonal Bias Reversal: When the direction of gender bias changes based on seasonal context. Christmas reduces female-framed recommendations by 28-61%, while Valentine's Day increases them by 8.8% (Gemini).

Cross-Model Agreement (Jaccard Index): The percentage of brands recommended by multiple AI platforms for the same query. Low agreement (10-45% in our studies) indicates platform-dependent visibility.

Cramér's V: A measure of association strength for chi-square tests, ranging 0 (no association) to 1 (perfect association). Values of 0.23-0.38 indicate medium to large gender-category associations in our studies.

Cohen's d: A standardized measure of difference between two groups. Values of 0.82-1.24 indicate large practical differences in brand counts between gender framings.

Brand Migration: When a brand's gender categorization changes between measurement periods. Theragun, Peloton, and Oura all shifted from gender-locked to neutral between studies.

People Also Ask

Continue Reading

research

70% of AI Gift Recommendations Are Gender-Locked

150 queries on Valentine's Day 2026 across three AI models reveal 70% of brand recommendations are gender-exclusive. Gemini names 11.4 brands per response vs 1.3 for Grok and 1.2 for OpenAI. Only Away and Ember transcend gender silos. Female-framed prompts now get 16.2% more brands than male - the opposite of Christmas.

research

Four-Prompt Gender Bias Analysis: Isolating Gender Effects in LLM Brand Recommendations

This study extends cross-prompt LLM brand consistency analysis by introducing a gender-neutral "partner" prompt to isolate gender bias effects. Comparing four prompt variations across 299 total samples, we provide strong evidence of gender-based brand recommendation bias in LLMs. Critical finding: Kindle appeared in 100% of wife AND partner iterations but 0% of husband iterations.

research

Only 2 Brands Are Visible Across All AI Platforms. Is Yours One of Them?

We ran 480 queries across three AI platforms. Cross-model agreement on brand recommendations was only 10-45%. We identified 26 gender-locked brands, discovered that only LEGO and Disney achieve universal visibility, and found that AI recommends 41% fewer brands for wife queries than husband queries.

research

Cross-Model Response Stability Analysis: A Comparative Study of Brand Mention Consistency in Large Language Models

This study investigates response variability and brand mention consistency across three leading LLMs: Google Gemini 3 Flash, OpenAI GPT-5.2, and xAI Grok-4-1. Using 239 total samples across three prompt variations (husband, wife, 2025), we measured brand mention frequency, response consistency, and cross-model agreement.

BeVisible Club

Get research like this before we publish it.

New AI visibility studies, ranking patterns, and source-stack data delivered to your inbox a week before they go public.

Join BeVisible Club

Want to Know How AI Sees Your Brand?

Rankfor.AI measures your brand's AI visibility across all major platforms and provides actionable recommendations.

About the Author

Dmitrij Żatuchin

Founder & Lead Researcher

Dmitrij Żatuchin is the founder of Rankfor.AI. A computer scientist with a PhD in semantic web technologies, he bridges the gap between how AI reasons about brands and how brands want to be understood. With over two decades of software architecture experience and academic roles at Estonian Business School, Dmitrij builds the measurement infrastructure brands need to transition from optimizing for search engines to becoming visible for reasoning engines.

© 2025-2026 Rankfor.AI™ All rights reserved.

Ask AI about Rankfor.AI

Rankfor.AI sp. z o.o., Skarbowcow 23B, 53-025 Wroclaw, Poland

KRS: 0001190083 | NIP: 8993033605

We use cookies

We use essential cookies to make our site work. With your consent, we may also use analytics cookies to understand how you use our tools (like the Dice Roller) so we can improve them. Learn more