What We Learned From 12 Weeks of Tracking 11 AI Search Engines
12 weeks of monitoring 11 AI engines across hundreds of brands produces patterns you don't see from one engine alone. Here's what the data actually shows about citation behavior in 2026.
By Julian Hernandez ยท
The short answer
Tracking 11 AI engines (ChatGPT, Perplexity, Gemini, Claude, Google AI Overviews, Copilot, Grok, DeepSeek, Mistral, Meta AI, Qwen) across the brands LiftRank monitors produced patterns no single-engine analysis would surface. The five biggest findings from the first 12 weeks of cross-engine tracking: engine behavior is more divergent than we expected, third-party citations dominate by even more than the public research suggested, stochastic variance is real but bounded, freshness penalties are larger than schema penalties, and brand mentions without citations dominate engine output by a factor of 2โ3ร. This post is the candid version of what the data actually showed โ including where we were wrong about the engines going in.
Why tracking 11 engines matters
A brand monitoring only ChatGPT misses 80% of the AI-search picture. A brand monitoring ChatGPT plus Google AI Overviews misses 50%. Even five-engine coverage leaves meaningful blind spots for brands whose audience extends into Claude (technical buyers), Copilot (Microsoft 365-bundled users), Meta AI (consumer audiences via WhatsApp and Instagram), or the long-tail engines (Grok, DeepSeek, Mistral, Qwen) where non-English and specific-platform users cluster.
The aggregate effect is that single-engine and few-engine measurement systematically produces partial pictures. The brands acting on partial pictures make optimization decisions that work for one engine and quietly miss the others. Cross-engine tracking surfaces the patterns that single-engine analyses can't see.
The five findings below are what 12 weeks of running the same prompt sets across all 11 engines produced as the actually-useful conclusions, separate from the obvious "AI search matters" claims everyone already agreed on.
Finding 1 โ Engine behavior diverges more than expected
Going into the cross-engine tracking, we expected meaningful per-engine differences but a broadly common core: same prompts producing roughly similar brand sets across engines, with some variation in ranking and framing.
The actual data showed much wider divergence. For a typical category prompt across the brands we monitor:
- ChatGPT mentioned 2โ3 brands per answer with cautious framing
- Perplexity mentioned 1โ3 brands with explicit citations and tighter recommendations
- Gemini mentioned 4โ7 brands per answer with longer, more descriptive responses
- Google AI Overviews mentioned 2โ4 brands inside the answer with 6โ10 cited sources below
- Claude mentioned 1โ3 brands with content drawn mostly from training data
- Copilot mentioned 3โ5 brands with high average position
- Grok, DeepSeek, Mistral, Meta AI, Qwen each showed distinct patterns we can't summarize neatly
The cross-engine overlap on category prompts โ same brand mentioned across multiple engines โ ran around 30โ40% for the loudest brands and dropped to under 15% for mid-tier brands. A brand dominant on Gemini was often weak on ChatGPT and vice versa. The divergence isn't noise; it's structural, driven by each engine's different source weighting and training data.
The implication: optimization strategies that win on one engine don't automatically win on the others. The brands seeing the strongest aggregate AI visibility lift are the ones investing in engine-specific patterns, not generic GEO tactics.
Finding 2 โ Third-party citations dominate by even more than expected
The public research we'd seen (Erlin's 500-brand study, GenOptima's 20-prompt cross-engine work) put third-party citations at around 68โ85% of AI brand mentions. We assumed our data would confirm that number.
The actual figure across the brands we monitor sits at the higher end of that range โ closer to 85% for B2B SaaS brands and over 90% for consumer brands. Owned-website content is a smaller share of AI citations than even the bullish public estimates suggested.
The third-party surfaces driving citations were also more concentrated than expected:
- For B2B SaaS: G2, Reddit, HubSpot's blog, and 1โ2 category-specific industry publications drove 70%+ of the third-party citations for a typical brand
- For consumer brands: Trustpilot, Reddit, Amazon reviews (even for brands not selling on Amazon), and 1โ2 curated review sites (Wirecutter or category equivalents) drove similarly high concentration
- For local businesses: Google Business Profile data and 1โ2 local-directory sites dominated
The strategic implication: brands optimizing only their own pages are working on 10โ15% of the citation graph. The largest available lift comes from earning coverage on the 4โ6 third-party surfaces most-cited in your category, not from publishing more on-domain content.
Finding 3 โ Stochastic variance is real but bounded
AI engines are probabilistic, which means the same prompt run on different days produces slightly different brand recommendations. We expected this would make weekly numbers noisy and force us toward longer reporting windows.
The actual variance across 12 weeks: week-over-week mention rate on a stable prompt set typically swings within ยฑ3โ5 percentage points without any underlying visibility change. Larger swings on a single engine in a single week are almost always either real events (engine updates, third-party content changes, content publishes) or methodology shifts in the monitoring tool.
The practical rule we landed on: aggregate to week-over-week deltas, trust trends that hold for 3+ consecutive weeks, and validate any single-week move over 10 points before treating it as a real signal. Below that discipline, teams interpret stochastic noise as meaningful change and over-react.
The variance is also engine-specific. Perplexity (live-search-driven) shows lower week-over-week variance because its real-time crawling tightens the answer to current content. ChatGPT (training-data-leaning for many queries) shows higher variance because the model samples from a distribution. Gemini sits in the middle.
Finding 4 โ Freshness penalties are larger than schema penalties
Going in, we expected schema markup deployment to produce the largest measurable single-pattern lift in AI citations. The public research from Erlin and others showed schema lifting citation rates roughly 3ร.
The actual pattern we observed across 12 weeks: freshness penalties (citation rates dropping on stale content) were larger and faster than schema benefits (citation rates rising on schema-marked content). Content older than 6 months on time-sensitive topics lost mention rate at roughly 1.5โ2% per month; the same content with stacked schema deployed but no freshness update didn't recover the lost ground.
The asymmetry has a strategic implication. For brands choosing where to invest scarce editorial time, refreshing existing high-value content quarterly produces faster measurable AI citation lift than deploying schema on un-refreshed pages. Schema still matters; it's just a slower lever than freshness on time-sensitive content.
Caveat: the freshness-over-schema pattern held strongest for time-sensitive categories (tooling, statistics, pricing comparisons, industry news). For evergreen conceptual content, schema's relative leverage was higher and freshness mattered less.
Finding 5 โ Mentions without citations dominate by 2โ3ร
Across the engines we monitor, the ratio of "brand mentioned in the answer text" to "brand cited via a clickable link" runs about 2.5:1 on average. For every clickable citation a brand earns, it gets roughly 2โ3 mentions inside answer prose without an attached link.
The split varies by engine:
- Perplexity runs close to 1:1 (cites almost every named source)
- Google AI Overviews runs about 1.5:1
- Gemini runs about 2:1
- ChatGPT runs about 3:1 (mentions brands frequently in prose, cites less often)
- Claude runs about 4:1 (especially in default mode without web search)
The implication: brands measuring only citation rate (clickable links) systematically underestimate their AI visibility. The mention-without-citation outcome drives brand recall, downstream branded search, and indirect funnel influence that doesn't show in referrer-based analytics. Measuring mention rate alongside citation rate is the only way to see the full impact.
For brands with strong branded-search trends in their analytics alongside flat AI referral traffic: that's the mention-without-citation pattern working in your favor. Reading it as "GEO isn't producing results" misses the actual signal.
What we got wrong going in
Three things we expected that the data didn't confirm.
Wrong expectation one: we thought engine convergence would happen faster. We assumed by mid-2026 the major engines would have converged on broadly similar source-selection patterns. The data shows the opposite โ divergence is widening, not narrowing, as each engine's product team optimizes for different user experiences.
Wrong expectation two: we thought sentiment classification would be solid by now. The third-party sentiment classifiers we evaluated, including our own, had error rates of 10โ20% on individual mentions. The aggregate sentiment numbers are still directionally useful, but single-mention sentiment classification isn't reliable enough to drive specific actions yet.
Wrong expectation three: we thought the long-tail engines (Grok, DeepSeek, Mistral, Meta AI, Qwen) would be aggregate noise. They're not. For brands with global exposure, the long-tail engines collectively account for 15โ25% of meaningful AI search activity, with specific engines dominating specific markets (Qwen for China, DeepSeek for technical audiences). The "ignore the long tail" position we considered going in turned out to be wrong for any brand with non-US exposure.
The candid version: tracking 11 engines surfaced patterns we wouldn't have predicted, and the patterns changed how we built the product and how we recommend brands invest. The next 12 weeks will surface more.