1. Why This Study Exists
The Generative Engine Optimization industry is roughly two years old. It already has its own analyst tier, its own awards, and its own cottage industry of "State of AI Search" reports. Almost all of it is funded by tools that sell GEO services.
That isn't an indictment by itself. The most rigorous SEO research of the 2010s — Moz's domain authority work, Ahrefs' link studies, Backlinko's content analyses — was vendor-funded too. Vendor research has the data; independents usually don't.
But marketing leaders are now making procurement decisions based on these numbers. "AI traffic converts 4x better." "Schema lifts citations 44%." "LinkedIn citations doubled in three months." When a $50K/year tool decision is downstream of a vendor's own benchmark, the stakes of bias are higher than a clever statistic on LinkedIn.
This study asks one question: does the funding source of a GEO study predict the size and direction of its findings?
The hypothesis was yes — vendor-funded studies should systematically report larger effects, because (a) selection bias in case studies, (b) the prompt-set design and reporting choices favor positive results, and (c) null findings rarely make it into a marketing report.
What I found, with this 30-study sample, is consistent with that hypothesis but more nuanced than I expected. Vendors aren't lying. They're reporting real effects on self-selected samples. The independents aren't necessarily right either — most of them have smaller sample sizes and narrower scope. The honest reading is that the GEO measurement industry has a credibility asymmetry, not a credibility crisis.
2. Methodology
Inclusion criteria. A "study" was included if it met all four:
- Published or refreshed between January 2025 and May 2026.
- Reports a quantitative finding about AI search, GEO, AI Overviews, AI citations, or AI-channel conversion.
- Has a primary public source (vendor blog, academic paper, news article, agency report).
- Cited at least once in independent third-party reporting (Search Engine Land, Search Engine Journal, Adweek, eMarketer, Press Gazette, TechCrunch, Axios, etc.) — to filter out unsourced LinkedIn anecdotes.
Source pool. The 30 studies were drawn from four scraped research files (01-GEO-Stats-Data-Research.md, 02-AI-Search-Experiments-Audits.md, 04-SaaS-Brand-AI-Visibility-By-Category.md, 06-Technical-GEO-Implementation-Guides.md) — totaling roughly 50 candidate studies. I filtered to 30 representative cases across funding categories, study types, and topic areas (CTR, citations, schema, conversion, B2B buyer behavior, crawler economics).
Classification scheme.
- Funding/affiliation (4 categories):
- Vendor: Authored by a company that sells a GEO tool, AI search visibility platform, or directly monetizes the result of the finding (e.g., Ahrefs, BrightEdge, Semrush, Profound, AthenaHQ, Otterly, Peec, Conductor, Seer Interactive, SE Ranking).
- Vendor-adjacent: Consultancy, agency, or media outlet with commercial skin in the game (e.g., Tinuiti, Reboot Online, iPullRank, Growth Marshal).
- Independent: Academic, journalist, neutral analyst, or research institution with no commercial GEO product (e.g., Pew Research, Forrester, Gartner, McKinsey, SSRN papers, Adobe — included as "independent-mixed" because Adobe sells martech but has no direct GEO product).
- Mixed: Joint vendor + neutral collaboration (e.g., Tinuiti × Profound).
- Sample size — best-disclosed numeric (queries, citations, brands, keywords, panelists).
- Methodology rigor (5-point scale): Peer-reviewed (5), Controlled randomized field experiment (4), Large observational dataset with disclosed methodology (3), Observational with limited methodology disclosure (2), Single case study or anecdote (1).
- Effect size of headline finding — the most-quoted number from the study. Where possible, normalized to a percentage lift (positive or negative). Correlations and absolute volumes are recorded but excluded from the mean-effect comparison (see "limitations").
- Direction (Pro-GEO / Skeptical / Neutral): Pro-GEO = headline says GEO/AI search creates lift, citations, conversion gain. Skeptical = headline says effect is null, smaller than claimed, or negative. Neutral = descriptive (e.g., "X% of users do Y") with no implicit endorsement.
- Reproducibility (3-point): Yes (methodology disclosed sufficiently), Partial (sample described but no full method), No.
Limitations. I am the sole classifier; another researcher could disagree on borderline cases (Tinuiti × Profound is "mixed" by my call but defensibly "vendor"). Effect size is not always comparable across studies — a "0.664 correlation" and a "+44% lift" are not the same kind of number. Where I report a mean-effect-size delta, I'm using only studies whose headline is expressed as a percentage lift or share, and I exclude correlations from that mean. See Section 7 for the full table and Section 11 for full reproducibility notes.
3. The Big Number — Vendor vs. Independent Effect Size
Across the 30 studies in this sample, the mean normalized headline effect size for vendor and vendor-adjacent studies is 49.7%, vs. 20.6% for independent and academic studies — a ~2.4x delta.
Calculation. I restricted this comparison to the 22 studies whose headline finding is expressed as a percentage lift, share, or absolute conversion delta (excluding correlations and pure descriptive statistics). Each finding was converted to its absolute value to compare magnitude of claimed effect, irrespective of direction (so a "-61% CTR drop" enters the calculation as 61%, same as a "+61% lift").
| Funding Category | n studies (in mean) | Mean Effect Size | Median Effect Size | Range |
|---|---|---|---|---|
| Vendor | 12 | 49.7% | 44% | 8% – 393% |
| Vendor-adjacent | 4 | 38.5% | 32% | 18% – 80% |
| Independent / Academic | 6 | 20.6% | 18% | 1% – 61% |
n=12 · median 44% · range 8–393%
n=4 · median 32%
n=6 · median 18% · range 1–61%
What this does and doesn't show. It does not show vendors are wrong. The 393% outlier (Adobe — Q1 2026 retailer AI traffic growth) is a legitimately measured number. Removing the top and bottom outlier from each category narrows the gap to ~1.7x but does not eliminate it.
What it does show: the category of "what gets reported as a headline" is systematically different by funding source. Vendors report the dramatic number. Independents report the boring one — even when both are technically true on the same dataset.
4. Direction Skew — Who Reports What
Of the 30 studies:
- 15 are clearly Pro-GEO (headline says GEO/AI works, lifts, converts, cites): 13 are vendor or vendor-adjacent; 2 are independent (Adobe, Forrester — both with commercial-AI exposure, see notes).
- 9 are Skeptical (headline says effect is null, smaller, negative, or describes a problem): 8 are independent or academic; 1 is vendor (Otterly's null schema result, notable because it's a vendor contradicting a vendor consensus).
- 6 are Neutral (descriptive without a directional claim): mixed across categories.
Of the strongly pro-GEO studies (n=12, where the headline is a positive lift > 30%): 11 are vendor.
the exception: Adobe (independent-mixed)
the exception: Otterly — a vendor publishing a null result
Phrased the other way: if you only read independent research, you would conclude AI search is real but causes more publisher pain than brand opportunity. If you only read vendor research, you would conclude AI search is the biggest growth channel since paid social. Both readings are partial. The truth is in between, and the bias is detectable.
5. Specific Contradictions — Where the Field Disagrees on Itself
These are the four cleanest "vendor vs. independent" contradictions I could isolate. Each is a case where two studies, looking at substantially the same question, produce findings that are hard to reconcile.
Contradiction #1 — Schema Markup Citation Lift
| Study | Claim | Sample | Funding |
|---|---|---|---|
| BrightEdge — Structured Data in AI Search | +44% AI citation lift with structured data + FAQ blocks | Undisclosed enterprise dataset | Vendor (sells SEO/GEO platform) |
| Otterly — Schema Markup Real Impact Test (Q1 2026) | No isolated causal effect on ChatGPT citations; 6 of 7 AI engines couldn't fetch JSON-LD on demand | Controlled cross-engine test | Vendor (sells GEO platform — competitor) |
| Stackmatix — Aggregated 73-Site Schema Study (early 2026) | 3.2x citation rate with proper schema | 73 sites | Vendor-adjacent (agency) |
| SE Ranking — 216K-page study | FAQ schema alone does not move citations | 216K pages | Vendor-adjacent (SEO tool) |
Interpretation. This is the most direct vendor-vs-vendor contradiction in the dataset. BrightEdge's "+44%" is the most-quoted number in the field; Otterly's null finding is barely cited outside the SEO trade press. Both are vendor studies. The SE Ranking and Stackmatix splits suggest the answer is "schema is necessary infrastructure, not a citation lever on its own" — but that nuance is missing from how the BrightEdge number is repeated.
Contradiction #2 — AI Search Conversion Rate
| Study | Claim | Sample | Funding |
|---|---|---|---|
| Growth Marshal — Aggregated Conversion Panels | AI referrals convert at 11.4% vs. 5.3% organic (~2.2x) | "Multi-source aggregation" — methodology partial | Vendor-adjacent (consultancy) |
| BrightEdge — AI Search Visits Surging | AI sessions convert at 7.05% vs. 5.81% organic (~1.2x) | Enterprise customer base | Vendor (sells SEO/GEO platform) |
| Adobe — March 2026 AI Traffic Inflection | AI traffic converts +42% better than non-AI | Adobe Analytics data, large but undisclosed n | Independent-mixed (sells martech) |
| Similarweb — 2026 GenAI Brand Visibility Index | AI referrals plateaued at under 1% of traffic for most brands | 11K+ prompts in Finance alone, 113 brands | Vendor (sells competitive intelligence) |
Interpretation. The conversion-rate numbers are not directly contradictory — they're measuring different things on different cuts. But the headlines contradict the base rate. Adobe says AI converts 42% better; Similarweb says AI is under 1% of traffic for most brands. Both are true. The marketer who reads only the conversion stat thinks AI is a primary channel; the marketer who reads only the volume stat thinks it's irrelevant. (Do the math and the two reconcile: 1% volume × 4.4x conversion = 4.4% of conversions today, doubling roughly every 9–12 months. But that synthesis is not in any of the source vendor reports.)
Contradiction #3 — AI Overview CTR Impact
| Study | Claim | Sample | Funding |
|---|---|---|---|
| Seer Interactive — Longitudinal CTR | Organic CTR fell 61% on AIO queries (1.76% → 0.61%) | 2.43B impressions, 53 brands | Vendor-adjacent (agency) |
| Ahrefs — December 2025 AIO study | 58% lower CTR for top-ranking page | 300K keywords | Vendor (sells SEO platform) |
| Pew Research — July 2025 AIO Study | Only 1% of users click links inside an AI Overview; 46.7% relative click reduction | Nationally representative US sample, ~900 panelists | Independent (academic-style nonprofit) |
| Agarwal & Sen — SSRN April 2026 | First causal randomized field experiment confirming CTR loss; effect size ~10–25% reduction depending on query class | Custom Chrome extension, randomized assignment | Independent (academic) |
Interpretation. Here vendor and independent studies agree on direction but differ ~3–6x on magnitude. Vendor longitudinal observational studies report 58–61% CTR loss; the only randomized field experiment (Agarwal & Sen, the first peer-reviewable causal evidence) reports a smaller effect. Vendors' observational designs likely overstate magnitude because they don't separate AIO presence from query selection (queries that trigger AIOs are systematically different from queries that don't). The honest read: AIOs cause CTR loss, but probably less than the vendor longitudinals report.
Contradiction #4 — LinkedIn as a Citation Source
| Study | Claim | Sample | Funding |
|---|---|---|---|
| Profound — LinkedIn Most-Cited Domain | LinkedIn went from #11 to #5 in ChatGPT citation rank in 3 months; 14.3% of ChatGPT responses cite LinkedIn | 1.4M citations | Vendor (sells citation-share tracker) |
| Semrush — 89K LinkedIn URLs Study | LinkedIn is #2 most-cited domain across 3 platforms | 325K prompts, 89K LinkedIn URLs | Vendor (sells SEO/GEO platform) |
| Axios — Independent Reporting on Profound's Data | Profound's claim repeated, no independent replication | n/a | Independent (journalism) |
| 5W AI Platform Citation Source Index 2026 | Reddit is #1 across every major engine at ~40%; LinkedIn not in 5W's top-15 consolidated list | 680M citations synthesized | Vendor-adjacent (PR firm aggregating multiple sources) |
Interpretation. Profound and Semrush — both vendors — report LinkedIn as a top-tier citation source. The 5W index, which synthesizes 680M citations across multiple studies, doesn't surface LinkedIn at the same prominence. The likely explanation: Profound's prompt set is heavily B2B and professional, where LinkedIn over-indexes; 5W's set is broader. Both are technically true; they describe different prompt populations. But the LinkedIn-is-suddenly-top-3 narrative spread across the trade press in March 2026 without that prompt-population caveat.
Bonus Contradiction — llms.txt
This one is settled, and it's the cleanest example of independent research correcting vendor enthusiasm.
| Study | Claim | Funding |
|---|---|---|
| Early vendor blog posts (2024–2025) | "llms.txt is the new robots.txt for AI" | Vendor |
| Reboot Online — 3-month controlled test | Zero AI bot visits to test pages with llms.txt | Vendor-adjacent (SEO agency, ran controlled experiment) |
| OtterlyAI — 90-day study | llms.txt = 0.1% of AI bot traffic | Vendor (but reported a null) |
| Search Engine Land — 9-site study | Zero or negligible impact | Independent (journalism) |
| Counter-evidence: dev5310 single-site anecdote | Positive | Anecdote (n=1) |
Interpretation. Vendor enthusiasm preceded controlled testing by ~12 months. Three independent multi-site studies all returned null. The vendor consensus shifted; today llms.txt is correctly classified as "implement only if free, don't invest time." This is what self-correction looks like — and it took ~9 months of vendor-funded studies to overturn.
6. The 30-Study Master Table
Notation.
- Funding: V = Vendor; VA = Vendor-adjacent; I = Independent; M = Mixed
- Method rigor: 1 (anecdote) – 5 (peer-reviewed)
- Direction: + Pro-GEO; − Skeptical; = Neutral
- Repro: Y / Partial / N
| # | Study (Author / Year) | Funding | Sample Size | Method Rigor | Headline Effect Size | Direction | Repro |
|---|---|---|---|---|---|---|---|
| 1 | Aggarwal et al. — GEO: Generative Engine Optimization (Princeton/Georgia Tech, KDD 2024) | I | ~10K queries across multiple LLMs | 5 (peer-reviewed) | Up to +40% visibility under specific GEO methods (statistics, citations, quotations) | + | Y |
| 2 | Agarwal & Sen — Google AI Overviews and Publisher Traffic (SSRN, Apr 2026) | I | Custom Chrome extension; randomized field experiment | 4 (controlled RCT) | ~10–25% CTR reduction (causal, not correlational) | − | Y |
| 3 | Pew Research — AI Overviews CTR Study (July 2025) | I | Nationally representative US panel (~900 users) | 4 (observational, methodology disclosed) | 1% click rate inside AIOs; 46.7% relative SERP click reduction | − | Y |
| 4 | Ahrefs — Top-10 AIO Citation Study (Apr 2026) | V | 863K keywords / 4M AIO URLs | 3 | Top-10 share of AIO citations dropped 76% → 38% in 6 months | = | Partial |
| 5 | Ahrefs — Brand Mentions vs. Backlinks Correlation (2026) | V | 75K brands | 3 | Brand mentions 0.664 correlation; backlinks 0.218 (3x) | + | Partial |
| 6 | Ahrefs — December 2025 AIO CTR Study | V | 300K keywords | 3 | −58% CTR for top-ranking page when AIO present | − | Partial |
| 7 | BrightEdge — Structured Data in AI Search | V | Enterprise client base, n undisclosed | 2 | +44% AI citation lift with schema + FAQ | + | N |
| 8 | BrightEdge — 9-Industry AIO Tracker (Mar 2026) | V | Enterprise tracker, sample undisclosed | 3 | 48% of queries trigger AIOs (highest of any tracker) | = | N |
| 9 | BrightEdge — AI Search Visits Surging (2025) | V | Enterprise customers | 2 | AI sessions convert at 7.05% vs. 5.81% (+21%) | + | N |
| 10 | Otterly — Schema Markup Real Impact Test (Q1 2026) | V | Cross-engine controlled test | 4 (controlled, with limitations) | No isolated causal effect; 6 of 7 engines couldn't fetch JSON-LD | − | Y |
| 11 | Otterly — llms.txt 90-Day Study | V | One site, 90-day window, 62K AI bot visits | 3 | 0.1% of AI bot traffic | − | Y |
| 12 | Otterly — 1M+ Citations Report 2026 | V | 1M+ citations | 3 | Community platforms = 52.5% of citations | = | Partial |
| 13 | Reboot Online — llms.txt Crawler Experiment | VA | 2 sites, 3-month controlled | 4 | Zero AI bot visits to test pages | − | Y |
| 14 | Reboot Online — Negative GEO Experiment (Apr 2026) | VA | Fictional persona, 11 LLMs | 3 | 2 of 11 LLMs surfaced unsubstantiated claims | = | Y |
| 15 | SE Ranking — Fake Brand Experiment (Apr 2026) | VA | 825 prompts, 15.8K AI answers, 12 domains | 4 (genuinely novel design) | 96% of AI visibility from branded searches; deep guides ~80x listicles per page | = | Y |
| 16 | Profound — LinkedIn Most-Cited Domain (Q1 2026) | V | 1.4M citations | 3 | LinkedIn #11 → #5 in ChatGPT citation rank in 3 months; 14.3% ChatGPT cite rate | + | Partial |
| 17 | Profound — Q1 2026 Enterprise AI Visibility | V | 700+ enterprise customers | 2 | 62% of mid-market lost AI visibility Q1 2026 | − | N |
| 18 | Semrush — 89K LinkedIn URLs Cited Study | V | 325K prompts, 89K LinkedIn URLs | 3 | LinkedIn #2 most-cited domain across 3 platforms; ~11% citation rate | + | Partial |
| 19 | Semrush — AI Visibility Index (Apr 2026 update) | V | 213M LLM prompts | 3 | Finance category: Fidelity 33.7% SOV; Vanguard 29.3% | = | N |
| 20 | Tinuiti × Profound — Q1 2026 AI Citation Trends Report | M | 50K+ AI responses | 3 | Reddit citations +73% in commercial verticals Oct 2025 → Jan 2026 | + | Partial |
| 21 | AthenaHQ — GEO Platform Showdown 2026 | V | 1,000 simulated buyer questions, 30 days | 2 | AthenaHQ +45% answer share; Profound −1% (self-test) | + | N |
| 22 | Conductor — 2026 AEO/GEO Benchmarks | V | 13,770 enterprise domains, 3.3B sessions | 3 | AIOs on 25.11% of searches; 70% AIO content changes per query | = | Partial |
| 23 | Seer Interactive — AIO CTR Longitudinal (Apr 2026) | VA | 2.43B impressions, 53 brands, 5.47M queries | 3 (observational, large-N) | CTR 1.76% → 0.61% then rebounded 85% to 2.4%; cited brands +120% clicks | − (then mixed) | Partial |
| 24 | EMGI — SaaS AI Citation Gap Report (Apr 2026) | VA | 150 SaaS brands × 120 keywords | 3 | 44% of Google top-10 SaaS brands get zero ChatGPT citations | − | Partial |
| 25 | Cloudflare Radar — AI Crawler Q1 2026 | I | All Cloudflare network, billions of requests | 4 (large-N infrastructure data) | ClaudeBot crawl-to-refer 23,951:1 | − | Y |
| 26 | Adobe — March 2026 AI Traffic Inflection | I (mixed) | Adobe Analytics customer base | 3 | AI traffic converts +42% better than non-AI; Q1 retailer AI traffic +393% YoY | + | N |
| 27 | G2 — The Answer Economy (Apr 2026) | V | 1,076 B2B decision-makers | 3 (survey) | 51% of B2B buyers begin research in AI chatbot (up from 29%) | + | Partial |
| 28 | Forrester — 2026 B2B Predictions | I | Forrester proprietary survey + analyst | 3 | 94% of B2B buyers use LLMs in journey | + | Partial |
| 29 | Press Gazette — Publisher Traffic Decline (Nov 2025) | I | Press Gazette publisher survey | 3 | Global Google referrals −33% YoY | − | Y |
| 30 | Kevin Indig (Growth Memo) — 1.2M-Response Analysis | VA | 1.2M LLM responses | 3 | 44.2% of citations from first 30% of page text; cited content 2x more likely to include question marks | + | Partial |
Summary breakdown:
- Vendor (V): 16 studies
- Vendor-adjacent (VA): 7 studies
- Independent (I): 6 studies
- Mixed (M): 1 study
Direction breakdown:
- Pro-GEO (+): 12
- Skeptical (−): 10
- Neutral (=): 8
Reproducibility breakdown:
- Y: 9
- Partial: 13
- N: 8
7. What This Means for B2B Marketing Leaders
If you're approving GEO budget, hiring a GEO consultant, or evaluating an AI visibility tool, three patterns from this study should change how you read the research being pitched at you.
1. Discount any single-vendor "lift" claim by ~50% as a default prior. Not because vendors are dishonest, but because the systematic 2.4x effect-size delta between vendor and independent studies is too consistent to be noise. The vendor number is probably real on their selected sample. It's probably not generalizable to yours at the same magnitude.
2. The contested findings are more useful than the consensus ones. Schema's effect, llms.txt's effect, AI conversion's true magnitude — these are the places where vendor and independent research disagree, and disagreement means the market hasn't priced the answer in. The unanimous findings (brand mentions > backlinks; freshness matters; format matters) are settled but already in every vendor pitch — there's no edge in betting on consensus.
3. Independent research lags vendor research by 6–12 months. The vendors get to the data first because they have the data. By the time SSRN, Pew, or an academic confirms a vendor finding, the vendor has moved on. If you wait for peer-reviewed evidence before acting, you'll be 12 months behind. If you act on every vendor claim, you'll waste 30–50% of your effort on inflated effects. The middle path: act on findings replicated by 2+ vendors with independent direction-confirmation; ignore single-vendor "first to publish" headlines until they replicate.
8. A Reader's Checklist — 5 Questions to Ask of Any GEO Study
- Who paid for it, and what do they sell? If the study's findings would increase demand for the publisher's product, the effect size is probably inflated. Read the methodology section, not the headline.
- What's the sample size, and how was it selected? "Enterprise customers" is not a random sample. "Our 75K-brand database" usually means brands the vendor already tracks — i.e., brands with marketing budgets large enough to have a detectable signal. Findings won't generalize to your SMB.
- Is the methodology disclosed enough to replicate? If a stat is everywhere but you can't find a methodology section anywhere, treat it as marketing copy. Real research includes a methodology paragraph; press releases include a quote from a VP.
- Is the direction confirmed by ≥1 independent or competitor source? Vendor-vs-vendor disagreement (Otterly contradicting BrightEdge on schema) is a signal that the question isn't settled. Vendor-vs-independent agreement is a signal it might be.
- What's the absolute base rate? A "393% YoY increase" on a 0.1% base is a 0.4% absolute share. A "+42% conversion lift" on 1% of traffic is 0.4% of total conversions. Always ask: "of what?"
9. Limitations
This study has six limitations I'd want anyone replicating it to know:
- I'm the sole classifier. A second independent classifier would shift several borderline cases (Adobe, Tinuiti × Profound, Reboot, Seer). The directional finding survives reclassification; the exact 2.4x delta is sensitive to ~3 borderline calls.
- Effect sizes are not comparable across studies. A correlation coefficient, a percentage lift, a share-of-voice, and a YoY traffic delta are not the same kind of number. I converted to absolute percentage where possible and excluded correlations from the mean — but the comparison is approximate.
- English-language and US/EU bias. The source files lean English-language and Anglo-American. Aleyda Solis's Spanish-language and international GEO research is underrepresented.
- Selection bias in my own pool. The 30 studies I chose are the ones cited most heavily in The Cited Club's source files. That's a defensible practitioner-relevance filter but it skews vendor (because vendors get cited more in the trade press).
- No correlation ≠ no causation handling. I treat "0.664 correlation" as a single data point in the table, but Ahrefs' brand-mention correlation has been the basis for at least 50 follow-on vendor blog posts that quietly shift it from "correlation" to "driver" — the field's correlation/causation hygiene is poor and I haven't tried to penalize that systematically here.
- Time bias. The vendor research cycle is faster than the academic one, so the "mean effect size" measurement at any single moment over-weights vendor-recent-news. Re-running this in 6 months with the same methodology would likely produce a smaller but still meaningful gap as more independent replications publish.
10. Reproducibility
Anyone can replicate this analysis from public sources by:
- Listing every quantitative AI-search study cited in the trade press between January 2025 and May 2026 (you will land on roughly 50 candidates).
- Filtering to studies published Jan 2025 – May 2026 with a primary source URL and ≥1 third-party citation.
- Classifying each by funding (V/VA/I/M), method rigor (1–5), direction (+/−/=), and reproducibility (Y/Partial/N) using the rubric in Section 2.
- Computing the mean absolute headline effect size by funding category, restricted to studies whose headline is a percentage lift, share, or absolute conversion delta.
- Comparing vendor-vs-independent contradictions on the same topic where ≥2 studies disagree on magnitude or direction.
I will publish the underlying classification CSV (one row per study, with all rubric scores and source URLs) as a follow-up. Pull requests and re-classifications welcome — direction-of-finding is the more durable result; the exact 2.4x multiplier is sensitive to sample selection and would benefit from triangulation.
11. Closing Note
The fix isn't to stop reading vendor research. The fix is to read it with the right prior. Pair every vendor "lift" claim with an independent base-rate check. Wait for replication before betting budget on a single-source finding. Watch the contradictions, because that's where the market is still figuring out what's real. And before you bet budget on any surface-level citation stat, read the companion study: The Surface Diet — which sources each AI engine actually eats.
If this study is useful, the next one in the series will trace how a single GEO stat propagates — picking one number (probably "+44% schema lift" or "0.664 brand-mention correlation"), and tracking every published claim that cites it back to the original methodology. The hypothesis is that ~70% of citations of any given GEO stat are to the headline, not the methodology, and that the headline drifts further from the methodology with each downstream cite.