Skip to content

The Vendor Bias Study: read the method before the headline

Before a research headline becomes a budget decision, check what was actually measured. A corrected evidence review and a practical checklist for buyers.

Shubham Bansal, founder of The Cited Club

· Founder

Published Updated 13 min readAgent Markdown

What this review can tell you

A founder reading a GEO proposal needs to make a fairly ordinary decision: is this work worth doing, and is this the right team to do it? GEO means generative engine optimization — work intended to improve how a business appears in AI-generated answers. A dramatic research headline can make the buying decision feel simpler than it is.

The problem starts when a narrow observation becomes a broad commercial promise. A report can describe the sources in a collection of answers. It cannot, from that observation alone, establish how much revenue a new customer will earn by buying the publisher's software. Those are different questions, with different evidence requirements.

This review gives you a way to keep those questions separate. It uses a small selection of named primary sources to demonstrate how to read a claim. The selection is illustrative, not exhaustive. It does not estimate how common bias is, rank research publishers, or establish that commercial funding causes exaggerated results.

The Cited Club sells GEO services, so we have a commercial interest in this subject too. Apply the same questions to our work. A source link should let you inspect a claim, not ask you to trust the person who added the link. The correction above is part of that responsibility.

Why the original comparison was invalid

The withdrawn version grouped studies by publisher affiliation and averaged percentage headlines. Some headlines described traffic growth. Others described a share of citations, a change in click-through rate, or a conversion difference. Making every number positive did not make the underlying measurements comparable.

A share answers how much of a defined total belongs to something. A relative change answers how far a measurement moved from its baseline. A percentage-point change answers the difference between two rates. Each can be written with a percent sign, but that shared notation does not make them observations of the same outcome.

The classification also mixed two decisions: identifying a publisher's commercial interests and deciding whether a finding sounded favourable to GEO. A result about reduced publisher clicks is not a test of whether an agency's service works. Labelling it favourable or sceptical adds an interpretation that needs its own justification.

Removing extreme values would not repair those problems. A smaller average of incompatible quantities is still incompatible. Nor would recruiting another classifier rescue the pooled effect without first defining a common outcome and a defensible selection method. We therefore removed the aggregate rather than replacing it with a different multiplier.

This changes the conclusion. We cannot use the original sample to say how much vendors overstate results, how frequently independent work disagrees, or how far a buyer should discount a commercial claim. In particular, there is no defensible blanket rule to halve every vendor's number.

Methodology for this corrected version

We checked the claims against identifiable primary publications and their stated methods. Where a source provided an observation, we retained the observation with its population and denominator. Where the available material did not establish a claim, we removed the claim. An inaccessible full paper is not treated as a completed full-paper review.

The evidence examples below serve different purposes. Pew provides an observational browsing comparison. The GEO paper describes a research benchmark. Agarwal and Sen describe a field experiment in their public abstract. Google's documentation explains requirements for its own search features. None is pooled with another to manufacture a single effectiveness score.

The buyer checklist and proposed test protocol are TCC's recommendations. They are not findings from a new experiment. We have not replicated the cited studies, audited their underlying participant data, or established a complete census of GEO research. Source links and access limitations are included so readers can distinguish inspection from replication.

Use this article to evaluate the next claim in front of you. For channel-planning questions, the companion Surface Diet review explains how to keep different products and citation measurements separate. For the actual work involved, see our practical GEO guide.

Read each finding at the scale it supports

An observational click comparison

Pew Research Center's July 2025 analysis used March 2025 browsing activity from a US panel. Traditional-result clicks occurred on 8% of visits with an AI summary, versus 15% without one. This is an observed association, not a randomized estimate of the summary's causal effect.

Visits with an AI summary8%

Traditional-result clicks

Visits without an AI summary15%

Traditional-result clicks

What the chart shows: Traditional-result clicks occurred on 8% of visits with an AI summary and 15% without one in Pew's sample. The comparison is observational.

fig. 1Traditional-result clicks in Pew's browsing samplePew Research Center, July 2025; browsing observed in March 2025

For a buyer, the distinction is between a useful warning and a budget forecast. A channel-level observation can justify inspecting your own discovery paths. It does not tell you which of your pages will lose visits, whether the missing visits would have become customers, or which proposed intervention would change that outcome.

A benchmark visibility result

Aggarwal and colleagues' GEO paper, accepted to KDD 2024, reports visibility improvements of up to 40% in its evaluation and says effectiveness varies by domain. That is a benchmark result about visibility in generated responses, not a measured revenue lift for a typical agency customer.

The practical question is whether the proposed work resembles the intervention that was evaluated. A salesperson citing a paper about changing source content still needs to explain what they will change on your site, how that change maps to your buyers, and how they will measure its effect. The paper is context, not a substitute for that plan.

A field experiment with an access limitation

Agarwal and Sen's SSRN abstract, revised July 8, 2026, reports a randomized Google AI Overviews experiment and a 39.8% reduction in outbound organic clicks conditional on an overview appearing. We verified the publicly indexed abstract, not the full paper or underlying data. The earlier article's smaller range is withdrawn.

Do not silently combine that estimate with an observational comparison. Before comparing two results, inspect whether they measure the same action, use the same unit of analysis and cover comparable conditions. A difference between headlines can arise from those design choices without establishing that either publisher manipulated the result.

A platform requirement, not a lift study

Google's guidance for AI features says normal SEO practices remain relevant and no special schema or new machine-readable file is required for AI Overviews or AI Mode. This is guidance about Google's features, not an experiment showing that a particular technical change increases citations.

Ask a provider to distinguish requirements, good maintenance and experimental tactics. Fixing an inaccessible page is a concrete technical task. Guaranteeing that the fixed page will appear in an answer is a separate claim. The first can be checked directly; the second depends on systems the provider does not control.

Five questions before you approve the budget

1. Who published it, and what do they sell?

Start with the author, organization and any disclosed sponsor. Note whether the company sells measurement, content production, implementation, advertising or consulting. That context helps you notice when the proposed solution closely follows the publisher's commercial offer. It does not establish that the data is wrong.

Apply the same test to an independent-looking article. A newsletter can have sponsors. A consultant can sell advice without selling software. An academic affiliation does not remove the need to inspect methods. Where funding is not disclosed, record it as unknown instead of assigning a category from the author's tone.

The useful procurement question is whether you can inspect the evidence without accepting the sales conclusion. A provider should be comfortable explaining which parts of a result are measured and which parts are their recommendation. If those keep changing during the discussion, ask for a written claim with a source and a boundary.

2. Who or what was included in the sample?

Ask whether the sample consists of prompts, answers, websites, visits, people or customers. Those units answer different questions. A large count of citations can come from repeated answers to a relatively narrow set of prompts. It does not automatically describe the range of questions your customers ask.

Then inspect selection. Were the questions chosen by customers, inferred from search terms, or written to feature known brands? Were unsuccessful answers retained? Did the report include small local businesses as well as recognizable software companies? A sample can be useful without being representative, provided the conclusion stays within its limits.

For your own brief, write down the buyer situation before looking at results. An accountant serving local businesses needs a different question set from an enterprise accounting platform. Geography, service boundaries, price sensitivity and switching concerns can all belong in that brief. They should not be added only after a favourable answer appears.

3. What exactly does the number count?

Ask for a numerator and denominator in plain language. For citation share, the numerator could be links to one domain; the denominator could be all links in the sampled answers. For answer coverage, the numerator could be answers mentioning your company; the denominator could be all eligible answers. These are not interchangeable rates.

A report should also distinguish a mention from a recommendation. An answer can name a company while describing a reason not to choose it. It can cite a page for an unrelated factual detail. Neither event, by itself, establishes a favourable shortlist position or a qualified lead.

When a provider uses a proprietary score, ask which observations feed it and what happens when the prompt set changes. The score can help organize work, but you need an intelligible path back to the answer, source and buyer question. Otherwise a change in the score is difficult to turn into a useful decision.

4. What is the comparison?

Look for a baseline, a comparison group and a defined period. A before-and-after chart can be informative, but other changes may have happened at the same time. A redesign, paid campaign, new product launch or tracking change should not disappear from the explanation because the report is about GEO.

Ask whether the same measurement process was used throughout. If the provider changed models, introduced new prompts or improved citation extraction, the later count may cover more than the earlier one. That can be a good product improvement while still breaking a direct trend comparison.

You do not need a perfect experiment before making every small improvement. You do need the confidence of the claim to match the design. A directional observation can justify a limited test. A contractual promise about incremental revenue requires much more than a favourable screenshot and an upward line.

5. What would make this useful for our company?

A finding becomes actionable when it points to a specific buyer question, a missing piece of evidence and a feasible change. Ask what will be produced, who will publish it, what access is needed and when it will be checked again. That is a stronger basis for a proposal than a broad statement that AI search is growing.

Agree what success means before work starts. Correcting inaccurate company information is useful even if it does not immediately produce a new lead. Winning more relevant citations is a different outcome. Improving qualified enquiries is another. Report them separately so progress on one does not quietly stand in for success on all of them.

Also define a stopping or revision rule. If the planned work is shipped but the expected pattern does not appear, the next step should be investigation, not a new definition of success. A useful partner can explain what they would change their mind about and what evidence would prompt that change.

Keep the commercial maths honest

A higher conversion rate is not the same as a large contribution to the business. You need both the rate and the number of visits it applies to. You also need a stable definition of conversion: a newsletter signup, an enquiry and a paying customer should not share one label in a commercial argument.

Consider a deliberately hypothetical example, not a TCC result or industry benchmark. A site receives 100 AI-referred visits and 9,900 other visits. Suppose 10 AI visits and 198 other visits convert. The AI conversion rate is 10%; the other rate is 2%. Yet AI contributes 10 of 208 total conversions, approximately 4.8%.

The higher rate is interesting. The absolute contribution is still small. Neither fact tells you whether the next batch of AI visits will behave the same way. The practical response is to inspect lead quality and acquisition cost, then test whether that channel can grow without assuming its initial rate will hold.

Keep incrementality separate too. A buyer who first learned about a company elsewhere may later arrive through an AI referral. Attribution records a path; it does not automatically establish that the referral created a sale that would otherwise never have happened. Report known paths honestly and label what the tracking cannot determine.

A test brief you can use with a provider

The following is a proposed workflow, not a study we have run. Its purpose is to make a modest engagement inspectable. It is intentionally narrow enough that a small business can understand the work without becoming a research department or operating a second analytics stack.

Begin with one commercial question family, such as choosing a supplier, understanding fees or switching from an existing provider. Record the exact prompts and why those questions matter. Separate branded questions from unbranded discovery questions. A company appearing when its name is in the prompt is a different event from being introduced to a buyer who has not named it.

Capture a baseline on the actual products included in the brief. Retain the answer text, visible links, run date and relevant settings. Record failed runs separately from successful answers that omit the company. Do not turn a timeout into a negative visibility result, or silently replace it with a more convenient response.

Choose one coherent intervention. That could be clarifying a service page and publishing the missing fee explanation, with any technical access problems fixed first. Record the pages changed and the date each became public. If several workstreams change at once, acknowledge that the next observation will not isolate the contribution of each one.

Recheck the same question family using the same collection rules. Keep exploratory questions in a separate group so adding easier prompts cannot inflate the trend. Inspect the answers as well as the counts: the company can gain mentions while being described inaccurately, or lose one link while gaining a more useful recommendation.

Close the review with a decision. Continue the work, correct an issue, expand the sample or stop the tactic. Include what shipped, what changed, what remained uncertain and who owns the next action. The report should make the next month easier to manage, not leave the founder translating a visibility score into a production brief.

What good evidence sharing looks like

Ask for an evidence record you can read without a sales call. At minimum it should identify the claim, source, collection period, measurement definition, important exclusions and the action being proposed. A chart should have enough context nearby that a screenshot of it does not turn a bounded observation into a universal statement.

For original work, ask whether representative raw answers can be inspected. Sensitive customer information may need redaction, and a public dataset is not always appropriate. That does not prevent a provider from explaining its method or showing how a reported count maps to retained observations under suitable access controls.

For secondary research, follow the link to the named publication. A link to a publisher's homepage is not a useful citation for a particular percentage. Check whether the report has been revised and whether the quoted finding belongs to the current version. If access is limited, say what was actually available for review.

Corrections should travel beyond the article body. Titles, summaries, charts, FAQs, structured data and downloadable versions can all repeat a claim. Leaving the old number in a share card after changing the paragraph creates two conflicting versions of the evidence. This revision removes the withdrawn findings from those connected surfaces too.

Limits and the next decision

This corrected review does not quantify vendor bias. It does not identify a best research publisher, prove that GEO works for every business or show that all studies disagree. It explains why several different kinds of evidence should not have been collapsed into one headline, and gives buyers a way to inspect future claims.

Its recommendations still require judgment. A local business with a small marketing budget may reasonably start with an obvious technical fix and a useful page. A larger investment in a new distribution channel deserves a more formal baseline and review. Match the cost of validation to the cost and reversibility of the decision.

The useful question is not whether a report makes GEO sound exciting. It is whether the evidence supports a concrete next move for your company. If you want help identifying that move, start with an audit. Ask us for the buyer questions, the supporting answers and the work we would prioritize — then apply this checklist to our recommendation.

Questions this study answers

Does vendor funding make a GEO study wrong?

No. Funding is relevant context, not a verdict on the result. Inspect the sample, measurement, comparison and limitations. Apply the same standard to agencies, independent commentators and TCC.

Why was the original funding comparison withdrawn?

It averaged unlike percentage measures and contained inconsistent classification counts. Those problems invalidate the aggregate. This corrected article does not quantify vendor bias or recommend a blanket discount for vendor claims.

What is the difference between citation share and a business result?

Citation share describes a source's portion of counted links in a defined answer sample. It does not establish visits, qualified enquiries or revenue. Each commercial outcome needs its own measurement and attribution limits.

How should a buyer vet a GEO benchmark?

Ask who published it, what was sampled, what the number counts, what the comparison is, and why it applies to your company. Request the original source and underlying counts before turning the headline into a forecast.

Sources & further reading

Every number in this study is attributed inline. The register below lists the primary datasets and reports it draws on.

  1. [1]Pew Research Center — Google users and AI summary clicks (July 2025) Observational US browsing comparison; not a randomized causal estimate.
  2. [2]Aggarwal et al. — GEO: Generative Engine Optimization (KDD 2024) Benchmark visibility research; not a forecast of client revenue.
  3. [3]Agarwal and Sen — AI Overviews field experiment (SSRN, revised July 2026) Publicly indexed abstract checked; full paper and underlying data not reviewed.
  4. [4]Google Search Central — AI features and your website Official guidance on Search eligibility and technical requirements.

From insight to execution

Make your company easier for AI and search to choose.

The Growth Retainer connects AI visibility research with GEO, SEO, content, technical fixes, authority, and ongoing measurement. Start with the service model or see your current gaps first.

Shubham Bansal, founder of The Cited Club

Written and reviewed by

Shubham Bansal

Founder of The Cited Club. Shubham leads the research and delivery behind its GEO, SEO, content, technical, and authority work.