How Should Publishers Judge the Evidence in a GEO Research Paper?
A practical way to assess GEO papers, white papers and reports without stretching a measured finding into a promise of AI visibility.

A GEO paper is useful to a publisher only when its conclusion survives a narrower question: What did the researchers change, what did they measure, and where did they test it? Start there, rather than with the headline percentage. Your decision may be to make no change, investigate a plausible idea, or run a limited test on your own site.
Rewrite the claim in one sentence
Try this template: “Under [tested conditions], [change or characteristic] was associated with or caused [measured difference] among [tested items].” Leave a blank wherever the paper does. Those blanks are questions for the authors or reasons to limit the conclusion—not invitations to fill in familiar assumptions.
Separate three statements that often sit close together in a report:
- Description: Some tested answers contained a particular type of source. This describes the collected answers.
- Association: Pages with a feature appeared more often in the tested answers. Other differences between those pages could also matter.
- Causal claim: Changing that feature produced a difference against a suitable comparison under the tested conditions. Ask how items were assigned, what else changed, and whether both groups faced comparable prompts and collection conditions.
A recommendation such as “publishers should add this feature” takes another step. Even if a test supports a causal claim in its setting, the recommendation still needs to fit your pages, audience and desired outcome.
Record who wrote and funded the work, whether the authors sell a related service, and whether the document is a reviewed publication, preprint, vendor report or presentation. These details help you decide which methods and data to scrutinize. None is a substitute for reading the methods: a commercial interest does not make a calculation wrong, and a publication label does not tell you whether the experiment matches your site.
Identify the population and the route to an answer
Make a small test-conditions record before applying the result:
| Question | What to record |
|---|---|
| What was sampled? | Where the prompts, pages or passages came from; how they were selected; and how many were excluded |
| What topics were covered? | Subject areas, page types, languages and locales |
| What produced the answers? | Named engine or model, version if available, product mode and collection dates |
| What material could it use? | Supplied passages, a fixed collection, or a live retrieval process; any documented access restrictions |
| What was scored? | The unit of analysis, scoring rule, denominator and treatment of missing or ambiguous answers |
The distinction between supplied passages and live web discovery is consequential. If a test hands an answer system passages to work with, it can examine what happens after those passages are available. It cannot, by that design alone, show that editing a public page would cause the page to be found. Likewise, a change in whether a source is used in a generated answer is not automatically a change in site visits.
Look for a diagram or methods description that follows the material from selection through scoring. If the collection dates, mode or retrieval conditions are missing, write “unknown” in the record. Do not silently substitute the product behavior you happen to see today.
Trace the improvement back to its denominator
Suppose a hypothetical report says an edit raised a source-appearance rate from 40 of 200 tested answers to 50 of 200. That is 20% versus 25%: an increase of 5 percentage points, or 25% relative to the original rate, representing 10 additional appearances in this hypothetical sample. All three descriptions refer to the same counts; none says that referrals rose by 25%.
Before using such a figure, check whether the denominator is prompts, answers, citations, sources or sites. Were the same prompts used before and after? Were repeated answers from one prompt counted as separate observations? Were failed or empty answers excluded? If a score combines several outcomes, find the formula and the underlying counts. A large relative change from a small baseline can sound more decisive than the number of affected observations warrants.
Keep the outcomes separate in your notes: page availability, appearance in an answer, citation, brand mention, recorded referral and attributed conversion are different observations. To claim an effect on visits or conversions, the study would need to measure those outcomes with an appropriate comparison; an answer-appearance metric cannot stand in for them. For a closer look at what a visibility report records, see what AEO and GEO tools can actually measure.
Decide what the design permits you to do
For an observational comparison, list plausible alternative explanations: perhaps the compared pages differed in subject, reputation, length or selection into the dataset as well as in the feature being studied. For an intervention, ask whether the change was isolated, whether a comparable group remained unchanged, and whether scoring was applied consistently. If multiple edits were made together, the result cannot isolate which edit mattered.
Finish with a short decision record:
- Verifiable: The exact claim, dataset, conditions, comparison, metric and arithmetic you can inspect.
- Uncertain: Missing methods, unresolved alternative explanations, or a mismatch between the test setting and your publishing context.
- Action: No change, a request for more detail, or a bounded local test with a defined outcome and stopping point.
A bounded test might ask whether a clearly specified content revision changes an outcome you can actually record for a selected set of pages and questions. Define that outcome before looking at results, retain the comparison, and keep any conclusion within the conditions you tested. The paper can help you choose a question worth testing; it cannot promise the answer your site will get.




