Go-to-market

A causal audit of an agentic search engine finds the raw 42.3 point gap between rank 1 and rank 5 shrinks to 0.0 points once you control for what the pages actually say

September 17, 2026 at 11:10 AM PT

Featured image from the Search Engine Journal report on the CITECHOICE study of agentic search citations

Image: Search Engine Journal

Why it mattersFor GEO work, position in a retrieval call is far less important than how a page is written, but the effect of structure is redistribution of credit among pages already cited rather than getting a new page cited for the first time.

Search Engine Journal reports on 17 September that Sriram Selvam and Anneswa Ghosh have posted an arXiv preprint titled "CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search". The paper, dated 14 September and not yet peer-reviewed, audits one agentic search product: a GPT-5.4 agent using Exa as its retrieval provider, replayed offline with no live web edits.

What the paper actually tested

The authors took 129 real multi-turn transcripts, and from those pulled 113 pairs of documents inside the same search call that independently supported the same pre-specified fact. A blinded human reviewer confirmed 103 pairs. For each pair, they replayed the saved conversation four ways: putting one page above or below the other, and showing that page either as plain prose or rewritten with headings and lists. Everything else in the transcript was frozen, and only the final answer was regenerated. Both prose and structured versions were AI rewrites of the same source, produced by Grok 4.3, with a separate Grok pass checking that the facts still matched.

The three numbers a working GEO team should carry

Raw position looks huge, controlled position is zero. In the saved transcripts, pages sitting in the first Exa result were cited 85.1% of the time and pages in the fifth were cited 42.8%, a raw gap of 42.3 percentage points. When the researchers moved the same page up inside its pair, the citation chance rose by 7.9 percentage points, and that lift was not statistically significant after adjustment. On a second set of 56 pairs where only the order was switched, the estimate was 0.0 percentage points with a 95% confidence interval from -5.4 to +5.4.

Structure raises the count for the page that has it, and does not raise the total. Rewriting a page with headings and lists gave it, on average, 0.50 more citation markers per answer, with a 95% CI from +0.20 to +0.84 and a Holm-adjusted p-value of 0.033. The total number of citations per answer did not go up, and the number on the other page barely changed. The authors read this as citation credit being redistributed toward the structured page rather than more pages being admitted overall.

Whether a page is cited at all is noisy. The pre-planned test for the structure effect on incidence, meaning does the page get cited at all, was +4.5 percentage points with a 95% CI from -1.4 to +10.4 and was inconclusive. Rerunning 120 responses on the same inputs changed the binary decision for 15% of cases.

Where the practical value lands

Search Engine Journal quotes the paper as calling this "an attribution-sensitivity warning". A raw benchmark that says pages at rank 1 get cited twice as often as pages at rank 5 is measuring page quality alongside position, since retrieval usually orders by relevance. On a like-for-like reorder inside one pair, the ordering effect on this one agent was zero.

The structure result gives a working page a clearer aim: moving prose into scannable headings and lists shifts citation credit toward that page without expanding the pool of pages cited overall. A rival page with the same facts and flatter formatting will get less of the credit on the same answer.

Source

Reporting: AI Citation Test Finds Source Order Matters Less Than It Looks, by Matt G. Southern at Search Engine Journal, 17 September 2026. Primary source: CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search, by Sriram Selvam and Anneswa Ghosh, arXiv preprint 2609.15164, 14 September 2026.

Reported by: Search Engine Journal

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

More from AI News

Illinois researchers audit Reddit Answers on 30,000 questions and find formal, already-popular comments dominate the summaries

A University of Illinois preprint audited Reddit Answers with 10,000 queries repeated three times and reports that formal, already-upvoted comments were 49% more likely to be picked, while first-person voice fell from 3.3% of quoted comments to 0.06% in the final answer.

Source: PressGo-to-market

Ahrefs tracked 963 French domains before and after AI Overviews launched, and the most exposed lost 23.1 percent of their clicks

Ahrefs matched Google Search Console data on 963 French domains across the 28 days before AI Overviews launched on 22 July 2026 and the 9 days after, and found the median domain lost 5.7 percent of its clicks while the most exposed lost 23.1 percent.

Source: PressGo-to-market

Zeeshan Yaseen ran two GEO experiments, and third-party listicles produced 85.8 percent of AI citations

Two hand-tracked GEO experiments by consultant Zeeshan Yaseen found third-party listicles produced 85.8 percent of AI citations, with owned listicles at 14 percent.

Source: PressGo-to-market