Illinois researchers audit Reddit Answers on 30,000 questions and find formal, already-popular comments dominate the summaries

Image: Search Engine Journal
Why it mattersA brand or contributor whose posts are casual, first-person, or newly written is much less visible inside a Reddit-grounded answer than the raw upvote counts on the platform would suggest.
Search Engine Journal reports on 17 September that researchers at the University of Illinois Urbana-Champaign have audited Reddit Answers, the generative search product built on top of Reddit content. The paper, "The Wisdom of the Loudest: A Large-Scale Audit of Generative Search on Reddit", was posted to arXiv on 13 September by Agam Goyal, Wang Claire and Eshwar Chandrasekharan, and is a preprint that has not yet been peer-reviewed.
The setup
The team submitted 10,000 questions to Reddit Answers, three times each roughly five hours apart, across 20 advice and support subreddits: ten large ones such as r/personalfinance and ten smaller ones such as r/UKJobs. That produced 30,000 answers, drawn from 14.68 million comments. The abstract says the goal was to see "which community voices survive retrieval and synthesis".
The three numbers a working GEO team should carry
Reddit Answers heavily favours highly upvoted comments. Search Engine Journal reports, drawing from the paper, that the median selected comment sat at the 91st percentile for score within its thread, and that 92% of the direct replies used in answers came from highly upvoted comments.
Formal, directive language was more likely to be picked, experiential language less so. Search Engine Journal cites the paper's odds ratios: formal language had 1.488 to one odds of selection, a 49% lift, while experiential voice had 0.789 to one odds against.
First-person singular language was cut back further during the summary step. Search Engine Journal reports that first-person share fell from 3.3% in the comments Reddit Answers quoted to 0.06% in the final synthesised answer. So the retrieval step already picks against experiential voice, and the summary step then removes even more of what remains.
What the authors say about the mechanism
The abstract states that differences between the three repeated runs were "driven primarily by retrieval", so run-to-run variance comes from which comments Reddit Answers surfaces in the first place. The authors also say answers "routinely combine evidence across communities" rather than staying inside the community the question is asked in. The paper labels itself observational rather than causal, and does not claim a mechanism, only the pattern.
The consequence for a team optimising for Reddit answers
Reddit is now one of the most-cited sources across generative answer surfaces, so the shape of what Reddit Answers itself does is close to a specification for how a brand or contributor gets into an AI-cited paragraph on someone else's product. Casual first-person accounts that read as authentic on Reddit itself are structurally penalised the moment retrieval and synthesis run over them. A contributor whose value is a real story, and a brand relying on organic community advocacy, both need to accept that the version of that story that reaches a summary strips the voice out. Practically, a comment that leads with a clean, formal claim and earns upvotes on the top-level reply is the shape that survives, and comments buried under a long thread do not.
Source
Reporting: Search Engine Journal, by Matt G. Southern. Primary source: The Wisdom of the Loudest: A Large-Scale Audit of Generative Search on Reddit, by Goyal, Claire and Chandrasekharan, arXiv preprint, 13 September 2026.
Reported by: Search Engine Journal
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually shipping with them. Short, and only when there is something worth reading.


