AI NewsGo-to-marketAnnouncement
Google has deployed SAFE, a four-agent LLM system that reads AI content spam by the intent of a policy
Google published a research paper describing SAFE, a deployed multi-agent LLM system that identifies AI-generated spam by the intent of a platform policy.

Image: Search Engine Journal
Why it mattersSEO teams using AI to write at scale now face a live detector that reads for coordination and policy intent, so batch-produced pages with shared infrastructure and burst timing are the ones most at risk.
Google is now running a spam detector on AI-generated content that reads for the intent of a policy rather than the wording of a rule. On 25 September Search Engine Journal reported that Google has published a three-page research paper describing the system, called Scaled Abuse Forensics Examiner, or SAFE, and confirmed it has been deployed.
The paper is titled The Synthetic Gap: Automating Forensic Investigation of "AI Slop" with the Scaled Abuse Forensics Examiner (SAFE) and sits as a PDF on Google's research storage. It states that "Early deployment results indicate that SAFE significantly accelerates the identification of novel synthetic threats, reducing forensic investigation time compared to human-in-the loop workflows." The paper does not share the test results behind that claim.
SAFE is the second AI spam detector Google has revealed this year. The first, Scalable Cluster Termination System (S-CTS), was named earlier in 2026. Roger Montti of Search Engine Journal writes that these systems may be part of the September Spam Update, which Google announced on 22 September and is still rolling out.
Four agents, one orchestrator
The paper describes SAFE as four specialised AI agents working together. A Root Agent coordinates the investigation and reaches the final conclusion. A Content Understanding Agent uses LLM methods to detect signs of AI-generated abuse, including content that evades existing classifiers while still violating "the spirit" of platform policy. A Behavior Understanding Agent looks for coordination in timing, infrastructure and posting patterns, such as synchronised uploads and burst publishing. A Channel Cluster Understanding Agent uses a graph of relationships between content producers to map wider spam networks by shared infrastructure.
The paper calls this "spirit of policy" enforcement, "primarily with a few-shot-trained LLM." The goal is to catch content that violates the intent of a rule without matching any existing violation-detection model.
Detection moves to the network behind the page
Two things stand out for a site whose search traffic depends on Google. First, the detection point sits at infrastructure, timing and relationships between producers, which are harder to disguise than a rewritten sentence. Second, an LLM judge that reads for policy intent will catch content earlier classifiers passed, so a page that survived every previous spam sweep on wording is not safe by history.
Google is keeping the specifics quiet. The paper is only three pages, mentions running tests but publishes no numbers behind the "significantly accelerates" claim, and does not name the surfaces where SAFE is running. Search Engine Journal calls the disclosure "highly unusual."
For a site producing AI-assisted content at scale, the pattern SAFE is built to find is batch generation across a shared host, with near-simultaneous publishing schedules and similar internal linking. Individual pages that pass a quality check on their own can still land inside a cluster the graph agent maps.
Source
- Search Engine Journal: Google Has Deployed A New AI Spam Detector Called SAFE, by Roger Montti, 25 September 2026.
- Google Research: The Synthetic Gap: Automating Forensic Investigation of "AI Slop" with the Scaled Abuse Forensics Examiner (SAFE) (PDF, 3 pages).
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


