AI NewsModels & agentsAnnouncement

Anthropic puts cyber safeguards on Sonnet 5.5, and API developers have to opt in to the Sonnet 5 fallback

Anthropic released Claude Sonnet 5.5 on Monday with the classifier-driven cyber safeguards and fallback routing that had only been running on Opus models before, and API developers who used to call Sonnet 5 have to enable fallback themselves or accept that a flagged request just stops.

AI News

Editorial3 min read

LinkedInX

Why it mattersA blocked request on Sonnet 5.5 does not reroute to Sonnet 5 on the API by default, so a team that treated the version bump as a plain replacement can now find requests failing that used to succeed.

An engineer changes claude-sonnet-5 to claude-sonnet-5.5 in the API call, releases the update, and starts seeing requests fail that used to succeed. That is the migration risk in the routing change Anthropic released with Sonnet 5.5 on Monday: the model that many teams use for production workloads is the first Sonnet to carry classifier-driven cyber safeguards, and the automatic fallback to Sonnet 5 that Anthropic's own apps use is not on by default on the API.

The system card shows why Anthropic put those safeguards on the model. With cyber safeguards turned off, Sonnet 5.5 achieved full arbitrary code execution in 178 of 410 ExploitBench runs. It completed 46.1% of the challenges on Irregular's CyScenarioBench, up from 0.7% for Sonnet 5, and managed 50 control-flow hijacks on a binary exploitation benchmark built on Google's OSS-Fuzz corpus, against three for Sonnet 5. Anthropic still considers the model less capable at cybersecurity than Opus 5.5 and Mythos 5.1, but the company says the improvement from the previous Sonnet was large enough to apply the same cyber policy it uses on Opus 5 and Opus 5.5.

How the block works

Enforcement runs in three stages. A probe reads the model's internal activations, a lightweight classifier runs on Sonnet 5.5 itself, and a separate trained LLM classifier weighs the probe's verdict when deciding whether to block a conversation. Anthropic says the classifiers catch harmful cyber requests at a rate comparable to those on Opus 5, but it also warns that Sonnet 5.5 users should expect more refusals than they saw with Sonnet 5, including on legitimate cybersecurity work. Blocks for biology, conventional weapons and anti-distillation end the request without any fallback model, and Anthropic says these blocks are transparent and do not change the response quietly.

The API fallback is an opt-in setting

Cyber blocks are the ones that can fall back. Higher-risk cybersecurity requests flagged by the safeguards route to Sonnet 5, and a narrow set of requests tied to frontier LLM development, such as kernel work on certain ML accelerators, also falls to Sonnet 5. Anthropic's own apps enable this automatically and show a notice when the model switches. On the API the fallback is an opt-in setting, and other platforms and providers may handle blocked requests differently. Without fallback enabled, a blocked request stops rather than being served by Sonnet 5.

The safeguards also apply to any content the model reads, including memory, connector output, web search results and files, so a page an agent pulls in can trigger a block on its own. Vulnerability discovery in source code is allowed, which keeps secure-coding workflows intact, but the same work on compiled binaries is not.

Fallback has a security cost

Anthropic's own indirect-prompt-injection tests inside coding environments show the cost of enabling fallback. 25% of Sonnet 5.5 requests were handed to Sonnet 5 after triggering a cyber block, often because injected instructions to wipe disks or delete files set off the classifier. Of the rerouted requests, 12.01% were successfully compromised. Of the 5,901 requests Sonnet 5.5 handled itself, four were compromised. Gray Swan, an AI security company, ran a separate indirect-prompt-injection benchmark and found no drop with fallback enabled, but Anthropic's coding-environment numbers say a team using fallback has to think about the older model's security as well as the new one's.

Anthropic says it is still tuning the classifiers to reduce false positives and plans to expand its Cyber Verification Program so verified defenders can use Sonnet 5.5 with fewer restrictions. Sonnet 5.5 is not in that program at launch.

Source

Primary: Anthropic Claude Sonnet 5.5 launch page and Sonnet 5.5 System Card (PDF). Reporting: The New Stack: You picked Claude Sonnet 5.5, but Anthropic may send your request to Sonnet 5 in "higher-risk" situations.

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX