AI NewsInfrastructureReported

Red Hat ran nine AI guardrails on the same tests and a 200M-parameter classifier matched a 35B model on prompt injection

The New Stack reports that Red Hat's AI Safety team ran nine guardrails on the same benchmarks; a 200 million parameter classifier on a laptop CPU matched a 35 billion parameter LLM on prompt injection and returned the answer in 54 milliseconds.

AI News

Editorial2 min read

LinkedInX

Why it mattersA guardrail that runs on a laptop CPU inside the request path is a different cost structure from a call to a 35 billion parameter model on remote GPUs, and the benchmark says you do not pay accuracy for the switch.

A guardrail call sits in front of every prompt the model sees, so its cost adds to every reply the user waits for.

The New Stack reports that Red Hat's AI Safety team put nine guardrail setups through the same two benchmarks, prompt injection and content safety, and published the figures on 2 October 2026. Amanda Caswell at The New Stack covered the result on 5 October. The underlying post, by Dr. Rob Geada, Dr. Mac Misiura and Shelton Cyril at Red Hat, ran every setup through NVIDIA's open-source NeMo Guardrails toolkit.

The prompt-injection result

Qwen3.6-35B used as an LLM judge led the prompt-injection test at 89.31 percent accuracy. Red Hat's DeBERTa-based classifier, which has about 200 million parameters, came in second at 89.01 percent. The decision model Jev, from TypeSafe AI, scored 86.35 percent.

Latency separated the three. Red Hat reports a median of 54.1 milliseconds for DeBERTa, 312.5 for Qwen, and 348.1 for Jev. DeBERTa ran on a MacBook Pro M1 CPU; Qwen ran through vLLM on an AWS node with 96 GB of VRAM; Jev was called through TypeSafe's API. The network hop to Jev inflates its number, so Red Hat subtracted an estimate for it and still had Qwen several times slower than DeBERTa on commodity hardware.

Content safety flipped the order

Jev topped content safety at 86.20 percent, followed by DiffusionGemma, an open decision-style model served through vLLM, at 85.53 percent, and Qwen at 85.47 percent. Red Hat's own 125-million-parameter Granite Guardian came sixth at 80.27 percent, about six points behind Jev, but was the fastest option at a 33.2 millisecond median.

Red Hat plans to ship both its classifiers, DeBERTa for prompt injection and Granite Guardian for content safety, as the default guardrails in OpenShift AI 3.6. The team notes that this gives it a stake in how its own models compare.

One policy change moved the result by 18 points

Nemotron-3.5-Content-Safety's prompt-injection accuracy went from 69.37 to 84.84 percent when Red Hat replaced NVIDIA's default risk definitions with its own. Laya, an open decision model with about 421 million parameters that Red Hat ran on a laptop CPU, scored 57.87 percent on content safety under Red Hat's original policy and 75.20 percent after the team wrote a policy for it. The same tuned policy lowered Jev's content-safety accuracy from 86.20 to 82.53 percent.

A policy that recovered 18 points for one model cost another 3.67 points on the same benchmark, so any single leaderboard position depends on who wrote the risk definitions and for which model.

The authors' conclusion, printed on the Red Hat page, is to pick the tool to the risk. A small task-specific classifier is the stronger default for a well-defined risk with labelled training data. A decision model or LLM judge pays for its overhead on broader policies where no strong classifier exists. The test ran on English-language datasets, so Red Hat says the accuracy results may not carry over to multilingual guardrails.

Source

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX