AI NewsInfrastructureAnnouncement
Archestra open sources OpenAPPA, a deterministic agent guardrail that reports zero successful attacks on its own benchmarks
Archestra released OpenAPPA this week, an MIT-licensed guardrail that sits between an agent and its tools and decides from a policy whether each action is allowed, and the project reports zero successful attacks across 1,320 of its own evaluations.
Image: GitHub
Why it mattersArchestra points out that a probabilistic guardrail at 99.3 percent still fails thousands of times across a million agent calls, so a deterministic policy engine that answers in milliseconds is a design worth evaluating against the probabilistic prompt-injection filters most teams rely on today.
Archestra's pitch for OpenAPPA, released this week under MIT, opens with arithmetic: on the project site, the company says probabilistic approaches "top out at 99.3 percent", and that "at millions of calls, 0.7 percent is a lot of breaches." Its answer is a deterministic engine that answers yes or no from a policy before every agent action. The project is on GitHub at archestra-ai/OpenAPPA with 1.4k stars and 737 commits on the main branch.
OpenAPPA sits between an agent and the tools it calls, and the engine answers one question before every action: is this data allowed to go to this destination? InfoQ reports that the formal research is attributed to Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov and Matvey Kukuy at Archestra, in a paper titled "APPA: Recoverable Information-Flow Control for Real-World LLM Agents" accepted at NeurIPS.
How the engine works
Policies are written in declarative TOML. The engine "decides from the event log alone, with no network or file calls", in the words of the repository, and the decisions are deterministic, so the same inputs give the same output on every run. It deploys in process or as a sidecar, and ships with a Claude Code plugin for teams already using Anthropic's coding agent. The project is marked preview and the README notes that config and wire surfaces may break without migration shims.
What the numbers measure
The numbers Archestra publishes are its own measurements against its own two benchmarks. Bench-Corp is 20 multi-step enterprise workflows and AgentThreatBench covers the OWASP Top 10 for Agentic Applications, with 1,320 total evaluations. On those tests Archestra reports zero successful attacks for OpenAPPA, 10 percent for Claude Code's auto mode, and 31 percent for Microsoft FIDES. Task completion was 89 percent for OpenAPPA, 90 percent for Claude Code and 41 percent for FIDES. The paper also reports an ablation: with recovery strategies disabled, task completion fell to 35 percent, so the recovery logic carries a lot of the completion number.
Treat those figures the way any vendor benchmark should be read. Archestra is comparing itself to two named competitors on tests Archestra wrote, and that is marketing, not a third-party audit. The direction of the result is the useful signal: a deterministic policy engine that answers in milliseconds is a different design from a probabilistic prompt-injection filter, and Archestra's own arithmetic argument, 0.7 percent at scale, is the reason to look at it.
Source
Primary source: archestra-ai/OpenAPPA on GitHub and the project site at openappa.com. Reporting: InfoQ, "New Archestra's OpenAPPA Saturates Two Major Security Benchmarks with a 0% Attack Success Rate".
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.