OpenAPPA is a new MIT-licensed permissions layer for AI agents, and it stopped every attack in its author's own AgentThreatBench run
OpenAPPA, an MIT-licensed permissions layer for AI agents that sits between an agent and its tools and refuses actions when the data was not cleared for that destination, was posted to Hacker News as a Show HN on 28 September and drew 23 points and 11 comments.
Image: GitHub
Why it mattersA team that runs an agent with real access to email, storage or paying tools now has an open-source layer to try that gives the same answer on the same request every time, and does not depend on the model to police itself.
Every team building with agents runs into the same problem: a model told to book a hotel with the company card can be steered by a hidden line in a webpage into buying flights instead. OpenAPPA is a new MIT-licensed library that puts a fixed rule between the model's next action and the tool call, and refuses the call when the data being passed is not cleared for that destination.
The project was posted to Hacker News as a Show HN on 28 September 2026 by user motakuk and reached 23 points and 11 comments that day. The GitHub repository at archestra-ai/openappa has 34 stars and 5 forks, per the GitHub API on 29 September, and was created on 18 August. The README calls the project a preview and an RFC, and says the config and wire surfaces may still change.
What the rule is
OpenAPPA implements what it calls the Agentic Permissions Policy Algebra: every piece of data an agent touches carries a label for how sensitive it is and how far it can be trusted, and every tool call is checked against those labels before it runs. The label can only get more restrictive, never less, and the decision is derived from the label rather than asked of the model. The README calls this "deterministic" because the same request gets the same answer every time, no matter what a prompt says.
The benchmark, and what it measures
The repository ships two benchmarks. Bench-Corp is 20 multi-step enterprise workflows run 200 episodes per model, and measures task completion. AgentThreatBench is a 24-task suite the README pairs with OWASP's Top 10 for Agentic Applications, and measures Attack Success Rate across 720 attack attempts on three models: GPT-5.6 Luna, DeepSeek V4 Flash and Gemini 3.7 Flash.
Archestra AI reports that OpenAPPA blocked all 720 attacks while keeping task completion at 89 percent, against Claude Auto at 90 percent completion with 10 percent of attacks succeeding, and FIDES at 41 percent completion with 31 percent succeeding. These are the project's own numbers on its own benchmark and no third party has reproduced them. The framework plugs into an existing agent loop in one place, and the README lists Claude Code, kAgent and Archestra as integrations already working.
Where the pitch stops
A permission layer only helps if the labels are correct in the first place, and OpenAPPA does not remove the work of writing them. The README acknowledges it is a preview and an RFC, so a team adopting it now is signing up for surface changes between versions. There is no independent reproduction of the benchmark, and Show HN points at 23 are real traction but modest. A team running paid tools from an agent still gets a real win: the decision to allow or refuse a call is the same every time, and does not depend on the model noticing the trick.
Source
- OpenAPPA GitHub repository, primary source: https://github.com/archestra-ai/openappa
- OpenAPPA project site: https://www.openappa.com/
- Show HN thread on Hacker News, 28 September 2026: https://news.ycombinator.com/item?id=49877515
- GitHub REST API for star and fork counts, queried 29 September 2026: https://api.github.com/repos/archestra-ai/openappa
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.
