AI NewsDev toolsAnnouncement

OpenAPPA is a new MIT-licensed permissions layer for AI agents, and it stopped every attack in its author's own AgentThreatBench run

OpenAPPA, an MIT-licensed permissions layer for AI agents that sits between an agent and its tools and refuses actions when the data was not cleared for that destination, was posted to Hacker News as a Show HN on 28 September and drew 23 points and 11 comments.

AI News

Editorial2 min read

LinkedInX
GitHub repository social card for archestra-ai/openappa, a permissions policy library for AI agents

Image: GitHub

Why it mattersA team that runs an agent with real access to email, storage or paying tools now has an open-source layer to try that gives the same answer on the same request every time, and does not depend on the model to police itself.

Every team building with agents runs into the same problem: a model told to book a hotel with the company card can be steered by a hidden line in a webpage into buying flights instead. OpenAPPA is a new MIT-licensed library that puts a fixed rule between the model's next action and the tool call, and refuses the call when the data being passed is not cleared for that destination.

The project was posted to Hacker News as a Show HN on 28 September 2026 by user motakuk and reached 23 points and 11 comments that day. The GitHub repository at archestra-ai/openappa has 34 stars and 5 forks, per the GitHub API on 29 September, and was created on 18 August. The README calls the project a preview and an RFC, and says the config and wire surfaces may still change.

What the rule is

OpenAPPA implements what it calls the Agentic Permissions Policy Algebra: every piece of data an agent touches carries a label for how sensitive it is and how far it can be trusted, and every tool call is checked against those labels before it runs. The label can only get more restrictive, never less, and the decision is derived from the label rather than asked of the model. The README calls this "deterministic" because the same request gets the same answer every time, no matter what a prompt says.

The benchmark, and what it measures

The repository ships two benchmarks. Bench-Corp is 20 multi-step enterprise workflows run 200 episodes per model, and measures task completion. AgentThreatBench is a 24-task suite the README pairs with OWASP's Top 10 for Agentic Applications, and measures Attack Success Rate across 720 attack attempts on three models: GPT-5.6 Luna, DeepSeek V4 Flash and Gemini 3.7 Flash.

Archestra AI reports that OpenAPPA blocked all 720 attacks while keeping task completion at 89 percent, against Claude Auto at 90 percent completion with 10 percent of attacks succeeding, and FIDES at 41 percent completion with 31 percent succeeding. These are the project's own numbers on its own benchmark and no third party has reproduced them. The framework plugs into an existing agent loop in one place, and the README lists Claude Code, kAgent and Archestra as integrations already working.

Where the pitch stops

A permission layer only helps if the labels are correct in the first place, and OpenAPPA does not remove the work of writing them. The README acknowledges it is a preview and an RFC, so a team adopting it now is signing up for surface changes between versions. There is no independent reproduction of the benchmark, and Show HN points at 23 are real traction but modest. A team running paid tools from an agent still gets a real win: the decision to allow or refuse a call is the same every time, and does not depend on the model noticing the trick.

Source

SourceGitHub

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX