
Image: Armature
Why it mattersBuyers of developer tools can now read what agents say worked and broke, and a product team can see how its tool scores with the agents that increasingly pick it.
A team shopping for a database, an auth provider or a CI service now has a second opinion to read next to the vendor's own copy. It comes from the agents that would use the tool.
Armature launched agent.reviews, a public site where coding agents post ratings of the software they used after real tasks. At launch, the catalog carried 3,981 tools and 189,144 reviews across 27 categories, from source control and databases to voice AI, sandboxes and browser automation. The reviewers are agents like Claude Code, Codex, Cursor, Antigravity, Gemini CLI and OpenCode. A Hacker News thread on the launch had 50 points at the time of writing.
How a review is made
Agent.reviews publishes three scores, each from 1 to 5: usefulness, which asks whether the tool did what the task needed; ease, which measures setup and daily use effort; and reliability, which asks whether the tool behaved the way the agent expected. The aggregate for a tool is the average of these scores across its reviews. A tool rated 4.3 or higher is labelled Excellent, 3.8 to 4.2 Great, 2.8 to 3.7 Average, and below that Poor or Bad. Tools with fewer than five reviews are listed separately so a single five-star rating cannot lead a category.
An agent writes a review after a task, naming the tool, how it connected, a short description of the work, the result and any friction. Setup is a one-line prompt pasted into the agent that installs a skill from agent.reviews/skills and signs the computer in once with Google or email. ChatGPT and Claude users can add https://agent.reviews/mcp as a connector instead. Reviews that trip a check for keys, tokens, emails or internal addresses are held before publishing. Code, prompts and conversation content never leave the agent.
Agent-native by construction
Every page on the site has a Markdown version at the same address plus .md, including category indexes, individual tool pages and the install guide. The site includes a sibling skill called tool-reviews that agents can use to check what other agents say about a tool before adding it to a workflow. Early signals, from the launch catalog, show Git at 4.7 and 20,025 reviews, FastAPI at 4.6 and 1,835 reviews, Vercel at 4.1 and 1,297 reviews, and Cursor at 3.7 and 576 reviews with 47 percent of tasks completed.
The company behind it, Armature, also runs a separate experiment called leaderboards that counts how often agents pick a tool when given a free choice on the same task across many runs. Reviews and picks are kept as different signals: reviews report day-to-day use, picks report first choices.
A product team shipping something agents will touch now has a second surface to watch next to its own analytics, and a signal about which setup steps agents report as problems. A team deciding between two libraries at the same job has a public shortlist with agent-reported friction attached.
Source
- agent.reviews, Armature
- How reviews work, agent.reviews
- Launch discussion, Hacker News
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.