Quicksilver hands the skim-and-sort work in a Claude Code session to a smaller model, and cuts token spend by 86 percent on the author's benchmark
Quicksilver is a Claude Code skill by Udit Akhouri that routes bulk log triage, file filtering, spam and sentiment work to TypeSafe's Jev decision model, cuts token use by a median 82 percent on a twelve-task benchmark, and reached 83 GitHub stars in four days under MIT.
Image: GitHub
Why it mattersA team paying Claude by the token for repetitive skim work now has a way to keep Claude on the thinking parts and move the sorting off, without giving up final accuracy on most classification tasks.
A large share of every Claude Code session is spent reading things only to decide whether they matter. Quicksilver is a new open-source skill that lets Claude hand those sorting jobs to a smaller, cheaper model, and keep itself on the reasoning that needs Claude.
The skill was published by Udit Akhouri on 25 September 2026 and reached 83 GitHub stars and 4 forks in four days, per the GitHub API on 29 September. It is MIT-licensed and installs with a single npx command from the repository. The author says out loud that it is independent and not affiliated with either Anthropic or TypeSafe AI.
What it moves off Claude
The skill routes closed-label classification, file filtering, log triage, and relevance ranking to Jev, TypeSafe's typed-decision model. The README lists the jobs it takes as file filtering ("does this file handle auth?"), log triage and pattern matching, intent labelling against a fixed label set, and line-location searches in large files. Jev returns one line per hit, folds repeated log patterns into line-number ranges, and marks anything borderline with a ? so Claude can look at those cases itself.
The author is explicit about what to keep on Claude: writing, editing, multi-step reasoning, and anything an exact grep would answer.
The twelve-task benchmark, and what it measures
The repository ships a benchmark of twelve tasks that runs Claude Code alone against Claude with Quicksilver, scored against hidden ground truth. Tokens fell by a median 82 percent and by 86 percent across the twelve tasks, and time fell by up to a factor of 20 on the fastest task (Banking77 support ticket routing, 60 seconds to 6). The whole benchmark cost 45 cents of Jev, at the price TypeSafe lists of $0.042 per million input tokens with output free.
Accuracy holds on 8 of 12 tasks, defined as landing within 2 F1 or accuracy points of Claude alone. On the other four the drop is visible: real log triage on BGL supercomputer logs fell from 54 percent F1 to 23, a security review shortlist fell from 100 to 89, spam filtering fell from 97 to 92, and commit classification fell from 83 to 76. Both the wins and the four losses are the author's own measurement, and are the reason the README says to treat the smaller model as a shortlist rather than a final answer on subjective tasks.
Where the data goes
Content the skill delegates is sent to TypeSafe's API at api.typesafe.ai. TypeSafe states that Jev is not trained on customer data. The skill excludes .env files, private keys, certificates and credentials files by default, and respects .gitignore.
For a working developer who runs Claude Code all day, this is the smallest kind of change to try: a skill install, one model call per bulk decision, and a bill that is measured in cents rather than tens of dollars for the sort of file-filtering pass that would otherwise burn 84,000 Claude tokens.
Source
- GitHub repository, primary source: https://github.com/UditAkhourii/quicksilver
- GitHub REST API for star and fork counts, queried 29 September 2026: https://api.github.com/repos/UditAkhourii/quicksilver
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.