AI NewsModels & agentsAnnouncement

Jevstiller runs a small local model in front of the paid Jev classification API and gives a mathematical guarantee that it agrees with Jev at least 98 percent of the time

Jevstiller is a new open-source tool by Tomer Glick that trains a tiny local model on TypeSafe's Jev classification API and answers most classification requests on your own hardware in about 15 milliseconds, with a Clopper-Pearson confidence bound that promises the local answers match Jev on at least 98 percent of requests.

AI News

Editorial3 min read

LinkedInX
GitHub social card for the tomerglick57/Jevstiller repository

Image: GitHub

Why it mattersA team paying for a hosted classifier can now offload most calls to a cheap local model with a mathematical statement of how often the two disagree, instead of a threshold set by hand that might quietly break the budget on the next data drift.

A new open-source project called Jevstiller sits between an application and TypeSafe AI's Jev classification API, trains a small local model on the labels Jev sends back, and starts answering easy requests locally in about 15 milliseconds instead of a 300-millisecond network call. It went up on Hacker News on 29 September 2026 as a Show HN, has 37 stars on GitHub as of the morning of 30 September, and is released under Apache 2.0 by an author who publishes under the GitHub handle tomerglick57.

Cascade proxies for hosted classifiers are already a common pattern: a small local model answers the easy queries and passes the harder ones on. The catch is the threshold that decides "easy". A team usually picks it from a calibration run and hopes traffic stays the same shape. Jevstiller replaces the hope with a Clopper-Pearson binomial confidence bound. The operator picks an agreement target, for example 98 percent, and the tool returns the label Jev would have returned on at least that share of requests, with 95 percent confidence.

The numbers on the five benchmark tasks

The author reports coverage and budget breaks on five public classification tasks, comparing the bound-based rule against the older point-estimate rule at a 98 percent agreement target. On Banking77 the bound-based rule covers 74.9 percent of requests locally and breaks the agreement budget 0 out of 20 test splits; the point-estimate rule covers a higher 79.8 percent but breaks the budget on 9. CLINC150 is 78.9 versus 84.3 percent, and 0 versus 12 breaks. AG News is 86.5 versus 90.3 percent, 0 versus 11. TweetEval sentiment is 24.0 versus 28.2 percent, 0 versus 8. TweetEval offensive is 28.7 versus 36.8 percent, 0 versus 6.

The trade the numbers describe is small extra caution on how much traffic goes local, in exchange for a promise the operator can actually hold to a finance team. On the two TweetEval tasks the safe local share is around a quarter to a third of traffic; on Banking77 and CLINC150 and AG News it is most of it.

What the guarantee does not cover

The author is direct about the limits. The bound holds only when the calibration rows are a random sample of the real traffic. During a distribution shift, the local share drops on purpose while the model retrains, so the guarantee holds by letting more requests through to Jev. The drift test in the post shows this: when the author changed Jev's answers on purpose, the local model's share fell from 90 percent to 9 percent in about 4 minutes, and recovered to 90 percent inside 49 minutes as the local model retrained. Budget accuracy costs latency during a drift event.

Two more limits are on the same page. The contract counts requests, so a rare class can carry more of the disagreement than its share of traffic. And the bound tracks how often Jevstiller matches Jev, which is a different question from how often either one matches the ground truth. If Jev is wrong, the local copy is wrong the same way.

For a team already paying for Jev on high-volume classification, that is a specific enough tool to try in a sandbox this week. The saving is a real number to measure. The claim to check is the disagreement rate the tool reports against the one seen on live traffic.

Source

Repository: tomerglick57/Jevstiller on GitHub. Author's write-up of the guarantee, benchmarks and drift test: The Guarantee. All coverage and disagreement numbers come from the author's own runs on the five public benchmark tasks named above.

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX