AI NewsModels & agentsAnnouncement
A decision model built from a 1.7B open LLM, 59% on CommonsenseQA
Nish Tahir shows how to build a Jev-style one-pass decision model from Qwen3-1.7B with constrained decoding, and reports 59.4 percent accuracy on a CommonsenseQA holdout, rising to 62.4 percent after a short fine-tune.

Why it mattersA team that pays per call for a hosted decision model can now test whether a small local LLM covers the task, with a short script and a public benchmark to measure the gap.
Paying per call for a hosted decision model only makes sense when a local one cannot do the same job. Nish Tahir published a short walkthrough that builds a Jev-style one-pass classifier from a 1.7B open-weight language model, with the code and the measurement side by side.
Decision models like TypeSafe AI's Jev pick between a fixed set of answers in a single forward pass. The post shows how to copy that behaviour from Qwen3-1.7B in about forty lines of Python. The script prefills the question and the option labels into the model, reads the output logits for the final position once, masks every token that is not one of the allowed option labels, applies softmax to what remains, and returns the top option.
The numbers on a public set
Tahir ran the script on a 1,221-question holdout of CommonsenseQA, with five options per question labelled A to E. The baseline Qwen3-1.7B scored 725 correct for 0.5938 accuracy and 0.5844 macro F1. A short fine-tune on the training split lifted that to 762 correct, 0.6241 accuracy and 0.6234 macro F1. The improvement is uneven across options: precision on D fell from 0.72 to 0.67 while recall on D climbed from 0.39 to 0.54, so the fine-tune reduced the model's bias away from the later options.
Treating the softmax as a confidence score
The second half of the piece covers a trap. The raw probability over the five allowed tokens is uncalibrated. In the top bin, confidence 0.9 to 1.0, 809 predictions averaged 0.9855 confidence and were correct 70.1 percent of the time. In the 0.8 to 0.9 bin, 121 predictions averaged 0.8555 confidence and were correct 47.1 percent of the time. A caller who routed on "confidence above 0.9" was wrong three times in ten.
Tahir fits a single temperature value against the eval, finds 3.797, and the bins realign: the 0.9 to 1.0 bin now holds 109 predictions, averages 0.9333 confidence, and is correct 95.4 percent of the time. The 0.3 to 0.4 bin holds 217 predictions, averages 0.3507 confidence, and is correct 39.2 percent of the time. The fix is one scalar and one eval run.
The whole thing is one script per step, under nishtahir/build-your-own-jev on GitHub (5 stars at the time of the post). The repo covers building a labelled dataset, running the baseline eval, a short fine-tune and the temperature calibration.
A team that already pays per call to a hosted decision model has a cheap diagnostic here: run the exact routing task against a 1.7B open model with the same script, read the per-bin calibration table, and see where the gap sits. When the top-1 accuracy is close enough, calibration is one scalar away. When the top-1 is far off, the hosted model is earning its price on something the public benchmark does not test.
Source
Build your own decision model, Nish Tahir, with the companion build-your-own-jev repository.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.

