AI NewsModels & agentsAnnouncement
Ollama can now run small decision models on your own machine, and answer many questions in one call
Ollama 0.35 can now run small decision models on your own machine and answer many yes-or-no questions in one request. It ships with three of them, starting with Bespoke Labs Nimble 9B, and reports about 91 milliseconds per question on a MacBook Pro M5 Max.

Why it mattersA team using a chat model to sort tickets, pick which model handles the next request, or check for unsafe content can now do it on its own hardware in about 91 milliseconds per decision, and stop paying a hosted service for every choice.
Asking a chat model to answer yes or no still makes it write the answer as text, one word at a time, and pay per call at the hosted rate. That is fine when the model has to explain itself, and wasteful when the whole answer is one word repeated a thousand times a day. Ollama shipped support on 29 September for a different kind of model, one built to return the answer as a small piece of data instead.
The new address, /v1/systemone, was added in Ollama 0.35. Send it a block of text and a set of named questions, and it answers all of them in one call. The post shows three kinds of question: pick from a list, yes or no, and a number score, each returned with the model's own probabilities and a separate confidence figure. Ollama says it fits sorting a support ticket into the right team, picking which chat model handles the next request, and checking whether a message breaks a content rule. The address follows the same shape TypeSafe AI uses for its Jev decision model, so any code already written for Jev works without changes.
Three models today, one from Bespoke Labs
The launch ships with three local models, all open. Nimble is a 9B parameter decision model that Ollama describes as "developed by Bespoke Labs". The other two are Tev1 4B and Tev1 0.8B, listed as experimental releases from Together AI. Ollama says "more decision models are coming soon, including models served by Ollama's cloud", but does not name them.
Anyone already using the Ollama CLI or its OpenAI-compatible client picks up Nimble by pulling the tag, without running a second local server.
What the post measures
Ollama reports one latency number, from a worked example that classifies a support ticket into a team, a refund choice and an urgency score in one call: "Nimble 9B averaged 91 ms per decision" running locally on an Apple M5 Max. That is per named question, not per request. The blog post's accuracy chart plots Nimble and Tev1 against a Jev 1.13 baseline that Ollama says is "from Bespoke Labs' published run on the same decisions" and covers 3,880 decisions across 13 public labelled datasets. Ollama does not print an overall accuracy number in the text, and the test set is Bespoke Labs' own, so treat the chart as the vendor's evidence rather than an independent audit.
The example response in the post carries every question's answer, its probabilities and a separate confidence score, plus a count of input and output tokens for that call. That last part matters for anyone watching the bill: a decision model still runs up costs when a hosted service serves it, and on your own hardware there is no per-call charge at all.
A team asking a chat model to decide which other model should answer a request, or to sort inbound support tickets, is doing a job a decision model was built for. Two things change once that decision runs on your own machine. The cost of each check stops rising with the number of requests, because you no longer pay for every call. And the answer arrives with a probability the app can read, so a low-confidence result can be sent to a slower, larger model or to a person, instead of being accepted as if the model were sure. A paid chat model writing its answer as prose gives you neither of those.
Source
Ollama, Ollama now supports Jev-style decision models, 29 September 2026.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.
