AI NewsModels & agentsReported

OpenAI opens a limited preview of a Decisions API built on its Luna model, returning classification answers in 150 milliseconds

The New Stack reports that OpenAI has opened a limited preview of a Decisions API built on its Luna model that returns predefined answers with confidence scores in 150 milliseconds, with an OpenAI spokesperson confirming a broad rollout in the coming days.

AI News

Editorial2 min read

LinkedInX

Why it mattersA team routing requests or classifying content with a chat model and a hand-written prompt now has a named OpenAI product to compare against, and the 150 millisecond figure is what decides whether it replaces a step that a chat model does more slowly.

The New Stack reported on 29 September that OpenAI has quietly opened a limited preview of a Decisions API for routing requests and classifying content, built on Luna, the smallest model in the company's current lineup.

Frederic Lardinois, The New Stack's senior AI editor, wrote that an OpenAI spokesperson confirmed the preview to the publication and said the company plans to share more "at broad rollout", which is planned for the coming days. The Decisions API returns a set of predefined answers with confidence scores rather than the free-form text a chat model gives back.

What the numbers say

OpenAI told The New Stack that the Decisions API returns results in 150 milliseconds, and compared that against GPT-6 Luna running the same job, which took 1.6 seconds. That is roughly ten times faster on the same task. The publication says pricing per call, the maximum number of candidate answers per request, and whether developers can tune the model on their own data are all still unknown at the time of writing.

Two other facts came from OpenAI directly to the publication. Luna is the smallest and cheapest model in OpenAI's current lineup. The Decisions API is a separate product from the Moderation API, which OpenAI has run for years, but with one important difference: the categories on the older API are set by OpenAI, while the Decisions API takes categories from the developer in the prompt.

Why a separate model exists

Most teams handle this classification job today with a regular chat model and a carefully worded prompt, asking it to pick from a list. Some read the model's token probabilities and treat the score for the top answer as a confidence. The other option is to train a small classifier: fast and cheap at runtime, but with labelled data to gather every time the label set changes.

A decision model sits between those two options. It takes new labels in the prompt, like the chat model does, and returns a confidence score that reflects the model's actual belief, closer to what a trained classifier gives back.

The Jev question the publication raises

Lardinois writes that the Decisions API "is likely a reaction to TypeSafe and Jev, and OpenAI probably rushed the announcement ahead of its DevDay". That framing belongs to the reporter, so treat it as one reader's read of the timing. The timing does line up: a run of Jev-like open source projects shipped through September, and OpenAI's own DevDay on 29 September led with Dots and GPT-6.1 Sol while leaving the Decisions API off the stage.

The reader's decision depends on the two prices The New Stack could not report: what OpenAI charges per call, and how many candidate answers each call can hold. Ten times faster than a chat model at the same task is only worth adopting if the per-call cost sits below what a small trained classifier would run, and if a single call can hold enough labels to be useful. The publication says those details will decide whether this becomes a standard building block in agent frameworks or stays a niche tool.

Source

Primary: OpenAI answers TypeSafe's Jev with a Decision API built on Luna, The New Stack, 29 September 2026, by Frederic Lardinois.

Reported byThe New Stack

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX