Nvidia says a 30B Nemotron fine-tuned on its supply-chain data scored 86.7 percent, versus 55.5 for the 550B Nemotron

Image: Nvidia via The New Stack
Why it mattersA team choosing between a large general model and a smaller model tuned on its own decisions now has a named 30-point gap on a narrow task, so specialization moves ahead of size as the default option to test first.
Nvidia and Palantir said on Thursday that they have fine-tuned Nvidia's 30-billion-parameter Nemotron 3.5 Lightning model on decisions made by Nvidia's own supply-chain team, and that the tuned model beat a version of Nemotron 3 Ultra with 550 billion parameters on the same internal task. The New Stack, which reported the result on 2026-09-09, quotes Nvidia's figures: 86.7 percent accuracy for the 30B model against 55.5 percent for the 550B model, on Nvidia's supply-allocation task.
How the system is put together
The setup pairs Palantir's Foundry and AIP platforms, plus its Ontology, with Nvidia's cuOpt optimization software and the tuned Nemotron. According to Nvidia's account in the report, the Ontology holds a live map of components, factories, capacity and production commitments. cuOpt works out how scarce parts could be distributed. Nemotron weighs the wider context and recommends what planners should do next.
Nvidia is starting with its own supply chain because it is unusually complex. The company says a single Vera Rubin rack contains about 1.3 million parts, and that one missing component can stall an assembly that is otherwise ready to go. Palantir CEO Alex Karp said in a separate statement that Nvidia has "arguably the most valuable, intricate, and complex supply chain in the world."
The number, and the caveat Nvidia states with it
The 86.7 versus 55.5 figure is Nvidia's own, published in a technical blog by Nvidia solutions architects Nell Barber, Rana Haber and Aastha Jhunjhunwala alongside the main announcement. It compares two Nvidia models on a task defined by Nvidia's own supply-chain operations team, so the number is an internal result, not an independent benchmark.
The Nvidia authors said in the same post that the gain does not make the smaller model more capable overall. "Its gains are concentrated in the domain it was post-trained on," they wrote. They add that future production risk forecasting remained hard even after fine-tuning, so specialization improved the decision task without solving every prediction problem attached to it.
The 30B model is Nemotron 3.5 Lightning, released in August. Nvidia publishes weights for the Nemotron 3 series and, for many of the models, training data and recipes. The 550B model is Nemotron 3 Ultra, the top of the current Nemotron 3 line that Nvidia launched in December.
Where the reader lands after this
The report is one company's read on one task, so the specific number should not be lifted out of its context. What is portable is the shape of the trade. Nvidia and Palantir are offering the same architecture (Foundry and AIP for the data, Ontology for the map, cuOpt for the optimizer, a fine-tuned open-weight model for the recommendation) to customers in manufacturing, energy, healthcare, automotive and aerospace. Customers can train Nemotron on their proprietary data and run the resulting system on-premises or through cloud and colocation providers, keeping data and weights inside their own environment.
For a team building software for that kind of operations problem, the default has been to reach for the biggest general-purpose model and add retrieval or tool use. The alternative in this playbook is to pick a smaller open-weight model, invest in labelled decisions from the people who currently do the job, and post-train. The Nvidia result gives that alternative a concrete 30-point margin to point at when the internal debate opens.
Source
Reported by: The New Stack
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually shipping with them. Short, and only when there is something worth reading.

