Models & agents

Nvidia says a 30B Nemotron fine-tuned on its supply-chain data scored 86.7 percent, versus 55.5 for the 550B Nemotron

September 10, 2026 at 4:20 AM PT

Bar chart comparing accuracy of a post-trained 30B Nemotron Lightning against the 550B Nemotron Ultra on Nvidia's supply-allocation task

Image: Nvidia via The New Stack

Why it mattersA team choosing between a large general model and a smaller model tuned on its own decisions now has a named 30-point gap on a narrow task, so specialization moves ahead of size as the default option to test first.

Nvidia and Palantir said on Thursday that they have fine-tuned Nvidia's 30-billion-parameter Nemotron 3.5 Lightning model on decisions made by Nvidia's own supply-chain team, and that the tuned model beat a version of Nemotron 3 Ultra with 550 billion parameters on the same internal task. The New Stack, which reported the result on 2026-09-09, quotes Nvidia's figures: 86.7 percent accuracy for the 30B model against 55.5 percent for the 550B model, on Nvidia's supply-allocation task.

How the system is put together

The setup pairs Palantir's Foundry and AIP platforms, plus its Ontology, with Nvidia's cuOpt optimization software and the tuned Nemotron. According to Nvidia's account in the report, the Ontology holds a live map of components, factories, capacity and production commitments. cuOpt works out how scarce parts could be distributed. Nemotron weighs the wider context and recommends what planners should do next.

Nvidia is starting with its own supply chain because it is unusually complex. The company says a single Vera Rubin rack contains about 1.3 million parts, and that one missing component can stall an assembly that is otherwise ready to go. Palantir CEO Alex Karp said in a separate statement that Nvidia has "arguably the most valuable, intricate, and complex supply chain in the world."

The number, and the caveat Nvidia states with it

The 86.7 versus 55.5 figure is Nvidia's own, published in a technical blog by Nvidia solutions architects Nell Barber, Rana Haber and Aastha Jhunjhunwala alongside the main announcement. It compares two Nvidia models on a task defined by Nvidia's own supply-chain operations team, so the number is an internal result, not an independent benchmark.

The Nvidia authors said in the same post that the gain does not make the smaller model more capable overall. "Its gains are concentrated in the domain it was post-trained on," they wrote. They add that future production risk forecasting remained hard even after fine-tuning, so specialization improved the decision task without solving every prediction problem attached to it.

The 30B model is Nemotron 3.5 Lightning, released in August. Nvidia publishes weights for the Nemotron 3 series and, for many of the models, training data and recipes. The 550B model is Nemotron 3 Ultra, the top of the current Nemotron 3 line that Nvidia launched in December.

Where the reader lands after this

The report is one company's read on one task, so the specific number should not be lifted out of its context. What is portable is the shape of the trade. Nvidia and Palantir are offering the same architecture (Foundry and AIP for the data, Ontology for the map, cuOpt for the optimizer, a fine-tuned open-weight model for the recommendation) to customers in manufacturing, energy, healthcare, automotive and aerospace. Customers can train Nemotron on their proprietary data and run the resulting system on-premises or through cloud and colocation providers, keeping data and weights inside their own environment.

For a team building software for that kind of operations problem, the default has been to reach for the biggest general-purpose model and add retrieval or tool use. The alternative in this playbook is to pick a smaller open-weight model, invest in labelled decisions from the people who currently do the job, and post-train. The Nvidia result gives that alternative a concrete 30-point margin to point at when the internal debate opens.

Source

Reported by: The New Stack

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

More from AI News

Hugo Vergnes trained a 3.8B language model to 0.384 on CORE for $998 in 43 hours

Solo engineer Hugo Vergnes trained a 3.8B-parameter language model on 65B tokens in 43 hours for $998, scoring 0.384 on the CORE benchmark and beating OpenAI's 2019 GPT-2 1.5B by a wide margin.

Source: Hacker NewsModels & agents

Meng Zhang's reasoning-prefill rerun shifts Qwen 3.8's answers 18 points toward GPT-5.5 Pro

Prefilling Qwen 3.8 A95B with the first one percent of GPT-5.5 Pro's reasoning tokens moved its answers 18 percentage points closer to GPT-5.5 Pro's, on a private set of 45 problems.

Source: Hacker NewsModels & agents

Nvidia agrees to buy Hugging Face for $12.9 billion, and says the platform stays open and multi-cloud

Nvidia has agreed to acquire Hugging Face for $12.93 billion, and says the platform will keep its brand, stay multi-cloud, and continue supporting open weights across the ecosystem.

Source: PressInfrastructure