AI NewsModels & agentsAnnouncement

Nemotron 3 cleared gold at IOI 2026 and IMO 2026, NVIDIA says

NVIDIA posted a Nemotron 3 build that it says scored 535.4 of 600 at IOI 2026 and 30 of 42 at IMO 2026, both above the gold threshold, with the fine-tuned checkpoints, the training recipe and a 200-problem benchmark released on Hugging Face.

AI News

Editorial2 min read

LinkedInX

Why it mattersA team that cares about the price of a very capable reasoning model now has an open NVIDIA checkpoint that cleared two olympiad golds, with the SFT and RL datasets and an inference loop published so the result can be reproduced instead of taken on trust.

A reasoning model's score on a human olympiad usually arrives as a bare number on a slide. NVIDIA posted something closer to a lab notebook on Hugging Face on 7 October: Nemotron 3 scores at IOI 2026 and IMO 2026, the recipe behind them, the checkpoints that produced them, and a 200-problem test set called Nemotron-IMO-Bench that anybody can run.

The two scores

At IOI 2026, the programming olympiad, NVIDIA says Nemotron-3-Ultra-CC with SFT and GenCorrect scored 535.4 out of 600, against a gold threshold of 361.12 and a top human score of 498.27. NVIDIA labels the attempt as a live, prospective run under the official competition's own constraints, as an unofficial benchmark, not an entry in the IOI ranking.

At IMO 2026, the mathematical olympiad, Nemotron 3 Ultra scored 30 out of 42 in a generate-verify-refine system that combines its SFT and RL checkpoints. The gold threshold at this year's IMO is 29. The model took full credit on four of the six problems. NVIDIA says the papers were graded by the official IMO graders.

The recipe NVIDIA says it used

The post walks through four steps the team applied to both tasks. Start with a strong Nemotron base. Curate domain-specific problems and reasoning traces. Apply SFT and RL post-training. Pair the specialist with an inference loop that generates, evaluates and refines.

For IOI, the training set is 22,000 curated problems with synthetic reasoning traces. The 30B total, 3B active Nano model used SFT and RL; the 550B total, 55B active Ultra model used SFT only. Inference ran under GenCorrect, an iterative generate-evaluate-refine strategy. The progression numbers NVIDIA reports on an IOI 2025 benchmark show the inference loop was the step that cleared gold: Nano SFT scored 280, Nano SFT plus RL 291, Nano with GenCorrect 468, Ultra with GenCorrect 502.

For IMO, the SFT corpus is 414,890 quality-filtered examples across 15,818 unique proof problems covering proof generation, refinement, verification and meta-verification. RL used 9,597 proof problems near the capability frontier. Inference used the generate-verify-refine system with complementary SFT and RL checkpoints.

What is on Hugging Face today

Nemotron-3-Ultra-CC is published as a model. The Nemotron Labs IMO 2026 Collection carries the SFT and RL checkpoints, the training datasets, and the Nemotron-IMO-Bench test set of 200 olympiad problems. The pipelines and quickstarts are in the NeMo-Skills repository on GitHub. Two papers are posted on arXiv, one for IOI 2026 and one for IMO 2026.

The reading in one line

NVIDIA writes that neither fine-tuning on its own nor brute-force sampling on its own produced a medal, and that only co-designing the model, the data and the inference loop did. A team picking a reasoning model today can read that line, open the paper, run the published benchmark and audit the claim against work it cares about, which is a different conversation from a leaderboard screenshot in a vendor deck.

Source

One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO, NVIDIA on the Hugging Face blog.

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX