NVIDIA says fine-tuned Nemotron 3 hit gold level at IMO 2026 and IOI 2026

NVIDIA says fine-tuned Nemotron 3 hit gold level at IMO 2026 and IOI 2026

NVIDIA's Nemotron team published a post on the Hugging Face blog saying that, starting from Nemotron 3, it used supervised fine-tuning (SFT), reinforcement learning (RL) and feedback-driven inference to build systems that reached gold-medal level at both IMO 2026 and IOI 2026. The authors frame it as a reusable recipe: start with a strong Nemotron base, curate domain problems and high-quality reasoning traces, apply standard post-training (SFT and, where useful, RL), and pair the specialist with an inference loop that generates, evaluates and improves candidate answers.

Two caveats on the results come from the post itself. The IOI result came from a live, prospective run under the same time, internet-access and submission constraints as human contestants, but it was an unofficial, unsupervised benchmark and was not included in the official IOI ranking. The IMO system's submitted proofs, by contrast, were graded by official IMO graders.

On competitive programming, the team curated 22,000 problems and generated synthetic reasoning traces to train two specialists. Nemotron-3-Nano-CC (30 billion total parameters, 3 billion active) received both SFT and RL. Nemotron-3-Ultra-CC (550 billion total, 55 billion active) received SFT. On IOI 2025, Nano went from 130 points before post-training to 280 after SFT and 291 after RL. With GenCorrect, the team's iterative generate-evaluate-refine strategy, Nano reached 468 points and crossed the gold threshold of 438.3. Ultra-CC reached 502 points with the same test-time strategy. SFT produced most of Nano's gain, with RL adding a smaller but consistent improvement. For Ultra, one SFT epoch was enough to outperform the fully post-trained Nano across IOI, ICPC and LiveCodeBench Pro. That finding guided the competition-specific Ultra-CC system used at IOI 2026, which scored 535.4 out of 600.

The IMO project applied the same idea to olympiad proofs. Starting from Nemotron 3 Ultra, the team trained one specialist with SFT and another with RL. The SFT corpus held 414,890 quality-filtered examples across 15,818 unique proof problems, and covered proof generation, refinement, verification and meta-verification, so the model learned to build arguments, find gaps, respond to critiques and judge whether a proof was complete. The RL model trained on 9,597 proof problems selected near the model's capability frontier. In development experiments both post-trained checkpoints outperformed the general-availability model; the SFT checkpoint was strongest in the first search round, while the RL checkpoint had the best overall single-checkpoint result. Because their strengths were complementary, the final system used both specialists alongside the general model.

For each IMO problem, the models generated candidate proofs, scored them, produced critiques and refined the most promising attempts, and a separate high-compute stage selected the final submission. The whole system worked in natural language, with no formal prover, external tools or internet access. It scored 30 out of 42, including full credit on four of the six problems, and exceeded the official gold-medal threshold.

The authors argue that fine-tuning and test-time compute work together: better specialization gives the inference loop better candidates, critics and refinements. At IOI, GenCorrect turned fine-tuning gains into larger improvements over multiple feedback rounds. At IMO, using complementary SFT and RL checkpoints was more valuable than simply drawing more samples from one checkpoint. In their words, the medals came from co-designing the model, the data and the inference loop, not from fine-tuning alone or brute-force sampling alone.

On release, the Nemotron Labs IMO 2026 collection on Hugging Face brings together the SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench, a new benchmark of 200 olympiad-level problems. An IMO paper describes the training approach and the generate-verify-refine system, and the NeMo-Skills repository includes the IMO inference pipeline, prompts, submitted proofs and a reproducible quickstart. For competitive programming, Nemotron-3-Ultra-CC is available on Hugging Face, an IOI paper gives the training recipe and the GenCorrect method, and the IOI evaluation and inference pipeline are also in NeMo-Skills.

Key facts

  • NVIDIA's Nemotron team says systems built from Nemotron 3 with SFT, RL and feedback-driven inference reached gold-medal level at IMO 2026 and IOI 2026.
  • The IMO system scored 30 out of 42, with full credit on four of six problems, and its proofs were graded by official IMO graders. It used natural language only, with no formal prover, external tools or internet access.
  • The IOI 2026 run was live but unofficial and unsupervised, outside the official ranking; the Ultra-CC system scored 535.4 out of 600.
  • On IOI 2025, Nemotron-3-Nano-CC rose from 130 points before post-training to 280 after SFT and 291 after RL, then to 468 with GenCorrect (gold threshold 438.3); Ultra-CC reached 502.
  • Released on Hugging Face: SFT and RL IMO checkpoints, both training datasets, the 200-problem Nemotron-IMO-Bench, and Nemotron-3-Ultra-CC, with pipelines in NeMo-Skills.

Why it matters

NVIDIA presents the two results as evidence that a general open model family can be turned into a frontier-level domain specialist with standard post-training plus a well-designed inference loop, rather than a new foundation model for each challenge. The post adds that fine-tuning and test-time compute reinforce each other: at IMO, combining complementary SFT and RL checkpoints was more valuable than drawing more samples from one checkpoint. It also finds that adaptation differs by scale: for Nano, SFT gave most of the gain and RL a smaller one, while for the much larger Ultra a single SFT epoch already beat the fully post-trained Nano on IOI, ICPC and LiveCodeBench Pro.

Who it affects

Researchers and engineers working on reasoning models, competitive programming and mathematical proof generation are the main audience, since the post releases checkpoints, datasets, a benchmark and inference pipelines. Teams weighing whether to specialise an existing base model for a demanding domain get a concrete recipe with named data sizes and stages. The post is also addressed to the Hugging Face community, which it hopes will build on the release.

How to use it

The Nemotron Labs IMO 2026 collection on Hugging Face has the SFT and RL checkpoints, both training datasets and Nemotron-IMO-Bench. The NeMo-Skills repository includes the IMO inference pipeline, prompts, submitted proofs and a reproducible quickstart, plus the IOI evaluation and inference pipeline. Nemotron-3-Ultra-CC is available on Hugging Face for competitive programming. The IMO paper covers the generate-verify-refine system and training approach; the IOI paper covers the training recipe and the GenCorrect method. The post does not give licence terms.

How solid is it

This is a first-party post by the team that built the systems, and every result is the authors' own account. The IMO score of 30 out of 42 was graded by official IMO graders, which is the stronger of the two claims. The IOI 2026 result came from a live run under human-contestant constraints, but the authors state it was unofficial, unsupervised and not in the official IOI ranking. The IOI 2025 numbers (468 and 502 against a 438.3 gold threshold) are from a past contest, not a live run. Papers, code and data are released, so the claims can be checked by others.

Risks and caveats

The gold-level claim for IOI rests on an unofficial benchmark, and the IOI 2026 gold threshold is not given in the post; 438.3 is the IOI 2025 threshold. The IMO cutoff is not stated as a number either, only that 30 out of 42 exceeded the official gold threshold. No compute budget, GPU counts or training time are given, and the post itself says the training and inference runs were substantial. No numbers are given for how the IMO checkpoints compared with the general-availability model. The post does not say how many human contestants or other AI systems competed, or compare with other labs' results. Nemotron-3-Nano-CC is described as getting SFT and RL, and Ultra-CC as getting SFT only.

“They came from co-designing the model, the data, and the inference loop.”

— NVIDIA Nemotron team, Hugging Face blog post