benchmarkNVIDIA

NVIDIA's fine-tuned Nemotron 3 Ultra scores 535.4/600 at IOI 2026 and 30/42 at IMO 2026, both above gold

TL;DR

NVIDIA reports that specialized versions of Nemotron 3 Ultra reached gold-medal level at both IOI 2026 (535.4/600) and IMO 2026 (30/42). The IOI run was unofficial, and the IMO proofs were graded by official IMO graders, according to NVIDIA. Checkpoints, datasets, and inference pipelines are published on Hugging Face.

3 min read
0

NVIDIA says fine-tuned versions of its Nemotron 3 family reached gold-medal level at both the International Olympiad in Informatics (IOI) 2026 and the International Mathematical Olympiad (IMO) 2026. The IOI system scored 535.4 out of 600, and the IMO system scored 30 out of 42. NVIDIA published the details on Hugging Face on October 7, 2026.

The results

Competition System Score Reference
IOI 2026 Nemotron-3-Ultra-CC with SFT and GenCorrect 535.4/600 Gold threshold 361.12; top human score 498.27
IMO 2026 Nemotron 3 Ultra general, SFT, and RL checkpoints in a generate-verify-refine system 30/42 Official gold threshold 29

The IOI result came from a live, prospective run under the same time, internet-access, and submission constraints as human contestants. NVIDIA describes it as an unofficial, unsupervised benchmark that is not part of the official IOI ranking. The IMO proofs were graded by official IMO graders, according to NVIDIA. The system earned full credit on four of the six problems.

How the IOI system was built

NVIDIA curated 22,000 competitive programming problems and generated synthetic reasoning traces to train two specialists:

  • Nemotron-3-Nano-CC: 30B total parameters, 3B active. It received both SFT and RL.
  • Nemotron-3-Ultra-CC: 550B total parameters, 55B active. It received SFT only.

On IOI 2025, Nano rose from 130 points before post-training to 280 after SFT and 291 after RL. With GenCorrect, NVIDIA's iterative generate-evaluate-refine strategy, Nano reached 468 points, above that year's 438.3 gold threshold. Ultra-CC reached 502 points with the same test-time strategy.

NVIDIA reports that a single SFT epoch on Ultra was enough to beat the fully post-trained Nano across IOI, ICPC, and LiveCodeBench Pro. That finding shaped the Ultra-CC system used for IOI 2026.

How the IMO system was built

The SFT corpus held 414,890 quality-filtered examples across 15,818 unique proof problems. It covered proof generation, refinement, verification, and meta-verification. The RL model trained on 9,597 proof problems selected near the model's capability frontier.

In development experiments, both post-trained checkpoints outperformed the general-availability model. The SFT checkpoint was strongest in the first search round, and the RL checkpoint had the best overall single-checkpoint result. The final system used both alongside the general model. It generated candidate proofs, scored them, produced critiques, and refined the best attempts, and a separate high-compute stage picked the final submission. NVIDIA says the system worked entirely in natural language, with no formal prover, external tools, or internet access.

According to NVIDIA, using complementary SFT and RL checkpoints beat drawing more samples from a single checkpoint.

What is released

  • The Nemotron Labs IMO 2026 collection on Hugging Face holds the SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench, a new 200-problem olympiad-level benchmark.
  • Nemotron-3-Ultra-CC is available on Hugging Face.
  • The NeMo-Skills repository includes the IMO inference pipeline, prompts, submitted proofs, the IOI evaluation and inference pipeline, and a reproducible quickstart.
  • Papers describe the IMO generate-verify-refine system and the IOI training recipe and GenCorrect method.

Context window, pricing, and training cutoff for these specialist models were not disclosed in the post.

What this means

The main finding is about method more than any single score. NVIDIA attributes the results to three things working together: domain-specific data, SFT and RL, and a feedback-driven inference loop. Neither fine-tuning alone nor brute-force sampling alone produced the medals. For teams building reasoning systems, that points to investing in verifiers and refinement loops alongside post-training.

Two caveats apply. The IOI score is unofficial and self-reported, although the live run mirrored human constraints. Both systems also relied on heavy test-time compute that NVIDIA does not quantify in the post, so the scores do not reflect single-pass model performance. Publishing checkpoints, datasets, and pipelines still lets outside researchers check the claims and test whether the recipe transfers to other domains.

Related Articles

model release

Nvidia Releases Nemotron 3 Diarization, a Free 100M-Parameter Model That Tracks 8 Speakers in Real Time

Nvidia released Nemotron 3 Diarization, a free 100-million-parameter model that identifies who is speaking in real time across up to eight participants. It leads the VoiceArena Diarization Benchmark v1 with a 14.7% error rate, cutting errors by 41% versus its predecessor.

model release

NVIDIA Releases Nemotron 3 Diarization, a 100M-Parameter Open-Weight Model Ranked #1 on Voice Arena's Diarization-Bench

NVIDIA released Nemotron 3 Diarization, a 100-million-parameter open-weight model that identifies who is speaking and when in audio conversations. It ranked #1 among 17 system configurations on Voice Arena's Diarization-Bench with a 14.72% diarization error rate, supporting up to eight speakers in both live and recorded audio.

product update

Nvidia and Palantir Deploy AI to Manage Supply Chains, Starting With Nvidia's Own 1.3-Million-Part Racks

Nvidia and Palantir are integrating Nvidia's open Nemotron models and cuOpt optimization engine into Palantir's Foundry platform to manage complex supply chains. The first deployment is Nvidia's own operation, where a single Vera Rubin server rack contains 1.3 million components.

research

Nvidia's SoL-Pi Cuts Coding Agent Token Usage by Up to 49% Through Automated Harness Optimization

A new Nvidia research system called SoL-Pi automatically rewrites the control logic of coding agents rather than the underlying model, cutting token usage by up to 49% while keeping performance nearly intact. The approach could shift efficiency gains in AI agents from model-level tricks to harness-level engineering.

Comments

Loading...