NVIDIA reported that specialized Nemotron 3 systems achieved gold-level results at IOI 2026 and IMO 2026, providing developers with checkpoints, datasets and implementation materials for competitive coding and mathematical proof workflows. The results, published by Hugging Face, show that the systems relied on both fine-tuning and multi-stage inference rather than a base model alone.

A reusable specialization recipe

NVIDIA described a four-part approach: begin with a capable Nemotron base model, curate domain problems and reasoning traces, apply supervised fine-tuning (SFT) and, where useful, reinforcement learning (RL), then run an inference loop that generates, evaluates and improves candidate answers. The company said the training and inference runs were substantial, but framed the approach as a reusable way to create specialists without building a new foundation model for each domain.

Competitive coding results

For IOI 2026, NVIDIA said Nemotron-3-Ultra-CC, using SFT and its GenCorrect generate-evaluate-refine process, scored 535.4 out of 600. That exceeded the stated gold threshold of 361.12 and the cited top human score of 498.27. The company said the run was prospective and subject to the same time, internet-access and submission constraints as contestants, but it was an unofficial, unsupervised benchmark and was not part of the official IOI ranking.

The coding effort used 22,000 curated problems and synthetic reasoning traces. NVIDIA said its 30-billion-parameter Nano model rose from 130 points on IOI 2025 before post-training to 280 after SFT and 291 after RL; with GenCorrect it reached 468, above that year’s 438.3 gold threshold. The 550-billion-parameter Ultra model received SFT and reached 502 in the same prior-year evaluation. NVIDIA said SFT supplied most of Nano’s gain, while a single SFT epoch on Ultra outperformed the fully post-trained Nano across IOI, ICPC and LiveCodeBench Pro.

Proof generation and verification

For IMO, NVIDIA combined general, SFT and RL checkpoints in a natural-language generate-verify-refine system. It scored 30 out of 42, including full credit on four of six problems, above the official gold threshold of 29. NVIDIA said official IMO graders assessed the submitted proofs.

The math SFT corpus contained 414,890 quality-filtered examples across 15,818 proof problems, including generation, refinement, verification and meta-verification tasks. The RL specialist trained on 9,597 problems selected near the model’s capability frontier. NVIDIA said the SFT checkpoint was strongest in the initial search round, while RL delivered the best result from a single checkpoint; the final system used both alongside the general model. It used no formal prover, external tools or internet access, although a separate high-compute stage chose the final submission.

Why test-time compute mattered

NVIDIA’s results do not attribute the scores to fine-tuning alone. GenCorrect amplified the coding specialists’ gains through feedback rounds, while the IMO system used complementary checkpoints to generate, critique, score and revise proofs. The reported outcome therefore reflects the combined design of training data, post-training and inference-time search.

Published materials

The Nemotron Labs IMO 2026 collection includes SFT and RL checkpoints, both training datasets and the 200-problem Nemotron-IMO-Bench. NVIDIA also published the IMO pipeline, prompts and submitted proofs in NeMo-Skills. The Ultra-CC model, IOI training recipe and GenCorrect methodology are available through the linked Hugging Face and paper materials. Source: Hugging Face

Definition. Nemotron’s specialization recipe combines curated domain data, post-training and inference-time generation, evaluation and refinement to create task-specific systems.

EvaluationReported result
IOI 2026535.4 out of 600; stated gold threshold: 361.12; unofficial run.
IMO 202630 out of 42; official gold threshold: 29; official graders assessed proofs.

Key takeaways

  • The reported IOI 2026 result was 535.4 out of 600, above the stated 361.12 gold threshold, but was unofficial.
  • The IMO 2026 system scored 30 out of 42, including full credit on four of six problems, above the official gold threshold of 29.
  • NVIDIA attributed the results to curated data, SFT, RL where useful, and multi-stage inference rather than a base model alone.
  • GenCorrect generated, evaluated and refined coding answers; the IMO system generated, critiqued, scored and revised proofs.
  • NVIDIA published checkpoints, datasets, benchmarks and implementation materials for the workflows.

FAQ

Was NVIDIA's IOI 2026 result official?

No. NVIDIA described the IOI run as prospective and subject to contestant-like constraints, but unofficial, unsupervised and outside the official IOI ranking.

What score did Nemotron report at IMO 2026?

NVIDIA reported a score of 30 out of 42, including full credit on four of six problems; it said official IMO graders assessed the submitted proofs.

What methods were used for the reported results?

The systems used curated problems and reasoning traces, supervised fine-tuning, reinforcement learning where useful, and iterative inference-time generation, evaluation and refinement.

Did the IMO system use external tools or a formal prover?

NVIDIA said the system used no formal prover, external tools or internet access, though a separate high-compute stage chose the final submission.

Sources