AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Hugging Face says specialized Nemotron systems reached gold-medal level at the 2026 International Olympiad in Informatics and International Mathematical Olympiad. The IMO proofs were graded by official graders; the IOI score came from an unofficial run and was not part of the official ranking.

Hugging Face says two systems built from its Nemotron 3 model family reached gold-medal level at the 2026 International Olympiad in Informatics (IOI) and International Mathematical Olympiad (IMO). The systems scored 535.4 out of 600 at IOI and 30 out of 42 at IMO, but the IOI result was an unofficial benchmark run and did not enter the competition’s official ranking.

For IOI, the company used a competition-specific version of Nemotron-3-Ultra-CC, trained with supervised fine-tuning (SFT) and paired with GenCorrect, an iterative process for generating, evaluating and refining solutions. Hugging Face reports that the system ran prospectively under the same time, internet-access and submission constraints as human contestants. Its score exceeded the stated 361.12 gold threshold and the top human score of 498.27, though the run was unsupervised and unofficial.

For IMO, Hugging Face combined the general Nemotron 3 Ultra model with SFT and reinforcement-learning checkpoints in a natural-language proof system. The models generated candidate proofs, scored and critiqued them, then revised promising attempts. The company says official IMO graders awarded the submitted proofs 30 points, above the stated gold threshold of 29, with full credit on four of six problems. The system used no formal prover, external tools or internet access.

The projects used different specialist models and data. The IOI work drew on 22,000 programming problems and synthetic reasoning traces. The IMO SFT data included 414,890 quality-filtered examples from 15,818 proof problems, while the RL model was trained on 9,597 problems selected near the model’s capability frontier. These figures and results are reported by Hugging Face.

At a glance
reportWhen: Reported after the 2026 competitions
The developmentHugging Face reported that systems fine-tuned from Nemotron 3 scored above the gold thresholds at IOI 2026 and IMO 2026.
One Model Family, Two Gold-Level Results

AI Research / Competition Report · 2026

One Model Family,
Two Gold-Level Results

Fine-tuning Nemotron for IOI and IMO produced two specialist systems at gold-medal level, according to Hugging Face. The results combine task-specific training with iterative search, while carrying different levels of official validation.

Reported2026After the competitions
IOI score535.4Of 600 points · unofficial
IMO score30Of 42 points · official grading
IMO full marks4 / 6Problems receiving full credit

01 / Two specialist tracks

One family, distinct methods

Both projects adapted Nemotron 3, but the tasks called for different training data, outputs and evaluation procedures.

IOI · Programming

Nemotron-3-Ultra-CC

A competition-specific model trained with supervised fine-tuning, paired with GenCorrect to generate, evaluate and refine code solutions.

IMO · Proof writing

Checkpoint combination

The general Nemotron 3 Ultra model was combined with SFT and reinforcement-learning checkpoints in a natural-language proof system.

Shared approach

Models supplied candidate answers and critiques; iterative feedback helped select and revise promising attempts.

02 / Scorecard

Gold-level scores, different evidence

The score totals are notable, but their competition status and validation are not equivalent.

IOI 2026 · Unofficial

Programming under contest constraints

535.4/ 600 POINTS

Hugging Face reports the run exceeded the 361.12 gold threshold and the top human score of 498.27. It ran prospectively under reported contestant constraints, but was unsupervised and excluded from official ranking.

IMO 2026 · Official grading

Written mathematical proofs

30/ 42 POINTS

Official graders awarded 30 points, above the stated 29-point gold threshold, with full credit on four of six problems. The system used no formal prover, external tools or internet access.

Reading the comparison: IMO is an officially graded result. IOI is a competition-like benchmark result, not an official placement. Neither score alone establishes performance across broader coding or mathematics tasks.

03 / Inference loop

From candidate to stronger answer

Hugging Face describes structured inference as a partner to post-training: candidates are generated, checked and improved before submission.

1 Generate

Propose

Create code candidates or written proof attempts from a specialist model.

2 Evaluate

Score

Assess candidate quality through task-specific checks and model feedback.

3 Critique

Find gaps

Use critiques to identify weak reasoning or opportunities to improve.

4 Refine

Revise

Iterate toward a stronger final solution or proof submission.

04 / Training footprint

Different data for different domains

Reported training sets reflect the distinct demands of executable programming and rigorous proof writing.

IOI data

22,000 problems

Programming problems and synthetic reasoning traces supported the competition-specific fine-tuning effort.

IMO · SFT

414,890 examples

Quality-filtered examples drawn from 15,818 proof problems formed the supervised fine-tuning data.

IMO · RL

9,597 problems

Selected near the model’s capability frontier for reinforcement-learning training.

05 / Development arc

From IOI 2025 to two subjects

Hugging Face describes the 2026 projects as an extension of earlier work on test-time computation. The new effort applied related ideas to code judged by executable tests and mathematics judged through written proofs.

2025Nemotron-3-Nano-CC: 130 points before post-training, 280 after SFT, and 291 after reinforcement learning.
+ SearchGenCorrect raised the reported score to 468, above that year’s stated gold threshold of 438.3. An Ultra-CC version scored 502 with the same test-time strategy.
2026Related specialization and inference methods were applied to separate IOI programming and IMO proof-writing systems.

06 / The team’s view

What the results suggest

Hugging Face frames the work as evidence that one model family can support specialists across demanding domains.

“Success at both points to something broader.”

— Hugging Face

07 / Limits & next steps

Promising results need scrutiny

The reported scores are meaningful evidence on specific competitions. Independent review would clarify how broadly the methods transfer.

What remains uncertain

  • IOI was unofficial; the supplied account does not describe independent verification or a detailed audit.
  • Broader performance on other competitions, unseen proof styles and real-world programming workloads is not established.
  • Data selection, compute budgets and evaluation details may affect competition scores.

What could strengthen the case

  • Independent groups reproducing the results.
  • External scrutiny of the IOI procedure.
  • Benchmark performance sustained on problems beyond the competitions.

Checkpoints and benchmarks ahead

Hugging Face says the Nemotron Labs IMO 2026 collection includes SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench: 200 olympiad-level problems. The company also points to an IMO paper and a NeMo-Skills repository. The supplied source leaves some repository release details incomplete.

Research materials

08 / Key questions

At a glance

What did the systems score?

Hugging Face reports 535.4 out of 600 at IOI 2026 and 30 out of 42 at IMO 2026.

Were both results official?

No. Official IMO graders evaluated the proofs. The IOI score came from an unofficial, unsupervised run outside the official ranking.

How were the models adapted?

Task-specific data and post-training included supervised fine-tuning and, for some models, reinforcement learning. Both efforts added processes to generate, evaluate and refine answers.

What is being released?

The IMO collection is reported to include SFT and RL checkpoints, two training datasets and the 200-problem Nemotron-IMO-Bench.

Specialization Paired With Search

The results point to a method that combines domain-specific fine-tuning with inference procedures that check and improve candidate answers. In Hugging Face’s account, fine-tuning helped supply stronger solutions and critiques, while iterative feedback improved the final output. The company argues that adapting a shared base model can support specialist systems across coding and proof writing without training a separate foundation model for each field.

The scores are notable evidence of capability on two demanding competitions, but they measure different things and carry different levels of official validation. IMO submissions received official grading. IOI’s score was produced under competition-like constraints but outside the official ranking. Neither result alone establishes how the systems would perform across broader mathematical or programming tasks.

Amazon

AI development training datasets

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From IOI 2025 to Two Subjects

Hugging Face describes the work as an extension of its earlier IOI 2025 experiments, where test-time computation helped open-weight models reach gold-level performance. For IOI 2025, the company reports that a Nemotron-3-Nano-CC model rose from 130 points before post-training to 280 after SFT and 291 after reinforcement learning; GenCorrect lifted it to 468, above that year’s stated gold threshold of 438.3. An Ultra-CC version scored 502 with the same test-time strategy.

The 2026 efforts applied related ideas to separate tasks: programming problems requiring executable code and hidden-test performance, and mathematics problems requiring rigorous written proofs. Hugging Face says the IMO project found complementary strengths across SFT and RL checkpoints, so its final system combined them with the general model rather than relying on one checkpoint alone.

“Success at both points to something broader.”

— Hugging Face

Amazon

programming problem sets for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Scores

The IOI score was unofficial and excluded from the official ranking; the source describes a prospective, unsupervised benchmark run but does not provide independent verification in the supplied material. The report does not specify how the run was audited or whether the evaluation has been independently replicated.

Hugging Face’s account also does not establish how well these systems generalize to other competitions, unseen proof styles or real-world programming workloads. The source excerpt says the company is releasing IMO checkpoints, datasets and a 200-problem benchmark, but its final sentence about the associated repository is incomplete, leaving some release details unclear.

Amazon

machine learning fine-tuning tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Checkpoints and Benchmarks Ahead

Hugging Face says its Nemotron Labs IMO 2026 collection includes the SFT and RL checkpoints, both training datasets and Nemotron-IMO-Bench, a benchmark of 200 olympiad-level problems. The company also points to an IMO paper describing the training and generate-verify-refine system, alongside a NeMo-Skills repository.

Those materials could let researchers inspect the methods and test the models against additional problems. The supplied source does not give a timetable for further releases or describe an independent evaluation of the IOI system.

Amazon

AI proof generation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where I land

I see these results as credible evidence that fine-tuning and structured inference can work together on demanding, clearly scored tasks. The IMO result carries added weight because official graders evaluated the proofs. The IOI run is informative under its reported constraints, but its unofficial status means I would describe it as a benchmark result, not an official medal.

The strongest counterargument is that the report comes from the team that built the systems, and competition scores can depend on the details of data selection, compute budgets and evaluation setup. I would become more confident in the broader claim if independent groups reproduced the results, the IOI procedure received external scrutiny, and the released benchmark showed sustained performance on problems outside the competitions.

Key Questions

What did the Nemotron systems score?

Hugging Face reports 535.4 out of 600 for its IOI 2026 system and 30 out of 42 for its IMO 2026 system.

Were both results official competition placements?

No. The IMO proofs were graded by official IMO graders. The IOI score came from an unofficial, unsupervised run and was not included in the official IOI ranking.

How were the systems adapted?

The teams used task-specific data and post-training, including supervised fine-tuning and, for some models, reinforcement learning. They paired the models with processes that generate, evaluate and refine candidate answers.

What is being released?

Hugging Face says the Nemotron Labs IMO 2026 collection includes SFT and RL checkpoints, two training datasets and a 200-problem benchmark called Nemotron-IMO-Bench.

Source: Hugging Face

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Market impact of the EU AI Act’s transparency obligations: vertical benefits and competitive dynamics

AIThis post was created with the assistance of artificial intelligence (AI). Prime…

Bridging the Global AI Talent Gap: Causes, Challenges and Pathways to Inclusive Workforce Development

AIThis post was created with the assistance of artificial intelligence (AI). Prime…

Inside OpenAI’s Enterprise Data Stack: What Happens to Your Company Data in 2026

AIThis post was created with the assistance of artificial intelligence (AI).OpenAI says…

Apple Is Reaching for Chinese Memory. Europe Doesn’t Even Have That Option.

AIThis post was created with the assistance of artificial intelligence (AI).This week…