TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Hugging Face says specialized Nemotron systems reached gold-medal level at the 2026 International Olympiad in Informatics and International Mathematical Olympiad. The IMO proofs were graded by official graders; the IOI score came from an unofficial run and was not part of the official ranking.
Hugging Face says two systems built from its Nemotron 3 model family reached gold-medal level at the 2026 International Olympiad in Informatics (IOI) and International Mathematical Olympiad (IMO). The systems scored 535.4 out of 600 at IOI and 30 out of 42 at IMO, but the IOI result was an unofficial benchmark run and did not enter the competition’s official ranking.
For IOI, the company used a competition-specific version of Nemotron-3-Ultra-CC, trained with supervised fine-tuning (SFT) and paired with GenCorrect, an iterative process for generating, evaluating and refining solutions. Hugging Face reports that the system ran prospectively under the same time, internet-access and submission constraints as human contestants. Its score exceeded the stated 361.12 gold threshold and the top human score of 498.27, though the run was unsupervised and unofficial.
For IMO, Hugging Face combined the general Nemotron 3 Ultra model with SFT and reinforcement-learning checkpoints in a natural-language proof system. The models generated candidate proofs, scored and critiqued them, then revised promising attempts. The company says official IMO graders awarded the submitted proofs 30 points, above the stated gold threshold of 29, with full credit on four of six problems. The system used no formal prover, external tools or internet access.
The projects used different specialist models and data. The IOI work drew on 22,000 programming problems and synthetic reasoning traces. The IMO SFT data included 414,890 quality-filtered examples from 15,818 proof problems, while the RL model was trained on 9,597 problems selected near the model’s capability frontier. These figures and results are reported by Hugging Face.
AI Research / Competition Report · 2026
One Model Family,
Two Gold-Level Results
Fine-tuning Nemotron for IOI and IMO produced two specialist systems at gold-medal level, according to Hugging Face. The results combine task-specific training with iterative search, while carrying different levels of official validation.
01 / Two specialist tracks
One family, distinct methods
Both projects adapted Nemotron 3, but the tasks called for different training data, outputs and evaluation procedures.
Nemotron-3-Ultra-CC
A competition-specific model trained with supervised fine-tuning, paired with GenCorrect to generate, evaluate and refine code solutions.
Checkpoint combination
The general Nemotron 3 Ultra model was combined with SFT and reinforcement-learning checkpoints in a natural-language proof system.
Specialize, then search
Models supplied candidate answers and critiques; iterative feedback helped select and revise promising attempts.
02 / Scorecard
Gold-level scores, different evidence
The score totals are notable, but their competition status and validation are not equivalent.
Programming under contest constraints
Hugging Face reports the run exceeded the 361.12 gold threshold and the top human score of 498.27. It ran prospectively under reported contestant constraints, but was unsupervised and excluded from official ranking.
Written mathematical proofs
Official graders awarded 30 points, above the stated 29-point gold threshold, with full credit on four of six problems. The system used no formal prover, external tools or internet access.
03 / Inference loop
From candidate to stronger answer
Hugging Face describes structured inference as a partner to post-training: candidates are generated, checked and improved before submission.
Propose
Create code candidates or written proof attempts from a specialist model.
Score
Assess candidate quality through task-specific checks and model feedback.
Find gaps
Use critiques to identify weak reasoning or opportunities to improve.
Revise
Iterate toward a stronger final solution or proof submission.
04 / Training footprint
Different data for different domains
Reported training sets reflect the distinct demands of executable programming and rigorous proof writing.
22,000 problems
Programming problems and synthetic reasoning traces supported the competition-specific fine-tuning effort.
414,890 examples
Quality-filtered examples drawn from 15,818 proof problems formed the supervised fine-tuning data.
9,597 problems
Selected near the model’s capability frontier for reinforcement-learning training.
05 / Development arc
From IOI 2025 to two subjects
Hugging Face describes the 2026 projects as an extension of earlier work on test-time computation. The new effort applied related ideas to code judged by executable tests and mathematics judged through written proofs.
06 / The team’s view
What the results suggest
Hugging Face frames the work as evidence that one model family can support specialists across demanding domains.
“Success at both points to something broader.”
— Hugging Face
07 / Limits & next steps
Promising results need scrutiny
The reported scores are meaningful evidence on specific competitions. Independent review would clarify how broadly the methods transfer.
What remains uncertain
- IOI was unofficial; the supplied account does not describe independent verification or a detailed audit.
- Broader performance on other competitions, unseen proof styles and real-world programming workloads is not established.
- Data selection, compute budgets and evaluation details may affect competition scores.
What could strengthen the case
- Independent groups reproducing the results.
- External scrutiny of the IOI procedure.
- Benchmark performance sustained on problems beyond the competitions.
Checkpoints and benchmarks ahead
Hugging Face says the Nemotron Labs IMO 2026 collection includes SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench: 200 olympiad-level problems. The company also points to an IMO paper and a NeMo-Skills repository. The supplied source leaves some repository release details incomplete.
08 / Key questions
At a glance
What did the systems score?
Hugging Face reports 535.4 out of 600 at IOI 2026 and 30 out of 42 at IMO 2026.
Were both results official?
No. Official IMO graders evaluated the proofs. The IOI score came from an unofficial, unsupervised run outside the official ranking.
How were the models adapted?
Task-specific data and post-training included supervised fine-tuning and, for some models, reinforcement learning. Both efforts added processes to generate, evaluate and refine answers.
What is being released?
The IMO collection is reported to include SFT and RL checkpoints, two training datasets and the 200-problem Nemotron-IMO-Bench.
Specialization Paired With Search
The results point to a method that combines domain-specific fine-tuning with inference procedures that check and improve candidate answers. In Hugging Face’s account, fine-tuning helped supply stronger solutions and critiques, while iterative feedback improved the final output. The company argues that adapting a shared base model can support specialist systems across coding and proof writing without training a separate foundation model for each field.
The scores are notable evidence of capability on two demanding competitions, but they measure different things and carry different levels of official validation. IMO submissions received official grading. IOI’s score was produced under competition-like constraints but outside the official ranking. Neither result alone establishes how the systems would perform across broader mathematical or programming tasks.
As an affiliate, we earn on qualifying purchases.
From IOI 2025 to Two Subjects
Hugging Face describes the work as an extension of its earlier IOI 2025 experiments, where test-time computation helped open-weight models reach gold-level performance. For IOI 2025, the company reports that a Nemotron-3-Nano-CC model rose from 130 points before post-training to 280 after SFT and 291 after reinforcement learning; GenCorrect lifted it to 468, above that year’s stated gold threshold of 438.3. An Ultra-CC version scored 502 with the same test-time strategy.
The 2026 efforts applied related ideas to separate tasks: programming problems requiring executable code and hidden-test performance, and mathematics problems requiring rigorous written proofs. Hugging Face says the IMO project found complementary strengths across SFT and RL checkpoints, so its final system combined them with the general model rather than relying on one checkpoint alone.
“Success at both points to something broader.”
— Hugging Face
programming problem sets for AI training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Scores
The IOI score was unofficial and excluded from the official ranking; the source describes a prospective, unsupervised benchmark run but does not provide independent verification in the supplied material. The report does not specify how the run was audited or whether the evaluation has been independently replicated.
Hugging Face’s account also does not establish how well these systems generalize to other competitions, unseen proof styles or real-world programming workloads. The source excerpt says the company is releasing IMO checkpoints, datasets and a 200-problem benchmark, but its final sentence about the associated repository is incomplete, leaving some release details unclear.
machine learning fine-tuning tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Checkpoints and Benchmarks Ahead
Hugging Face says its Nemotron Labs IMO 2026 collection includes the SFT and RL checkpoints, both training datasets and Nemotron-IMO-Bench, a benchmark of 200 olympiad-level problems. The company also points to an IMO paper describing the training and generate-verify-refine system, alongside a NeMo-Skills repository.
Those materials could let researchers inspect the methods and test the models against additional problems. The supplied source does not give a timetable for further releases or describe an independent evaluation of the IOI system.
As an affiliate, we earn on qualifying purchases.
Where I land
I see these results as credible evidence that fine-tuning and structured inference can work together on demanding, clearly scored tasks. The IMO result carries added weight because official graders evaluated the proofs. The IOI run is informative under its reported constraints, but its unofficial status means I would describe it as a benchmark result, not an official medal.
The strongest counterargument is that the report comes from the team that built the systems, and competition scores can depend on the details of data selection, compute budgets and evaluation setup. I would become more confident in the broader claim if independent groups reproduced the results, the IOI procedure received external scrutiny, and the released benchmark showed sustained performance on problems outside the competitions.
Key Questions
What did the Nemotron systems score?
Hugging Face reports 535.4 out of 600 for its IOI 2026 system and 30 out of 42 for its IMO 2026 system.
Were both results official competition placements?
No. The IMO proofs were graded by official IMO graders. The IOI score came from an unofficial, unsupervised run and was not included in the official IOI ranking.
How were the systems adapted?
The teams used task-specific data and post-training, including supervised fine-tuning and, for some models, reinforcement learning. They paired the models with processes that generate, evaluate and refine candidate answers.
What is being released?
Hugging Face says the Nemotron Labs IMO 2026 collection includes SFT and RL checkpoints, two training datasets and a 200-problem benchmark called Nemotron-IMO-Bench.
Source: Hugging Face
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
