SOOFI revises report after GPQA removal, drawing new eval-leakage criticism
Julius Jitsev said SOOFI removed GPQA and its capability index after feedback but still compared against Nemotron 3 Nano using benchmarks seen in training. He argued the remaining English and German scores are compromised.

TL;DR
- SOOFI’s revision removed GPQA and the capability index, but the remaining benchmark dispute shifted to the Nemotron 3 Nano comparison, according to JJitsev’s GPQA note and JJitsev’s broader critique.
- QA-base is the crux: JJitsev’s dataset comparison pointed to MBPP, SocialIQA, MMLU-Pro, and PIQA as common eval sets that SOOFI had seen while Nemotron 3 Nano had not.
- The English code split is the cleanest example: JJitsev’s Table 5 note said SOOFI looked stronger on MBPP, which it trained on, while LBPP and HumanEval favored Nemotron.
- The German results raised the same issue with extra repetition, since JJitsev’s German comparison said SOOFI trained on all German rephrased evals 10 times while Nemotron had seen none of them.
The official arXiv v3 report now says both English and German aggregate charts exclude GPQA and a withdrawn held-out benchmark group. The QA-base dataset card is unusually direct: 9.6M rows, 4.82 GB, 25 benchmark collections, and 122 JSONL files per language, with questions, options, and answers formatted as standalone passages. The Decoder’s update reported the pipeline bug: GPQA had no separate training set on Hugging Face, so test material under the default train label entered practice data.
GPQA removal
SOOFI’s visible concession was narrow: JJitsev said the revision removed GPQA from evaluation and gave up the capability index. The official v3 report also says its main aggregate charts exclude GPQA and a withdrawn held-out benchmark group.
JJitsev’s complaint is that the revised match against Nemotron 3 Nano still remains invalid, because SOOFI-S had seen many more eval sets during training, according to his broader critique.
QA-base overlap
The QA-base dataset card describes normalized and paraphrased splits of 25 NLP benchmarks across English, German, French, Spanish, and Italian, intended for base model pretraining. JJitsev singled out four common eval sets that appeared in QA-base while Nemotron 3 Nano allegedly did not see them:
- MBPP
- SocialIQA
- MMLU-Pro
- PIQA
The comparison is unusually sensitive because Soofi S adopts the Nemotron 3 Nano architecture without modification, a design choice the arXiv report frames as scientific control for measuring the data recipe. The Nemotron model card also describes Nvidia’s release as open weights, training data, and recipes, which makes the baseline’s disclosed data history part of the argument.
English code split
JJitsev pointed to Table 5 as the visible pattern: benchmark exposure tracks the direction of the score gap.
- MBPP: SOOFI trained on the eval set, and SOOFI scores higher.
- LBPP: SOOFI did not train on the eval set, and scores are lower.
- HumanEval: SOOFI did not train on the eval set, and Nemotron scores higher.
That is the bookmarkable bit for eval hygiene: the suspicious signal is not one contaminated benchmark, but a score split aligned with which benchmark families entered pretraining.
German repeated evals
The German side adds a repetition detail. JJitsev said SOOFI trained on all German rephrased evals with 10x repetition, while Nemotron 3 Nano had not seen any of them.
His example mirrors the English code split:
- MBPP-DE: SOOFI trained on the German rephrased evals, and SOOFI leads.
- HumanEval-DE: Nemotron leads on the example JJitsev gave, where the contamination advantage does not apply.
Frontier-level claim
JJitsev said the updated report still uses eval-leakage-contaminated numbers to support the repeated “frontier-level champion” framing. He also argued it remains unclear whether SOOFI-S actually beats its Nemotron 3 Nano origin after roughly 2T more tokens.
The official arXiv abstract still says Soofi S obtains the highest English and German evaluation scores among fully open models, ahead of OLMo 3 32B and Apertus 70B. The Decoder’s update reported SOOFI’s position that GPQA was removed, all 16 models were recalculated, rankings did not change, and roughly 152,000 individual results were made available for verification.
Open-release framing
JJitsev also objected to the report’s taxonomy, saying Nemotron 3 Nano was cast as open weights even though Nvidia disclosed both data and training stack. Nvidia’s model card describes Nemotron as a family of open models with open weights, training data, and recipes.
SOOFI’s own release status is still gated. The Soofi-S model card calls the checkpoint a closed beta research artifact, says it is not an open release, and says the final model will be released openly under a permissive license later.