DeepSeek V4.1 Flash benchmark results draw analyst questions over possible contamination
Two analysts argue that DeepSeek V4.1 Flash's results across benchmark vintages are consistent with public-benchmark contamination. The model also leads Artificial Analysis's new private evaluation, complicating the assessment.

TL;DR
- DeepSeek reports that V4.1 Flash scored 90.6% on Terminal-Bench 2.1, then 30.0% on 3.0 and 31.2% on 4.0. onusoz's chart makes that cross-version drop visible.
- The contamination allegation remains deliberately tentative: onusoz wrote that the model "looks" contaminated, then his clarification said the available data did not justify an accusation.
- Terminal-Bench version changes also change the test. onusoz's reply confirms that his visual uses ordinal rank rather than score distance, while the v4 benchmark revised tasks, compute allowances, instructions, environments, and verifiers.
- A private evaluation complicates a public-benchmark-only explanation: a LocalLLaMA post reported that V4.1 Flash beat Astra on Artificial Analysis's new test, whose test set is private.
- A separate two-round reliability check put V4.1 Flash at 91.6 in round one and 92.8 in round two, according to teortaxesTex's evaluation.
DeepSeek's model card says V4.1 Flash was trained from scratch on a 45T-token multimodal corpus and extended to a one-million-token context window at 34T tokens. Artificial Analysis's v4.3 update added AutomationBench-AA, an agentic workflow test with a private set. The Terminal-Bench v4.0 leaderboard describes a harder 66-task suite with recalibrated time and compute allowances.
Terminal-Bench versions
DeepSeek's model card reports all three Terminal-Bench results under its own harness, with minimal mode, one-million-token context, and maximum reasoning effort. The raw scores are 90.6 on v2.1, 30.0 on v3.0, and 31.2 on v4.0; among the seven models in the chart, V4.1 Flash moves from first to fourth.
The plot does not show the magnitude of those gaps. In a follow-up, onusoz said the vertical axis represents ranking rather than value, so a one-point gap and a 20-point gap occupy the same visual step.
Version 4.0 also changes more than the calendar. The benchmark's own description calls it a harder 66-task evaluation with recalibrated resource allowances and revised instructions, environments, and verifiers. A score change across these versions therefore combines task-set changes, evaluation conditions, and model behavior.
The contamination claim
Onusoz's theory is that V4.1 Flash performs unusually well on older public benchmark vintages, then loses ground on later ones that were less likely to have appeared in training data. He also raised the possibility for Kimi K3.
The claim stops short of attributing intent. onusoz's clarification says that "contaminated" was his weaker description of a visible pattern and did not prove anything, which is why he wrote "looks" rather than "is."
Before the launch, onusoz's request for private-benchmark results singled out Terminal-Bench 2.1 and asked users running private tests how the model compared with earlier DeepSeek releases. That supplies a provenance question, not evidence that DeepSeek knowingly trained on benchmark tasks.
The private counterpoint
Artificial Analysis says AutomationBench-AA uses a private test set and is one of 10 evaluations in Intelligence Index v4.3. That makes its result a distinct datapoint from the public Terminal-Bench versions in the chart.
DeepSeek V4.1 Flash beats Astra on AA's new benchmark
0 comments
The Reddit post reported that V4.1 Flash beat Astra on the new AA benchmark. A private test result does not establish what appeared in the model's broader training corpus, but it means the available evidence is not limited to performance on older public task sets.
Reliability and planning
A separate evaluation from teortaxesTex tested output stability across two rounds. Its weighted reference score put V4.1 Flash at 92.2, versus 82.9 for V4 Pro 0813 after Pro's coding score collapsed in the second round; the accompanying methodology calls that number a routing reference rather than a new capability benchmark.
Early qualitative reports remain mixed. Onusoz said V4.1 Flash felt worse at planning but equally good at long-horizon agentic work in his early impression.
The Together AI listing says the hosted model exposes one-million-token context and native multimodal input, providing a public endpoint for the architecture at the center of the benchmark debate.