Skip to content
AI Primer
breaking

Reports rank Gemini 4 Argon highly on four engineering benchmarks

Reports place Gemini 4 Argon at 77.9% on DeepSWE, 57.6% on Terminal-Bench, 77.5% on AutomationBench-AA, and 68% on CWE-Bench. The reports also cite lower cost or token use than rivals.

6 min read
Reports rank Gemini 4 Argon highly on four engineering benchmarks
Reports rank Gemini 4 Argon highly on four engineering benchmarks

TL;DR

  • Argon’s strongest reported engineering results are 77.9% on DeepSWE v1.1, 57.6% on Terminal-Bench 4.0, 77.5% on AutomationBench-AA, and 68% on CWE-bench v1, where it ties for first. ArtificialAnlys and ValsAI provide the benchmark results.
  • Google’s own comparison shows a broad lead in knowledge work and long-context tests, alongside losses on FrontierSWE, Terminal-Bench, science, ML engineering, and OSWorld. GoogleDeepMind
  • The cost advantage comes mainly from launch pricing: Artificial Analysis measured $1.99 per Intelligence Index task, against $3.26 for GPT-6 Astra and $5.98 for Claude Opus 5.5. ArtificialAnlys
  • Argon remains restricted to trusted cyber defenders while Google iterates on safeguards and prepares access for paid API customers and AI Ultra subscribers. Google

Google’s official announcement pairs a 1M output ceiling with a launch price that later rises from $2/$10 to $4/$20 per million input and output tokens. Artificial Analysis’s independent report finds a 15% hallucination rate but only 50% answer accuracy, while a Hacker News discussion is already questioning how much of the result is benchmark selection. The release surprised kimmonismus, who said no rumors or benchmark leaks preceded it.

Headline benchmark scores

The four headline numbers come from multiple scorecards, with Google’s launch table and Artificial Analysis using related but not identical benchmark variants.

  • DeepSWE v1.1: 77.9%. Google describes it as a test of real-world, long-horizon software engineering in its launch post.
  • Terminal-Bench 4.0: 57.6%, up from 19.0% for the earlier Gemini result, according to ValsAI’s comparison.
  • AutomationBench-AA: 77.5%, the top score in Artificial Analysis’s independent run. Google’s own Zapier AutomationBench result is 51.3%, a separately named scorecard in the official announcement.
  • CWE-bench v1: 68%, tied for first on vulnerability remediation, according to haider1 and Google’s results.

The benchmark spread

Google’s comparison table gives Argon clear leads, but it also records specific losses that disappear in a simple “beats the competition” summary. The largest reported advantages are in professional and long-context work:

  • Vals Index: 68.9%, versus 67.0% for Claude Opus 5.5 and 63.1% for GPT-6 Astra.
  • Vals Finance Agent v2: 65.4%, versus 58.6% for Opus 5.5 and 53.5% for Astra.
  • Harvey’s Legal Agent Benchmark: 19.6%, versus 6.7% for Claude Fable 5.1 and 3.8% for Opus 5.5.
  • GraphWalks from 256K to 1M tokens: 84.2%, versus 71.8% for Astra and 66.8% for Opus 5.5.
  • LVBench: 91.7%, versus 87.5% for Astra and 83.7% for Opus 5.5.

The same table shows a 55.0% FrontierSWE v2 score, 10.5 points below Astra’s 65.5%, and 57.4% on Terminal-Bench 4.0, 9 points below Opus 5.5’s 66.4%. Argon also trails Astra on Terminal-Bench Science, 57.6% to 68.1%, and OSWorld-2.0, 69.2% to 72.6%.

Argon’s 45.3% PostTrainBench result beats Astra’s 44.3% but trails Opus 5.5’s 49.3%, while karinanguyen reported the score as more than double Gemini 3.1 Pro’s 21.99%. BlackHC separately compared Argon’s 19.6% Harvey score with Muse Spark 1.2’s 25.42% and questioned what the cross-model numbers represented.

Independent scorecards

Artificial Analysis placed Argon at 53 on its Intelligence Index, tied with GPT-6 Astra and Claude Fable 5.1, one point above GPT-6.1 Sol, but below Claude Sonnet 5.5 at 56 and Opus 5.5 at 58. Its AA-Omniscience result is unusually asymmetric: a 15% hallucination rate versus 51% for Astra, paired with 50% answer accuracy versus Astra’s 63%.

The token profile complicates the cheap-model story. Argon averaged 62,000 output tokens per Intelligence Index task, compared with 27,000 for Astra and 119,000 for Opus 5.5. Artificial Analysis attributed the $1.99 launch-period cost per task primarily to lower token prices, not lower token consumption; its standard-price estimate is $3.98 after the promotion ends.

Arena’s scorecards split by task type. arena put Argon first in Text Arena at 1,525, while its Code Arena WebDev result landed eighth at 1,679. A separate arena report placed it eighth in Agent Arena at a preliminary +7.92% net improvement across 3,000 real-world sessions, with a $0.62 median cost per task.

That practical split also appears in early hands-on testing. arena cited 33% lower cost per task than GPT-6.1 Sol and 70% lower cost than Opus 5.5, while Theo’s independent Terminal-Bench thread reported a result better than Opus 5.5 at roughly one-thirtieth the price, then still called Opus his default coding model.

Output limits and rollout

Argon’s headline capacity is an output limit, not a 1M-token context-window claim. Google raised the maximum output from 64K to 1M tokens, and rohanpaul_ai compared that per-response ceiling with 128K for Astra, Opus 5.5, and Fable 5.1.

The introductory rate is $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%, as OfficialLoganK announced. Google’s launch post says the later rate is $4 per million input tokens and $20 per million output tokens.

Access is deliberately narrow. Google said the first cohort would be trusted cyber defenders in its Fairwind Program, while tulseedoshi and koraykv described the release as an early testing phase. testingcatalog reported that paid API customers and Google AI Ultra subscribers were next in line, and demishassabis framed broader availability as dependent on the safety rollout.

Internal engineering workloads

The most concrete usage evidence comes from inside Google. The company says Argon agents analyzed fleet-wide profiling telemetry and freed more than 300 TiB of data-center memory after rollout, with estimated total savings of 500 TiB to 1 PiB. They also worked on C and C++ to Rust migrations ranging from core libraries to more than 800,000 lines of the Fuchsia Zircon kernel.

For libgav1, Google says the agents replaced 32,000 lines of SIMD code with safe Rust after repeated profile-guided experiments. The resulting decoder ran 2.7 times faster than the existing Rust port while preserving identical video output, according to the official account and ai_for_success.

The internal workload also extends beyond code migration. Google says Argon helped quantum researchers beat a published quantum algorithm optimization baseline by 40% in minutes. mirrokni said Google was running internal RSI loops and agents, including an internal version of /teamwork in AGY, for memory optimization, code migration, mathematics, and other large software engineering tasks.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
TL;DR1 post
Headline benchmark scores2 posts
The benchmark spread1 post
Independent scorecards1 post
Output limits and rollout6 posts
Share on X