Google rolls out Gemini 4 Argon to cyber defenders
Google is rolling out Gemini 4 Argon through Fairwind to government users, vetted cyber defenders, and trusted testers. The model supports up to 1M output tokens and costs $2/M input and $10/M output.

TL;DR
- Google is rolling out Gemini 4 Argon to trusted cyber defenders through Fairwind; paid API customers and Google AI Ultra subscribers come later, according to Google's rollout post.
- Output headroom rises from 64K to 1M tokens per trajectory, while Google DeepMind's announcement describes the capacity for long, multi-step work.
- Introductory pricing is $2 per million input tokens and $10 per million output tokens, with a later increase to $4 and $20, per Logan's launch post and the pricing follow-up.
- The coding results split sharply: Google's comparison puts Argon ahead on DeepSWE but behind on FrontierSWE and Terminal-bench 4.0, as the benchmark table shows.
- Independent testing puts Argon at 53 on the Artificial Analysis Intelligence Index, up from Gemini 3.8 Flash's 41, while Artificial Analysis's results put its discounted cost at $1.99 per task.
Google's launch post describes agents rewriting 32,000 lines of video-decoder SIMD code. The evaluation methodology reveals that the DeepSWE and Terminal-bench numbers in Google's comparison were computed with different sourcing arrangements for Argon and its rivals. A Vals leaderboard lists a higher legal-agent score than any model in Google's four-model chart.
What shipped
- Initial access: Trusted cyber defenders through Fairwind, alongside Google internal use and a U.S. government pre-release access process, per Google's rollout announcement. Fairwind's program description includes government agencies, selected Google Cloud customers and cybersecurity partners.
- Next access: Paid API customers and Google AI Ultra subscribers are next, with no date in Google's launch post. Developers, enterprises and consumers follow the initial testing phase.
- Token ceiling: 1M output tokens per trajectory, up from 64K, according to Google's model announcement.
- Prices: $2 input and $10 output per million tokens initially; $4 and $20 after the promotion. Cached input gets a 95% discount, according to Logan's pricing announcement and Google's pricing footnote.
- Cyber configuration: Google says trusted defenders and internal teams get the model without cyber guardrails for defensive work, according to Google's announcement.
Benchmarks that moved
First-party
- DeepSWE v1.1: Claude Opus 5.5's 74.2% → Argon's 77.9%, +3.7 points, in Google's comparison table.
- AutomationBench: Claude Opus 5.5's 42.5% → Argon's 51.3%, +8.8 points, in Google's chart.
- GraphWalks at 256K to 1M context: GPT-6 Astra's 71.8% → Argon's 84.2% F1, +12.4 points, in Google's table.
- LVBench long-video understanding: GPT-6 Astra's 87.5% → Argon's 91.7%, +4.2 points, in Google's table.
Third-party evaluators
- Artificial Analysis Intelligence Index: Gemini 3.8 Flash's 41 → Argon's 53, +12 points, per Artificial Analysis's evaluation; Argon ties GPT-6 Astra at 53 and trails Opus 5.5 at 58.
- AutomationBench-AA: Claude Sonnet 5.5's 71.3% → Argon's 77.5%, +6.2 points, per Artificial Analysis's testing.
- Vals Index: Claude Opus 5.5's 66.97% → Argon's 68.90%, +1.93 points, per Vals AI's leaderboard.
- Vibe Code Bench, perfect apps: Gemini 3.8 Flash's 16 → Argon's 30, +14 apps, per Vals AI's results.
- Text Arena: Claude Opus 4.6 High's 1,505 → Argon High's 1,525, +20 Elo, per Arena's ranking.
Customer-reported
No scored customer before-and-after evaluation appears in Google's launch post; its Wiz vulnerability case is qualitative.
Google's methodology says Argon's DeepSWE score used a mini-swe agent harness, while the other DeepSWE scores came from a public leaderboard and rival system cards. Google's Terminal-bench 4.0 score was also self-computed, while rival figures came from its public leaderboard. Those sourcing differences belong alongside the headline comparisons.
Where it regressed
Google's own chart contains the coding and computer-use losses behind the broad frontier-performance claim:
- FrontierSWE v2: GPT-6 Astra's 65.5% → Argon's 55.0%, −10.5 points, in Google's benchmark table.
- Terminal-bench 4.0: Claude Opus 5.5's 66.4% → Argon's 57.4%, −9.0 points, in the same table.
- Terminal-Bench Science 0.1: GPT-6 Astra's 68.1% → Argon's 57.6%, −10.5 points, in Google's results.
- OSWorld-2.0 offline-subset partial score: GPT-6 Astra's 72.6% → Argon's 69.2%, −3.4 points, in Google's results.
Arena's separate WebDev ranking puts Argon High at eighth with 1,679 points, even though it leads Text Arena; Arena's release-day results report a 96-point WebDev improvement over Gemini 3.8 Flash. Artificial Analysis also found that AA-Omniscience accuracy fell from Gemini 3.1 Pro Preview's 55% to Argon's 50%, −5 points, alongside a 15% hallucination rate for Argon, per its independent analysis.
Under the hood
- Long decode: The advertised 1M is an output ceiling, separate from the 1M input context reported by Artificial Analysis. Its testers used Long Decode Continuation to pause and resume generation across follow-up API calls without request timeouts; Google's launch description calls the million-token span a single trajectory. Vals's evaluation used a 262,144-token maximum output setting, per its configuration disclosure, so that run did not exercise the advertised ceiling. Google has not published a public Argon API contract for the continuation feature in the launch post.
- Task economics: Artificial Analysis measured 62K output tokens per task for Argon versus GPT-6 Astra's 27K, +35K tokens. Discounted Argon costs $1.99 per Index task versus Astra's $3.26, but the estimated post-promotion Argon cost rises to $3.98, per Artificial Analysis's cost analysis. A Hacker News discussion also pointed to GPT-6.1 Sol's lower $0.72 cost per task despite its one-point-lower Index score.
- Model identity: A Google spokesperson told Reuters Argon is larger than earlier top-tier Pro models. Neither that statement nor Google's launch post specifies its parameter count or confirms the training lineage claimed in one observer's pretraining speculation.
- Safeguards: Google's launch post describes cyber and CBRN misuse refusals for broader use, adversarial training for indirect prompt injection, monitoring of reasoning and actions that can stop a run, and sealed sandboxes for high-risk tests. The cyber-guardrail exception for trusted defenders is part of that staged deployment.
Contested claims
Claim: Google's launch calls Argon a leading model on Harvey's Legal Agent Benchmark, and its four-model chart gives Argon 19.6% against Claude Fable 5.1's 6.7%.
Cited by: Google's announcement and evaluation methodology, which says its Harvey figures came from Vals AI.
Counter: BlackHC's comparison found Muse Spark 1.2 at 25.42% on the Vals Harvey leaderboard, above Argon's reported 19.6%.
Evidence so far: The Vals page describes a held-out set where a task passes only when every grading criterion passes, but its displayed leaderboard does not yet list Argon. The dates and model coverage therefore leave the all-model rank unresolved; Google's chart establishes a lead over the three rivals it selected.
Vibe Check
- On FrontierSWE, one evaluator found Argon repeatedly saying “Eureka!” and criticizing its own mistakes, including a stray
if (1)left in a Game Boy music app. - In another hands-on use, a Google tester reported that Argon played board games well, without publishing a scored comparison in that post.
- Inside Google, Mirrokni described using agents built on Argon for memory optimization, code migration, math and long-horizon engineering through an internal
/teamworkworkflow.
Google's production runs
Google's launch post offers unusually specific internal engineering examples:
- Data-center agents analyzed fleet-wide profiling telemetry and found changes expected to free more than 300 TiB of memory when rolled out; Google estimates 500 TiB to 1 PiB in eventual savings.
- For
libgav1, agents replaced 32,000 lines of SIMD in an existing Rust port with safe Rust that the compiler can vectorize. Google reports identical video output and 2.7× the existing Rust port's speed, after profile-guided experiments. - C/C++-to-Rust migration work extends from
re2andlibgav1to more than 800,000 lines of the Fuchsia Zircon kernel. Google says the critical rewrites still face automated and manual audits, emulation and review before production. - A quantum-computing subroutine optimization beat a published spacetime-resource baseline by 40% in minutes, according to Google.