Google opens Gemini 4 Argon to trusted testers with 1M-token output limit
Google says Gemini 4 Argon is rolling out first to trusted testers, including cyber defenders. The model advertises a one-million-token output limit, and broader access has not been announced.

TL;DR
- Gemini 4 Argon is rolling out first to trusted cyber defenders through Fairwind, Google DeepMind's announcement says; Google has given no date for general access.
- Its output-token ceiling rises from 64K to 1M, according to minchoi's breakdown. That is the amount it can generate, separate from how much it can read.
- Google's benchmark table puts Argon ahead on legal work and behind rivals on some coding tests. Independent scores add a less uniform picture.
- Introductory API pricing is $2 per million input tokens and $10 per million output tokens, as Logan's launch thread notes; Google's announced post-promotion rates double both figures.
A quantum optimization beat a published baseline by 40% in Google's internal testing. The same launch account describes agents rewriting a video decoder's SIMD code, while Artificial Analysis tested a way to pause and resume unusually long outputs across API calls.
What shipped
Google announced Argon on September 30 for long-running coding, enterprise knowledge work and cyber defense. The official announcement distinguishes today's restricted rollout from a later release for developers, enterprises and consumers.
- Access now: trusted cyber defenders and other initial testers through Fairwind. Logan's reply promises wider rollout “as soon as possible,” without a date.
- Access next: paid API customers and Google AI Ultra subscribers are first in line when broader release begins, according to Google's rollout plan.
- Price when it opens: $2 input / $10 output per million tokens initially, rising to $4 / $20 after the introductory period. Cached input is 95% off the introductory input price. Google gives no promotion end date in its announcement.
One paying user in a Hacker News discussion reported that the Gemini app still offered an earlier model on launch day. Public hands-on reports of Argon remain scarce while access is restricted.
Benchmarks that moved
The comparisons below identify the older model or rival before the arrow. Google's evaluation notes distinguish its own test runs from public leaderboards and say Argon generally used the highest thinking setting.
First-party
- DeepSWE v1.1, GPT-6 Astra → Argon: 74.1% → 77.9%, +3.8 points, in Google's comparison; Google computed Argon's score with a mini-swe agent harness.
- GraphWalks at 256K to 1M input tokens, GPT-6 Astra → Argon: 71.8% → 84.2% F1, +12.4 points, in Google's table; Google computed all four models' scores on this test.
Third-party evaluators
- Harvey's Legal Agent Benchmark, Claude Fable 5.1 → Argon: 6.7% → 19.6%, +12.9 points, in the launch chart; Google's methodology attributes these results to Vals AI.
- AutomationBench, Claude Opus 5.5 → Argon: 42.5% → 51.3%, +8.8 points, in Google's chart; the scores come from Zapier's public leaderboard.
- Artificial Analysis Intelligence Index, Gemini 3.1 Pro Preview → Argon: 30 → 53, +23 points, in Artificial Analysis's independent evaluation. Argon tied GPT-6 Astra at 53 on that index.
- Artificial Analysis Terminal Bench 4, Gemini 3.1 Pro Preview → Argon: 4% → 57%, +53 points, in the same evaluation.
Customer-reported
No numerical customer benchmark accompanies Google's account of Wiz's early use.
Where it regressed
Google's own table has conspicuous losses alongside its wins. On FrontierSWE v2, GPT-6 Astra → Argon is 65.5% → 55.0%, −10.5 points; on Terminal-bench 4.0, Claude Opus 5.5 → Argon is 66.4% → 57.4%, −9.0 points, according to the published comparison.
Against Google's prior non-Flash model, Argon's AA-Omniscience accuracy fell from 55% to 50%, −5 points, while its hallucination rate was 15%, versus GPT-6 Astra's 51%, according to Artificial Analysis. Google's methodology says comparison scores often come from providers' self-reported results and that some tests use different harnesses or settings.
Under the hood
The advertised 1M-token limit applies to output, up from a previous 64K ceiling, as Google's launch thread specifies. Artificial Analysis separately lists a 1M-token input context window and tested Long Decode Continuation, which resumes generation across follow-up calls when a long request would time out.
In the evaluator's tests, Argon used an average of 62K output tokens per task versus GPT-6 Astra's 27K, +35K tokens. Its reported cost per Intelligence Index task was $1.99 at promotional rates and would rise to $3.98 at the announced standard rates, according to Artificial Analysis. Google's announcement gives no date for that change.
Google also says its models undergo broad internal testing: Logan's reply described thousands of software engineers testing new Gemini revisions for weeks, while framing that figure as an assumption about typical releases rather than an Argon-specific count.
Google agent runs
Google's internal examples go beyond benchmark tasks:
- Fleet memory: agents analyzed profiling telemetry and found optimizations expected to free more than 300 TiB once rolled out, with 500 TiB to 1 PiB in estimated total savings, a figure Linus Ekenstam highlighted.
- Code migration: agents are working on C/C++ to Rust conversions from small libraries up to Fuchsia's 800K-plus-line Zircon kernel. Google says these rewrites still face automated and manual auditing, emulation tests and review before production.
- Video decoding: agents replaced 32K lines of SIMD code in an existing Rust libgav1 port after profile-guided experiments. The resulting decoder ran 2.7 times as fast as that Rust port with identical video output, according to Google's write-up.
Cyber defense
Google says trusted defenders and its own internal teams will get Argon without cyber guardrails so they can use its full defensive capabilities. For wider access, it describes work on misuse refusals, prompt-injection robustness, monitoring of model reasoning and actions, and sealed test environments in its safeguards account.
One launch reaction said government review preceded release. Google's wording is narrower: it says it is actively engaged in the U.S. government's voluntary pre-release access process while expanding access.
Wiz has used Argon through its Scan for Good program and found a critical vulnerability exposing personal information in hospital software that previous frontier models had missed, according to Google's account. The company also reports a 68% score on CWE-bench v1, tied for first on vulnerability remediation in its evaluation table.