Google releases Gemini 3.8 Flash at Gemini 3.7 Flash pricing
Google released Gemini 3.8 Flash across its developer and consumer products. Google says it matches Gemini 3.7 Flash pricing and scored 73.7% on DeepSWE 1.1.

TL;DR
- Gemini 3.8 Flash is rolling into Google’s developer, enterprise, and consumer surfaces, as GoogleDeepMind's rollout lists Antigravity, the API, AI Studio, Android Studio, the Gemini app, and AI Mode.
- List pricing stays at $0.75 per million input tokens and $3.75 per million output tokens through December 31, while OfficialLoganK's post also claims roughly the same speed as 3.7 Flash.
- Google’s DeepSWE chart puts long-horizon software engineering at 65.3% → 73.7%, +8.4 points, in OfficialLoganK's DeepSWE chart.
- Higher-effort runs can consume more tokens: OfficialLoganK's reply calls the extra intelligence a deliberate design choice, and Google’s launch post says complex jobs take extra reasoning steps and iterative tool calls.
- A creator-run Video QC test puts 3.8 Flash at the top of its detection-versus-false-positive chart, according to kaigani's Video QC post.
Google’s launch post includes a one-prompt, playable DOS Google Maps with locations, directions, and Street View. It also describes a USGS-data topographic-map build that produces real-time cross-sections and 2D projections.
What shipped
- Gemini 3.8 Flash is Google’s general-purpose model, and 3.8 Flash Cyber is its cybersecurity variant, with both sharing foundational intelligence in Google’s launch post.
- Developer access covers the Gemini API, AI Studio, Android Studio, Antigravity, and Stitch, while Gemini Enterprise gets the standard Flash model, according to GoogleDeepMind's rollout.
- Google AI Pro and Ultra subscribers get Flash in the Gemini app, AI Mode in Search, and Gemini in Google Sheets, per Google’s launch post.
- Introductory API pricing is $0.75 per million input tokens and $3.75 per million output tokens, then rises to $1.50 and $7.50 on January 1, 2027, in the launch post.
Benchmarks that moved
First-party
- Terminal-Bench 4.0: 11.2% → 19.1%, +7.9 points, in OfficialLoganK's benchmark table.
- OSWorld-2.0: 50.6% → 59.0%, +8.4 points, in OfficialLoganK's benchmark table.
- BioMysteryBench, Human Difficult: 43.5% → 56.5%, +13.0 points, in OfficialLoganK's benchmark table.
- Vals Finance Agent v2: 59.0% → 61.4%, +2.4 points, in OfficialLoganK's benchmark table.
- HLE-Verified: 53.6% → 54.9%, +1.3 points, in OfficialLoganK's benchmark table.
Third-party evaluators
- Datacurve’s DeepSWE v1.1: 65.3% → 73.7%, +8.4 points, in OfficialLoganK's chart that credits Datacurve.
Customer-reported
Google’s launch post publishes partner results in relative terms, rather than raw 3.7-to-3.8 scores, so it provides no customer-specific point delta.
Where it regressed
Google’s model card reports a +5.4-point change in Multilingual Safety against 3.7 Flash, where lower is better. Unjustified refusals also rose by 1.1 points, another lower-is-better metric, in the same safety table.
Google says manual review found the apparent losses were overwhelmingly false positives or non-egregious content, while its human red-team results were similar or better than 3.7 Flash in the model card. Google’s product account also positioned the model as strong rather than the world’s best in OfficialLoganK's reply.
Under the hood
The model card identifies 3.7 Flash as 3.8 Flash’s model dependency and sends readers to 3.7 documentation for its architecture, data, hardware, and software.
Google describes the higher-effort approach as extracting more intelligence from Flash in OfficialLoganK's reply.
- Effort: On complex work, Google says 3.8 takes extra reasoning steps and makes iterative tool calls. It may use more tokens at higher effort, while lower effort levels remain available and 3.7 Flash remains supported for efficiency-first work in the launch post.
- Input and output: The model accepts text, images, audio, and video in up to a 1 million-token context, then returns text up to 64K tokens, according to the model card.
- Training claim: Google attributes gains across the shared Flash and Flash Cyber core to long-running agentic loops and training innovations in cybersecurity in its launch post.
Vibe Check
Ozan Sihay, who said he had early access, reported shorter waits before answers, less need to spell out instructions, and fewer small misses when reading image and text together in ozansihay's early-access notes.
In a separate ozansihay's video post, Sihay said a 480p, 15-second video completed in five seconds without speed-up.
The reported 480p, 15-second generation
The four additions in AIandDesign's gallery post put Gemini 3.8 Flash designs into a one-page portfolio project, with the creator singling out “The Collider.”
kaigani’s Video QC Bench runs 14 videos through six role prompts, for 84 API calls per system, and subtracts false-positive penalties from detection credit. Their Video QC post calls 3.8 Flash the best value and says it tops that detection-versus-false-positive chart.
Where it shows up
LiteLLM’s day-zero support routes gemini-3.8-flash through both Google AI Studio and Vertex AI, with recent LiteLLM configurations requiring no Docker-image update.
Google’s launch post also uses AI Studio for Hardware Anatomy, a Three.js device-teardown visualizer with an explode-and-inspect slider.
Flash Cyber has a separate, limited path through Fairwind for trusted government authorities, critical-infrastructure operators, and software maintainers. Google says its Chrome Security team saw 2.6 times more correct vulnerability patches than the best larger commercial models, in GoogleDeepMind's Fairwind post.