Skip to content
AI Primer
release

Google releases Gemini 3.8 Live with visual understanding

Google introduced Gemini 3.8 Live and Live Extended Thinking in Gemini Live and its API. The models add near-real-time visual understanding, support for 97 languages, and background tool use.

3 min read
Google releases Gemini 3.8 Live with visual understanding
Google releases Gemini 3.8 Live with visual understanding

TL;DR

  • Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are rolling out in Gemini Live and the developer API, according to GoogleDeepMind's announcement.
  • The base model combines near-real-time visual understanding, automatic detection across 97 languages, and tool calls that run without stopping the conversation, as GoogleDeepMind's feature list describes.
  • Extended Thinking adds deeper background reasoning for multi-step work and can narrate its progress while it keeps speaking, according to GoogleDeepMind's launch post.
  • Google lists audio input at $0.005 per minute and output at $0.018 per minute in its developer release; OfficialLoganK's announcement calls the pair a frontier price-and-performance offering.

A developer post calls out alphanumeric precision, aimed at spoken confirmation codes, claim numbers, and technical data. The model card identifies Gemini 3 Pro as the foundation and gives the live models a 128K-token input context.

Continuous conversation and background tools

Google's developer announcement positions Gemini 3.8 Live as the faster, scale-oriented model for real-time dialogue. Its Extended Thinking sibling targets complex, multi-step requests.

The API model page says Extended Thinking processes configurable background reasoning and asynchronous tool calls while it continues streaming audio. Google describes progress narration as part of that flow, using early acknowledgements while longer work continues.

Visual context and language switching

Google says the Live model can ground a conversation in near-real-time visual input and switch languages during the same conversation across 97 supported languages in its launch post.

The developer release adds incremental content updates, which merge live audio with structured data, plus alphanumeric parsing for details that frequently get lost in speech interfaces.

The voice front door sketch

gregisenberg imagines a contractor reporting a problem aloud on site while an agent prepares a quote, checks inventory, updates the CRM, texts the customer, and surfaces risky items. The same post names nurses, dispatchers, recruiters, brokers, and insurance agents as other settings for a voice-led administrative workflow.

Price and scorecard

Google's developer post puts both models at $0.005 per minute of input audio and $0.018 per minute of output audio, an estimate based on $3 per million input tokens and $12 per million output tokens.

The vendor's launch page reports these Extended Thinking results:

  • 82.6 on Artificial Analysis' Speech-to-Speech Quality Index.
  • 68.6% on τ-Voice agentic task completion.
  • 35.1% on Sierra's τ-Voice-banking benchmark.
  • 97.7% on Big Bench Audio.

Where it is rolling out

Google's rollout list puts both models in the Gemini API and Google AI Studio. Gemini 3.8 Live is also arriving in Search Live, while Extended Thinking is available in Gemini Live and selected Workspace surfaces.

Extended Thinking is in private preview for Gemini Enterprise, with broader Enterprise, Customer Experience, and Workspace business availability described as coming soon. Google AI Pro and Ultra subscribers get it in Docs, while Gmail and Keep availability covers Google AI subscribers.

Model card limits

The model card specifies audio, image, video, and text inputs, with up to 128K input tokens and 64K output tokens. It also lists hallucinations, occasional slowness or timeouts, and a January 2025 knowledge cutoff among the known limitations.

Google says all AI-generated audio from its products carries SynthID watermarking in its launch post.

Share on X