Gemini 3.1 Flash TTS Preview
Low-latency, controllable Gemini text-to-speech with expressive audio tags.
A preview Gemini text-to-speech model release that takes text input and generates audio output, optimized for low-latency, controllable single- and multi-speaker speech generation with natural delivery, multilingual support, steerable prompts, and expressive audio tags.
Pricing
Model profile · Current snapshot
Input / 1M
$0.50
Output / 1M
$3.00
Blended / 1M
$1.13
Output TPS
185
TTFT (s)
5.52
Model Intelligence
Context window
8,192 tokens
Arena ranking
46
Benchmarkable
Yes
Model level
release
Intelligence Index
46.4
Coding Index
42.6
Math Index
97
MMLU Pro
0.89
GPQA
0.9
HLE
0.35
LiveCodeBench
0.91
SciCode
0.51
AIME 2025
0.97
IFBench
0.78
LCR
0.66
TerminalBench Hard
0.39
TAU2
0.8
Recent stories
1 linked story