Skip to content
AI Primer

Gemini 3.1 Flash TTS Preview

Controllable text-to-speech with Audio Tags

Preview release of Google's Gemini 3.1 Flash text-to-speech model for controllable speech generation, including inline Audio Tags, multilingual output, and SynthID watermarking.

Pricing

Model profile · Current snapshot
Input / 1M
$0.50
Output / 1M
$3.00
Blended / 1M
$1.13
Output TPS
185
TTFT (s)
5.52

Model Intelligence

Arena ranking
46
Benchmarkable
Yes
Model level
release
Intelligence Index
46.4
Coding Index
42.6
Math Index
97
MMLU Pro
0.89
GPQA
0.9
HLE
0.35
LiveCodeBench
0.91
SciCode
0.51
AIME 2025
0.97
IFBench
0.78
LCR
0.66
TerminalBench Hard
0.39
TAU2
0.8

Recent stories

1 linked story
AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.