Skip to content
AI Primer
TOPIC9 stories

Voice AI

Stories, products, and related signals connected to this tag in Explore.

RELEASE2w ago
xAI says Grok Voice Transcribe 2.0 doubles accuracy at the same price

The API release keeps batch transcription at $0.10 per hour and streaming at $0.20 per hour. xAI says the model improves transcription of support calls, spoken credentials, and short commands.

RELEASE3w ago
Google releases Gemini 3.8 Live with visual understanding

Google introduced Gemini 3.8 Live and Live Extended Thinking in Gemini Live and its API. The models add near-real-time visual understanding, support for 97 languages, and background tool use.

RELEASE1mo ago
Gemini 3.5 Transcribe adds real-time speech-to-text in 85 languages

Google released Gemini 3.5 Transcribe with streaming transcription, speaker identification, custom vocabulary, and function calling. It is available through Gemini and Google's developer platforms.

NEWS2mo ago
Seedance 2.5 demo clones a creator’s voice and likeness from one prompt

Venture Twins gave Dreamina a photo, a talking clip, and office images to make an a16z SF walkthrough in their voice and likeness. Replies said the clip was one continuous generation and noted European IP access limits.

WORKFLOW2mo ago
Codex Voice Mode supports hands-free trackers, email, reminders, and scheduling

Allie Miller described using Codex Voice Mode for trackers, email, reminders, scheduling, and document work. Min Choi’s ElevenLabs setup shows a browser agent that listens, streams replies, and handles interruptions.

RELEASE4mo ago
ElevenLabs claims Speech Engine adds 70-plus voice languages to agents

A sponsored explainer thread described Speech Engine as a WebSocket layer that adds speech-to-text, turn detection, interruption handling, and text-to-speech to existing LLM agents. The pitch is that teams can keep their current model stack and add voice without rebuilding the whole agent.

RELEASE4mo ago
Supertone opens Supertonic with ONNX on-device TTS

Supertone open-sourced Supertonic, a local TTS engine that runs faster than real time on phone CPUs with ONNX models and cross-language runtimes. Voice apps and audiobook workflows can use it to avoid per-character API billing and keep audio generation private.

RELEASE5mo ago
Grok Voice Think Fast 1.0 adds concierge and reservation tool templates

xAI rolled out Grok Voice Think Fast 1.0 with ready-made tool schemas for medical offices, restaurants, help desks, real estate, appointments, and hotel concierge tasks. The release lowers setup work because common service actions arrive pre-wired as callable tools.

RELEASE5mo ago
VoxCPM releases 2B voice model with 3-second cloning and 30-language support

OpenBMB released VoxCPM on GitHub with text-described voice design, 3-second cloning, 48kHz audio, and 30-language support. The Apache 2.0 release makes multilingual voice work and local self-hosting cheaper.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.