Skip to content
AI Primer
TOPIC50 stories

Voice

Stories, products, and related signals connected to this tag in Explore.

WORKFLOW24th September
Creator demos show Claude Opus 5.5 generating voiced animated videos from prompts

Creator demos show Claude Opus 5.5 producing short narrated explainers, opening credits, and marketing videos from detailed prompts. One report says a 30-second voiced animation ran inside Claude's VM from a single prompt.

WORKFLOW22nd September
InVideo Editor agents add synchronized sound effects to a trailer

Creators describe using InVideo Editor with GPT-6 Astra to turn raw footage into edits while retaining control over story and style. In one test, the editor added atmosphere, impacts, transitions, and synchronized effects to the edit.

RELEASE1w ago
xAI says Grok Voice Transcribe 2.0 doubles accuracy at the same price

The API release keeps batch transcription at $0.10 per hour and streaming at $0.20 per hour. xAI says the model improves transcription of support calls, spoken credentials, and short commands.

WORKFLOW2w ago
Meng To says he built a multiplayer Catan-inspired game with GPT-6 Astra in four days

Meng To says he built a free, mobile-friendly multiplayer Catan-inspired game with GPT-6 Astra in four days. He reports spending hundreds of dollars in tokens on narration, guides, expansions, and voice-chat lobbies.

RELEASE4w ago
Gemini 3.5 Transcribe adds real-time speech-to-text in 85 languages

Google released Gemini 3.5 Transcribe with streaming transcription, speaker identification, custom vocabulary, and function calling. It is available through Gemini and Google's developer platforms.

RELEASE1mo ago
Pika Music launches 4-input diffusion model via API Club

Pika Music accepts text, lyrics, voice, and music references separately or together in one diffusion decoder. It is available through Pika API Club, and Pika claims up to 10 times better cost efficiency than other music models.

WORKFLOW1mo ago
Codex Voice Mode supports hands-free trackers, email, reminders, and scheduling

Allie Miller described using Codex Voice Mode for trackers, email, reminders, scheduling, and document work. Min Choi’s ElevenLabs setup shows a browser agent that listens, streams replies, and handles interruptions.

RELEASE2mo ago
OpenAI rolls out ChatGPT Voice with GPT-Live-1

OpenAI staff said the new ChatGPT Voice is powered by GPT-Live-1 for more natural conversations. The launch also refreshed the ChatGPT shader with Blender prototypes translated to Metal and WebGL using Codex.

RELEASE2mo ago
Seed Audio 1.0 generates 120-second scene audio from one prompt

A BytePlus thread says Seed Audio 1.0 can create dialogue, effects, background music, and reference voices for clips up to 120 seconds. The workflow can pass audio timing into Seedance for video generation.

NEWS2mo ago
Pictory reports 1.53M AI videos in 2026 State of Video Report

Pictory's 2026 State of Video Report analyzed 1.53M videos, including 9–10pm creation peaks, Denmark's voiceover rate, and UAE per-capita adoption. Use the benchmark to compare team workflows: professional teams used URL-to-video 9–10x more than personal users.

RELEASE2mo ago
Higgsfield launches Seed Audio 1.0 with 18-language dubbing and Claude MCP

Higgsfield launched Seed Audio 1.0 with voice replacement, text narration, 18-language dubbing and Claude access through Higgsfield MCP. Early tests and BeatBandit integrations show it being used to audition performances before sending audio-guided shots into Seedance.

RELEASE3mo ago
Wan-Streamer v0.1 opens real-time video agents with live voice demos

Wan-Streamer v0.1 surfaced with paper links and demos showing real-time video conversations, live recording, and spoken avatar responses. That matters for interactive characters and live creator tools because multimodal generation moves from rendered clips to low-latency back-and-forth.

RELEASE3mo ago
OpenArt Director launches 5-minute video generation with characters, voiceovers, music, and captions

OpenArt rolled out Director as a conversational video mode that can generate projects up to 5 minutes with recurring characters, props, voiceovers, music, captions, and templates. The release pushes the product toward full sequence assembly inside one tool instead of short demo clips.

RELEASE3mo ago
STAGES AI releases Resolve Audio Studio v1.01 for paid users

STAGES AI said Resolve Audio Studio v1.01 shipped for paid users after earlier preview posts, with multi-project sound editing as the main change. Posts also kept v1.02 stem separation on the roadmap, moving the audio module from preview into live use.

NEWS3mo ago
Stages AI previews Audio Studio v1.01 with stem-separation in v1.02 plans

Stages AI previewed Audio Studio and said v1.01 is under checks while v1.02 is slated to add stem separation, alongside voice-design and mobile demos. The rollout points to an in-platform audio workflow for editing, narration, and music tasks that usually live in separate tools.

WORKFLOW3mo ago
Pika supports six-language Language Swap demos through MCP

Pika expanded its Language Swap story with official multi-language demos and creator tests of own-voice Japanese dubbing through MCP. The evidence matters because localization is shifting from voice replacement toward performance-preserving video translation inside creator workflows.

WORKFLOW3mo ago
Pika adds Language Swap to MCP for own-voice video dubbing

Pika launched a Language Swap skill in Pika MCP that re-dubs talking-head videos into other languages while keeping the speaker’s voice profile and lip sync. Creator demos in Japanese and Mandarin show it is already usable, though facial hair and loose framing can still distort mouth motion.

NEWS4mo ago
MiniMax Speech 2.8 powers Storyverse crime-drama voices at Cannes

Hailuo said Storyverse used MiniMax Speech 2.8 on the Italian series Il Cinese to recreate regional accents and character-specific vocal traits at Cannes. The showcase moves AI voice closer to film-market pitching and entertainment trade coverage rather than standalone TTS demos.

NEWS4mo ago
GLITCH introduces speech-to-film workflow in an 8-minute Michael Shamberg demo

Starks ARQ introduced GLITCH as a speech-to-film system and used it to generate a short with producer Michael Shamberg in under eight minutes. The demo points to AI filmmaking directed out loud, but access appears limited to direct project inquiries.

RELEASE4mo ago
ElevenLabs claims Speech Engine adds 70-plus voice languages to agents

A sponsored explainer thread described Speech Engine as a WebSocket layer that adds speech-to-text, turn detection, interruption handling, and text-to-speech to existing LLM agents. The pitch is that teams can keep their current model stack and add voice without rebuilding the whole agent.

RELEASE4mo ago
OmniVoice Studio opens local dubbing for 600 languages from one MP4

A community post spotlights OmniVoice Studio, an open-source local dubbing pipeline that transcribes, translates, clones voice from 3 seconds and remixes dubbed audio back into video. Running locally keeps voice data on device and removes subscription costs, so it may fit privacy-sensitive dubbing workflows.

RELEASE4mo ago
Harbor releases v0.4.18 with Open Design and Voicebox

Harbor 0.4.18 added one-command access to Open Design and Voicebox, bundling a local-first design app and a voice cloning and TTS studio inside one homelab layer. The release cuts setup friction, so users can migrate both tools into a single local install path.

RELEASE4mo ago
Grok Voice Think Fast 1.0 adds concierge and reservation tool templates

xAI rolled out Grok Voice Think Fast 1.0 with ready-made tool schemas for medical offices, restaurants, help desks, real estate, appointments, and hotel concierge tasks. The release lowers setup work because common service actions arrive pre-wired as callable tools.

WORKFLOW4mo ago
Seedance 2.0 supports multi-speaker lip sync in live-action and animation

Curious Refuge posted tests showing Seedance 2.0 syncing multiple speakers from a reference image plus blacked-out video or audio, using shot-by-shot dialogue prompts. The workflow moves Seedance closer to directed dialogue scenes, but prompt wording and voice guidance still affect stability.

RELEASE4mo ago
Runway launches Characters with 24fps HD video agents

Runway launched Characters, a real-time system that turns one image into a conversational HD video agent. The company says replies start in 1.75 seconds and stream above 24 fps, so live avatar workflows are moving closer to production use.

RELEASE4mo ago
Apocalypse Drone adds 128 AI players and ElevenLabs radio voices

Apocalypse Drone added 128 AI players, squad leader reassignment, and ElevenLabs radio chatter with location callouts in weekend dev updates. It matters for solo game builders because the project is simulating large-team coordination and voice comms on a lightweight stack instead of a bigger live-ops setup.

RELEASE4mo ago
Pika launches Claude MCP with podcast, explainer and UGC ad skills

Pika launched a Claude connector that turns prompts, URLs and repos into explainers, podcast clips and UGC ads. The update keeps face, voice and identity controls inside Claude workflows, so creators can build video assets without switching apps.

RELEASE5mo ago
OpenClaw adds voice personas with 43ms first output benchmarks

OpenClaw contributors posted a voice-persona feature and fresh performance numbers that cut first output from 1s to 43ms. Separate posts describe 300-user sandboxed deployments and stronger PR, CI, and testing workflows, pointing to team-scale use beyond hobby demos.

RELEASE5mo ago
Grok Imagine adds lip sync and multi-speaker audio to video clips

Creator posts say Grok Imagine's video update can make one-shot clips with spoken audio, stronger lip sync and support for multiple speakers, pets and varied face angles. The demos also show selfie-to-scene transforms and timeline prompting, but the rollout is documented mainly through independent testing.

RELEASE5mo ago
Cappy launches video editing in iMessage and RCS

Cappy launched as a text-message video editor that plans, cuts, captions, voices, and revises clips inside iMessage or RCS threads. Creators can start from raw footage, photos, audio, or URLs without opening a conventional timeline.

RELEASE5mo ago
BytePlus launches Seedance 2.0 API with multimodal inputs and scene extension

BytePlus launched the Seedance 2.0 API, and creator tests showed image, video, audio, and text inputs, scene extension, voice-synced delivery, and steadier physics. The move brings Seedance from app-only access into repeatable production pipelines and custom workflows.

RELEASE5mo ago
Gemini 3.1 Flash TTS adds Audio Tags, 70-language support, and SynthID

Gemini 3.1 Flash TTS added Audio Tags, 70-plus language support, and SynthID watermarking for generated speech. The preview spans Gemini API, AI Studio, Vertex AI, and Google Vids, so teams can test delivery control before adopting it.

RELEASE5mo ago
Runway adds Character video-call links for Zoom, Meet, and Teams

Runway now lets a Character join video meetings from a pasted Zoom, Google Meet, or Teams link. The feature extends Runway Characters from rendered clips into live meeting stand-ins, so watch the launch demo and early reactions for reliability.

RELEASE5mo ago
HeyGen launches CLI for one-command avatar video and batch translation

HeyGen released a Mac and Linux CLI that creates avatar videos, lip-synced translations, voice matches, and photo avatars from terminal commands. The binary returns structured JSON and wait flags, which makes video generation scriptable for localization and agent workflows.

RELEASE5mo ago
Runway Characters adds custom voices from text prompts with API access

Runway added prompt-generated custom voices for Characters in the web app and API. Creators can now define tone and persona from text instead of recording or cloning a source voice first, which should speed up voice setup.

WORKFLOW5mo ago
Suno users report v5.5 misses duet tags and instrument cues despite stronger vocals

Reddit posts said v5.5 improved voice tone but still ignores gender-labeled sections, switches singers mid-part, and struggles with detailed instrument instructions. Creators are iterating on renders until the emotion fits, then generating lipsync video to work around the gaps.

RELEASE5mo ago
Pika launches PikaStream 1.0 video chat skill for Google Meet and any agent

Pika released a beta skill that lets Pika AI Selfs and third-party agents join Google Meet with real-time face and voice, and published the integration on GitHub. Pika says memory and personality persist across calls, while beta notes and user posts report glitches as the feature expands beyond Pika’s own agents.

RELEASE6mo ago
Cohere opens Transcribe 2B weights with a browser demo

Browser demo posts and a Hugging Face release surfaced Cohere Transcribe 2B as part of a wider open-audio week that also featured Voxtral 4B TTS. The model gives creators a multilingual ASR option that can live closer to local or browser workflows.

NEWS6mo ago
KittenTTS supports 25MB ONNX voice models as HN debates prosody

Hacker News discussion around KittenTTS has shifted to edge deployment, streaming latency, expressive control, and prosody rather than new model changes. The 25MB ONNX footprint keeps it attractive for CPU and on-device use, but voice quality is still the production boundary.

RELEASE6mo ago
Gemini 3.1 Flash Live launches with 90.8% ComplexFuncBench audio score

Google says its new realtime voice model improves noisy-environment understanding, long conversations and function calling, and it's rolling into Gemini Live, Search Live and AI Studio. Voice creators can test it for lower-latency spoken interactions.

RELEASE6mo ago
KittenTTS releases 25MB nano model for CPU text-to-speech

KittenTTS now offers nano, micro and mini text-to-speech models, with the smallest int8 build under 25MB and built for ONNX CPU inference. Creators can run local voice tools without a cloud round trip.

RELEASE6mo ago
Lightning V3.1 releases 10-second voice cloning with 44.1kHz output and sub-100ms latency

Smallest says Lightning V3.1 can clone a voice from about 10 seconds of audio with 44.1kHz output, sub-100ms latency and 50-plus languages on Waves. Test it for multilingual narration and dubbing, but get explicit permission before cloning any voice.

RELEASE6mo ago
KittenTTS releases 25MB nano voice model with CPU-only ONNX runtime

KittenTTS 0.8 ships new 15M, 40M and 80M models, including an int8 nano model around 25MB that runs on CPU without GPU. It is a fit for narration, character voices and lightweight assistants that need offline or edge-friendly speech.

RELEASE6mo ago
KittenTTS releases v0.8 with a 25MB int8 model and CPU-only speech synthesis

KittenML's latest open-source TTS release spans 15M to 80M models, with the smallest coming in under 25MB and the larger one reportedly running faster than realtime on CPU. Audio creators should test pronunciation and install overhead before betting on it for edge or local voice tools.

NEWS6mo ago
Variety reports As Deep as the Grave used generative AI for Val Kilmer's performance

Variety reports that As Deep as the Grave used generative AI to create Val Kilmer's performance, with material supplied by his family and their backing for the release. For filmmakers, it is an early consent-based case study in digital resurrection where rights and audience expectations matter.

RELEASE6mo ago
Fun-CineForge opens multi-speaker dubbing with temporal modality and a dataset pipeline

Tongyi Lab opened Fun-CineForge with multi-speaker dubbing, temporal modality for off-screen or blocked faces, and a full dataset-building pipeline. It matters for dialogue and localization workflows that break on hard cuts, overlapping speech, or missing lip cues.

RELEASE6mo ago
Grok launches Text-to-Speech API with expressive controls and LiveKit support

xAI released Grok's Text-to-Speech API with natural voices, expressive controls, and LiveKit support; creators are also using Grok Imagine in reference-image and cartoon animation workflows. Try it if you want Grok in a broader voice-and-motion stack instead of chat alone.

WORKFLOW6mo ago
Seedance 2.0 supports wildlife-documentary narration and character SFX, creators report

Creators report Seedance 2.0 is being used for wildlife-documentary scenes with built-in narration prompts and character clips with sound effects. Test it if you want a faster path from prompt to finished short without a separate voice pass.

RELEASE6mo ago
Freepik launches Speak: lip-synced videos in 30+ languages, up to 5 minutes

Freepik launched Speak, which turns an image plus text or audio into a lip-synced talking video with 30+ languages and a 5-minute cap. Use it for UGC ads, localized product demos, and fast talking-head tests without reshoots.

RELEASE6mo ago
Runway launches Characters API: real-time avatars with custom voices and knowledge banks

Runway opened Characters on its developer platform with API access, custom voices, embedded knowledge, and a free starter allowance. Use it to build interactive hosts, guides, and assistants that can talk through tasks instead of relying on passive video.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.