Skip to content
AI Primer
release

xAI says Grok Voice Transcribe 2.0 doubles accuracy at the same price

The API release keeps batch transcription at $0.10 per hour and streaming at $0.20 per hour. xAI says the model improves transcription of support calls, spoken credentials, and short commands.

4 min read
xAI says Grok Voice Transcribe 2.0 doubles accuracy at the same price
xAI says Grok Voice Transcribe 2.0 doubles accuracy at the same price

TL;DR

  • Grok Voice Transcribe 2.0 keeps batch transcription at $0.10 per audio hour and streaming at $0.20, while minchoi's post relays xAI's claim that accuracy doubled versus 1.0.
  • xAI's SpaceXAI's accuracy post flags customer calls, credentials, and short commands as the target cases; the official launch post puts short-phrase word error rate at 20.6% in 1.0 and 6.8% in 2.0, a -13.8 percentage point change.
  • Atlassian put the model into Loom, where SpaceXAI's Loom post describes recorded change requests flowing from a transcript into Cursor.
  • Version 2.0 is in the Grok Voice API now, while SpaceXAI's availability post says 1.0 will be deprecated after 2.0 becomes the default.

The feature list includes up to eight separately transcribed channels, 100 domain-specific key terms, and language switching within a recording. Atlassian's Loom example offers a more concrete loop: capture the explanation once, then turn its transcript into a code change.

What shipped

  • Accuracy claim: Transcribe 1.0 to 2.0, +100% claimed accuracy, according to SpaceXAI's accuracy post.
  • Batch price: $0.10 per audio hour in 1.0 to $0.10 in 2.0, $0.00 per-hour change, per minchoi's price post.
  • Streaming price: $0.20 per audio hour in 1.0 to $0.20 in 2.0, $0.00 per-hour change, as SpaceXAI's availability post states.
  • Migration: 1.0 is the current default, while the official announcement says 2.0 will become the default in coming weeks; callers can pin grok-voice-transcribe-1.0 during the transition.
  • Included features: diarization, word timestamps, and key-term biasing remain included in the published per-hour price, according to the launch post.

Benchmarks that moved

First-party

  • Telephony WER, xAI's 8 kHz English support-call set: 10.6% in 1.0 to 7.1% in 2.0, -3.5 percentage points, as reported in OrcaRouter's transcription of xAI's chart.
  • Conversational WER, xAI's English Grok-conversation set: 8.7% to 3.3%, -5.4 percentage points, in the same reported chart.
  • Spoken-credential WER, xAI's English account-code, email, and phone-number set: 7.2% to 3.2%, -4.0 percentage points, according to OrcaRouter's comparison.
  • Short-phrase WER, xAI's 19-language voice-command set: 20.6% to 6.8%, -13.8 percentage points, in the official launch post.

Third-party evaluators

Customer-reported

  • Loom transcription: an unnamed existing solution to Transcribe 2.0, reported as more accurate but with no score delta disclosed, in SpaceXAI's Loom post.

Where it regressed

  • Time to final transcript rose from 0.37 seconds in 1.0 to 0.49 seconds in 2.0, +0.12 seconds, according to OrcaRouter's side-by-side.
  • Time to first partial rose from 0.25 seconds to 0.49 seconds, +0.24 seconds, in the same comparison.
  • xAI calls 2.0 first among 32 streaming models, while OrcaRouter's audit says the leaderboard page displayed 27 of 33 models. The company count and the board display are not the same field size.
  • The public streaming score uses AA-AgentTalk for 50% of its mix and VoxPopuli and Earnings22 for 25% each, per the Artificial Analysis methodology. xAI's four version-to-version sets are drawn from its own production traffic.

Under the hood

The official feature list keeps batch files and URLs alongside real-time streaming in the same speech-to-text API.

  • Word-level start and end timestamps include confidence scores.
  • Speaker diarization labels speakers without an added charge.
  • Multichannel transcription handles up to eight channels independently.
  • Key-term biasing accepts up to 100 domain terms per request.
  • Text normalization formats numbers, dates, currencies, emails, and phone numbers.
  • Automatic language detection follows mid-recording language switches.
  • Filler-word removal and smart turn detection target voice-agent handoffs.

Loom to Cursor

Atlassian says it now uses Transcribe 2.0 to transcribe every Loom video, after finding it more accurate than its prior solution in the launch account. Sanchan Saxena, SVP of Teamwork Collection at Atlassian, described the intended chain as captured context flowing through the tools rather than being re-explained.

  1. Record an action plan or change request in Loom.
  2. Generate the transcript with Grok Voice Transcribe 2.0.
  3. Send the transcript to Cursor, which Atlassian says can make the code updates directly.

Grok bot voice mode

The API release lands alongside voice interfaces in the Grok product itself. jenny_wen condensed the shift to “grok bot → talk bot,” while SpaceXAI's fal post pitches Grok Voice for low-latency agents that resolve customer issues.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 3 threads
TL;DR1 post
What shipped1 post
Grok bot voice mode2 posts
Share on X