SpaceXAI has released Grok Voice Transcribe 2.0 , its newest speech-to-text (STT) model. The development team claims it to be twice as accurate as Grok Voice Transcribe 1.0 at the same price. The model targets hard audio: noisy phone lines, competing voices, local accents, and spoken credentials. It runs in batch and real-time streaming modes through the Speech to Text API . Is it deployable? Yes, as a hosted API. It is live today under the model ID grok-voice-transcribe-2.0 . SpaceXAI has not announced open weights, so self-hosting is not an option. What is Grok Voice Transcribe 2.0? Grok Voice Transcribe 2.0 is built on the audio foundation model behind Grok Voice. SpaceXAI team states that Grok Voice already handles tens of thousands of customer-support calls a day. It also transcribes millions of hours of video narration and runs the Grok assistant in Tesla vehicles. The training data is live, noisy, multilingual audio recorded across diverse environments. SpaceXAI then refined the model with post-training. Benchmarks: What SpaceXAI Reports SpaceXAI reports a first-place accuracy rank among 32 streaming models on the public Artificial Analysis leaderboard . That benchmark, AA-WER Streaming, uses about 8 hours of audio. It weights AA-AgentTalk at 50%, VoxPopuli at 25%, and Earnings22 at 25%. See the methodology for details. SpaceXAI also measures word error rate (WER) on 4 internal sets drawn from production traffic: Telephony (8 kHz): English customer support calls Conversational: English conversations with Grok Credentials: phone numbers, emails, and addresses in Engli
Source: MarkTechPost
