OpenAI's GPT Transcribe Cuts Word Error Rate to 3.31% but Trails ElevenLabs, Google, and Mistral
OpenAI released GPT Transcribe and GPT Live Transcribe, improving word error rate to 3.31 percent and cutting prices 25 percent to $0.0045 per minute. Independent benchmarks still place OpenAI behind ElevenLabs, Google, and Mistral on transcription accuracy.
OpenAI has released two new speech recognition models through its API: GPT Transcribe, for pre-recorded audio, and GPT Live Transcribe, for real-time streaming. Both improve on last year's GPT-4o Transcribe but still fall short of the top scores posted by ElevenLabs, Google, and Mistral on the AA-WER benchmark run by Artificial Analysis.
Performance and pricing
According to Artificial Analysis, GPT Transcribe achieves a word error rate (WER) of 3.31 percent, a 0.7 percentage point improvement over GPT-4o Transcribe's roughly 4 percent rate. OpenAI also cut pricing by 25 percent, bringing the cost to $0.0045 per minute of audio.
GPT Transcribe processes pre-recorded audio files approximately 34 times faster than real time, according to OpenAI. GPT Live Transcribe is designed specifically for low-latency streaming use cases. Both models accept text context, keyword lists, and multiple input languages to improve transcription accuracy.
Where OpenAI ranks
Despite the improvements, GPT Transcribe does not lead the AA-WER benchmark. ElevenLabs Scribe v2 currently tops the ranking with a 2.3 percent error rate. Google's Gemini 3 Pro follows at 2.9 percent, and Mistral's Voxtral Small comes in at 3 percent — all ahead of OpenAI's 3.31 percent.
Mistral has also been aggressive on price. Its recently launched Voxtral Transcribe V2 starts at $0.003 per minute, undercutting OpenAI's new rate by roughly 33 percent even before accounting for OpenAI's accuracy gap.
Context
The new transcription models arrive alongside OpenAI's broader Realtime model lineup, which also includes GPT-Realtime-Whisper, a dedicated real-time transcription model. Full technical details, including supported languages and API parameters, are available in OpenAI's Transcription Guide.
What this means
Speech-to-text has become a crowded, fast-moving segment where accuracy gaps between leading providers are now measured in fractions of a percentage point. OpenAI's improvement — trimming WER by 0.7 points while cutting price 25 percent — is a solid iteration, but it doesn't change the competitive picture: ElevenLabs, Google, and Mistral all currently transcribe more accurately, and Mistral is also cheaper.
For developers choosing a transcription API, the decision now hinges on trade-offs between raw accuracy, latency, price, and existing platform integration rather than any single provider having a clear lead. OpenAI's advantage is likely to come from bundling — pairing transcription with its broader Realtime and voice model ecosystem — rather than from topping accuracy leaderboards. Whether that bundling advantage outweighs a rival's better error rate will depend on the specific use case, particularly for high-volume transcription where cost and accuracy compound quickly.
Related Articles
OpenAI Report Claims Coding Agents Sped Up Eight Scientific Computing Projects
OpenAI has published a field report documenting eight scientific computing projects that used its Codex coding agent — alone or alongside Anthropic's Claude Code — to reduce software build times. The report is a vendor-authored survey, not an independent study.
ChatGPT Now Refuses to Mimic Famous Authors' Writing Styles, Citing Copyright
ChatGPT is refusing prompts to write in the style of specific authors, both living and dead, citing copyright concerns. The change marks a shift from earlier behavior documented by researchers, though the chatbot still offers to write content with similar stylistic qualities.
Microsoft Launches MAI-Cyber-1-Flash Security Model, Still Routes Hard Cases to OpenAI's GPT-5.4
Microsoft has released MAI-Cyber-1-Flash, a compact cybersecurity model built into its MDASH multi-agent system that scores 96 percent on the CyberGym benchmark. The setup handles 90 percent of security tasks in-house but still hands off difficult cases to OpenAI's GPT-5.4.
Anthropic's Claude Opus 4.7 Completes Robot Tasks 20x Faster Than Prior Model, New Benchmark Shows Week-Long Coding Feat
A new Epoch/METR benchmark called MirrorCode shows Claude Opus 4.7 reimplementing large software programs from scratch in tasks estimated to take humans 2-17 weeks, for $251 in inference cost. Separately, Anthropic reports Opus 4.7 completed a suite of quadruped robot tasks in 9 minutes 35 seconds, down from 181 minutes with an earlier model assisting humans.
Comments
Loading...