OpenAI's GPT Transcribe Cuts Word Error Rate to 3.31% but Trails ElevenLabs, Google, and Mistral
OpenAI released GPT Transcribe and GPT Live Transcribe, improving word error rate to 3.31 percent and cutting prices 25 percent to $0.0045 per minute. Independent benchmarks still place OpenAI behind ElevenLabs, Google, and Mistral on transcription accuracy.
OpenAI has released two new speech recognition models through its API: GPT Transcribe, for pre-recorded audio, and GPT Live Transcribe, for real-time streaming. Both improve on last year's GPT-4o Transcribe but still fall short of the top scores posted by ElevenLabs, Google, and Mistral on the AA-WER benchmark run by Artificial Analysis.
Performance and pricing
According to Artificial Analysis, GPT Transcribe achieves a word error rate (WER) of 3.31 percent, a 0.7 percentage point improvement over GPT-4o Transcribe's roughly 4 percent rate. OpenAI also cut pricing by 25 percent, bringing the cost to $0.0045 per minute of audio.
GPT Transcribe processes pre-recorded audio files approximately 34 times faster than real time, according to OpenAI. GPT Live Transcribe is designed specifically for low-latency streaming use cases. Both models accept text context, keyword lists, and multiple input languages to improve transcription accuracy.
Where OpenAI ranks
Despite the improvements, GPT Transcribe does not lead the AA-WER benchmark. ElevenLabs Scribe v2 currently tops the ranking with a 2.3 percent error rate. Google's Gemini 3 Pro follows at 2.9 percent, and Mistral's Voxtral Small comes in at 3 percent — all ahead of OpenAI's 3.31 percent.
Mistral has also been aggressive on price. Its recently launched Voxtral Transcribe V2 starts at $0.003 per minute, undercutting OpenAI's new rate by roughly 33 percent even before accounting for OpenAI's accuracy gap.
Context
The new transcription models arrive alongside OpenAI's broader Realtime model lineup, which also includes GPT-Realtime-Whisper, a dedicated real-time transcription model. Full technical details, including supported languages and API parameters, are available in OpenAI's Transcription Guide.
What this means
Speech-to-text has become a crowded, fast-moving segment where accuracy gaps between leading providers are now measured in fractions of a percentage point. OpenAI's improvement — trimming WER by 0.7 points while cutting price 25 percent — is a solid iteration, but it doesn't change the competitive picture: ElevenLabs, Google, and Mistral all currently transcribe more accurately, and Mistral is also cheaper.
For developers choosing a transcription API, the decision now hinges on trade-offs between raw accuracy, latency, price, and existing platform integration rather than any single provider having a clear lead. OpenAI's advantage is likely to come from bundling — pairing transcription with its broader Realtime and voice model ecosystem — rather than from topping accuracy leaderboards. Whether that bundling advantage outweighs a rival's better error rate will depend on the specific use case, particularly for high-volume transcription where cost and accuracy compound quickly.
Related Articles
OpenAI's GPT-6 Astra Beats Claude Fable 5.1 Nearly 3-to-1 in Autonomous Business Benchmark, Tops Drone Navigation Tests
Independent testing lab Andon Labs found OpenAI's GPT-6 Astra nearly triples Claude Fable 5.1's performance running a simulated vending machine business, averaging $15,515 versus $5,422. Astra also became the first model to beat human-AI baseline performance across all five Drone-Bench subtasks, including autonomous person-tracking via drone.
GPT-6 Astra Beats Ai2's MolmoAct2 on New Robotics Benchmark, Researcher Calls It a 'Step Change'
A new robotics benchmark called StationeryBench shows OpenAI's GPT-6 Astra completing 7 of 100 desk-object manipulation tasks versus zero for Ai2's MolmoAct2, with a median progress score of 46 against 12. Cornell/DeepMind researcher Yoav Artzi calls the result a 'step change in spatial reasoning.'
AWS Benchmark: OpenAI's GPT-5.6 Luna Beats GPT-5.4 Mini on Cost-Per-Correct-Answer Despite Similar List Price
An AWS blog post using an open-source benchmarking harness finds that GPT-5.6 Luna, Terra, and Sol on Amazon Bedrock deliver lower cost-per-correct-answer than OpenAI's cost-optimized GPT-5.4 Mini and Nano, once accuracy, token efficiency, and agent turn counts are factored in. The analysis also cites a July 30, 2026 price cut of up to 80% for GPT-5.6 Luna on Amazon Bedrock.
Google Releases TimesFM-3, a 330M-Parameter Model That Forecasts Sales Using Weather and Discount Data
Google Research has released TimesFM-3, a 330-million-parameter time series forecasting model that predicts outcomes like sales by combining related variables, historical data, and known future events such as discounts or weather. The model claims top rankings on three benchmarks against Amazon's Chronos-2 and the Toto-2.0 family.
Comments
Loading...