ElevenLabs and Google lead Artificial Analysis speech-to-text benchmark
Artificial Analysis has released an updated speech-to-text benchmark showing ElevenLabs and Google as top performers. The benchmark provides comparative analysis of current speech recognition systems across multiple models.
ElevenLabs and Google Lead Updated Speech-to-Text Benchmark
Artificial Analysis has released an updated speech-to-text benchmark, with ElevenLabs and Google emerging as the dominant performers in speech recognition accuracy and reliability.
The benchmark evaluation tests current speech-to-text systems across multiple dimensions, comparing how well different models convert spoken audio into written text. Both ElevenLabs and Google demonstrate competitive performance at the top tier of the rankings.
Benchmark Details
The updated benchmark from Artificial Analysis provides comparative metrics across speech recognition providers. However, specific accuracy scores, model names tested, dataset composition, and detailed performance differentials between ElevenLabs and Google have not been disclosed in available sources.
ElevenLabs, primarily known for text-to-speech synthesis, has expanded into speech recognition capabilities. Google's speech-to-text service has been a standard offering through Google Cloud and native Android/web integration for years.
What This Means
The benchmark reinforces that speech recognition quality remains concentrated among well-resourced companies with large training datasets. ElevenLabs' competitive positioning in both directions of audio-text conversion suggests the company is building comprehensive speech processing capabilities. Google's continued dominance reflects its extensive audio data access through YouTube, Google Assistant, and cloud service users.
For developers and enterprises selecting speech-to-text providers, this benchmark offers independent evaluation data to guide integration decisions. The full benchmark details would be critical for understanding which system performs better in specific use cases (noise conditions, language support, latency requirements, accuracy thresholds).
Access the full Artificial Analysis benchmark for detailed scoring metrics and comparative analysis across all tested providers.
Related Articles
ServiceNow Releases First Code-Switching ASR Benchmark: ElevenLabs Scribe V2 Leads with Lowest WER Across Four Language
ServiceNow released AU-Harness, the first comprehensive benchmark for code-switched speech recognition in enterprise voice agents, testing seven ASR systems including ElevenLabs, Gemini, and AssemblyAI. The benchmark covers 918 utterances across Spanish-English, French-English, Canadian French-English, and German-English, measuring Word Error Rate (WER), Semantic WER (SWER), and Answer Error Rate (AER). ElevenLabs Scribe V2 achieved the lowest WER across all language pairs, followed closely by AssemblyAI Universal-3 Pro.
Gemini 3.5 Flash ranks 6th in Android coding benchmark at 3x cost of Gemini 3.1 Pro
Google's latest Android Bench results show Gemini 3.5 Flash ranking 6th with a 63.7% success rate, despite averaging $147.10 per benchmark run compared to Gemini 3.1 Pro Preview's $47.90. The newer model used 355.9 tokens per run versus 73.3 for its predecessor, while GPT 5.5 leads the benchmark at 74% success rate.
Composio Benchmark: Claude Code Fastest Agent Framework, But Costs Nearly 3x More Than OpenCode
Composio benchmarked DeepSeek V4 Flash across four agent frameworks—Claude Code, Codex, OpenCode, and Oh My Pi—on 30 real-world tasks. Claude Code finished fastest at 122 seconds per task but cost $0.195, nearly three times OpenCode's $0.073, while Oh My Pi had the highest success rate at 17/30 but took 272 seconds per task.
Qwen3.8 Max Matches Claude Opus 4.8 on Intelligence Index, But Costs 2x More Per Task Than Predecessor
Alibaba's Qwen3.8 Max jumps 10 points to 56 on the Artificial Analysis Intelligence Index, putting it on par with Claude Opus 4.8. But Kimi K3 still edges it out at a lower per-task cost, and Qwen3.8 Max shows a sharp rise in hallucination rate.
Comments
Loading...