model releaseCohere

Cohere releases 2B parameter Arabic speech recognition model with 25.9% average WER

TL;DR

Cohere and Cohere Labs released Cohere Transcribe Arabic, a 2B parameter automatic speech recognition model optimized for Arabic dialects and Arabic-English code-switching. The open-source model achieves a 25.9% average word error rate across major Arabic ASR benchmarks, outperforming models up to 30B parameters.

2 min read
0

Cohere releases 2B parameter Arabic speech recognition model with 25.9% average WER

Cohere and Cohere Labs released Cohere Transcribe Arabic, a 2B parameter automatic speech recognition (ASR) model optimized for Arabic dialects and Arabic-English code-switching. The model tops the Open Universal Arabic ASR Leaderboard with a 25.87% average word error rate (WER) and 11.80% character error rate (CER).

Architecture and capabilities

The model uses a Conformer-based encoder-decoder architecture. A large Conformer encoder extracts acoustic representations from audio input, followed by a lightweight Transformer decoder for text generation. Audio inputs are automatically resampled to 16kHz and converted to log-Mel spectrograms during preprocessing.

Cohere Transcribe Arabic supports both Arabic and English transcription. For long-form audio, the feature extractor automatically segments waveforms into chunks, with the processor reassembling transcriptions using audio chunk indices.

Benchmark performance

According to Cohere, the model achieves the following WER scores across Arabic ASR benchmarks:

  • SADA: 37.47%
  • Common Voice: 5.82%
  • MASC clean: 19.60%
  • MASC noisy: 27.07%
  • MGB-2: 15.54%
  • Casablanca: 49.71%

The model outperforms larger competitors including OmniASR LLM 7B (28.32% average WER), Qwen3-Omni 30B (30.71% average WER), and Whisper Large v3 (36.86% average WER).

Integration and deployment

The model is natively supported in Transformers 5.4.0+ and available under Apache 2.0 license. For production deployment, Cohere recommends using vLLM 0.19.0 for serving.

Cohere provides code examples for single-file transcription, long-form audio processing, and vLLM server setup. The model requires the processor to specify language ("ar" or "en") during inference.

Limitations

Cohere notes the model performs best with single-language audio and lacks automatic language detection. Performance on code-switched audio is inconsistent. The model does not support timestamps or speaker diarization. Like most audio encoder-decoder models, it may hallucinate transcriptions for non-speech audio without a voice activity detection preprocessor.

What this means

This release provides the first competitive open-source ASR model specifically optimized for Arabic dialects, an underserved segment in speech recognition. The 2B parameter count makes it deployable on consumer hardware while matching or exceeding models 15x larger. The Apache 2.0 license enables commercial use without restrictions, potentially accelerating Arabic speech applications in regions where proprietary APIs have limited availability or high costs.

Related Articles

model release

TII releases 1.6B Falcon-ASR, claims 20.92% Arabic WER against best listed 23.17%

The Technology Innovation Institute (TII) released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic and the Emirati dialect. TII claims a 20.92% average word error rate across six Arabic test sets, versus 23.17% for the next-best system on the leaderboard snapshot it used. A demo is live on Hugging Face. Pricing and API availability have not been disclosed.

model release

Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta

Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.

model release

Microsoft releases FrogNano-4B, an Apache 2.0 coding agent trained with RL on 1,500 synthetic tasks

Microsoft has released FrogNano-4B-2609, a repository-level coding agent derived from Qwen3.5-4B and published under Apache 2.0 with open weights. Microsoft says it was post-trained only with reinforcement learning on about 1,500 synthetic software-engineering tasks, with no stronger-model trajectories. It is evaluated at roughly 131K tokens of context.

model release

Qwen releases Qwen-Image-2.1-Turbo: 8-step text-to-image and editing checkpoint on a 7B architecture

Qwen has published Qwen-Image-2.1-Turbo on Hugging Face, an accelerated checkpoint of Qwen-Image-2.1 that runs text-to-image generation and image editing in 8 denoising steps. It keeps the same 7B visual generation architecture and loads through a new QwenImage21Pipeline in Diffusers.

Comments

Loading...