model release

TII releases 1.6B Falcon-ASR, claims 20.92% Arabic WER against best listed 23.17%

TL;DR

The Technology Innovation Institute (TII) released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic and the Emirati dialect. TII claims a 20.92% average word error rate across six Arabic test sets, versus 23.17% for the next-best system on the leaderboard snapshot it used. A demo is live on Hugging Face. Pricing and API availability have not been disclosed.

3 min read
0

The Technology Innovation Institute (TII) in Abu Dhabi has released Falcon-ASR, a 1.6B-parameter speech recognition model for Arabic with a focus on the Emirati dialect. According to TII, it reaches a 20.92% average word error rate (WER) across the six test sets of the Open Universal Arabic ASR Leaderboard protocol. The best published result in the snapshot TII used was 23.17%.

All figures below come from TII's own evaluations and have not been independently verified.

Key specs

  • Parameters: 1.6B
  • Languages: Arabic (Emirati, Modern Standard Arabic, other Gulf and Arabic dialects), English, French, Spanish, Portuguese
  • Language handling: all five languages share the same weights, with no language flag required
  • Output: transcripts in the spoken language, with word-level timestamps
  • Foundation: builds on TII's Falcon3-Audio work
  • Access: Hugging Face demo Space; API and native applications are "planned"
  • Pricing, context window, training cutoff: not yet disclosed

Arabic benchmark results

The leaderboard is maintained by the ELM Research Center and ranks systems by equal-weight average WER across six test sets. TII says it used the leaderboard's pinned manifests. Competitor figures are published leaderboard averages checked on 30 September 2026.

Model Params Avg WER (%) Avg CER (%)
Falcon-ASR 1.6B 20.92 8.79
Audar-ASR-V1-Turbo 2.35B 23.17 9.23
Cohere Transcribe Arabic (07-2026) 2.0B 25.87 11.80
omniASR LLM 7B 7.0B 28.32 12.52

The gap to the runner-up is 2.25 percentage points of WER.

Emirati dialect evaluation

TII supplemented public data, including the UAE subset of Casablanca, with an internal evaluation. It uses held-out Emirati and Gulf recordings with human-validated transcripts. The dataset is not public, so the results cannot be independently reproduced.

Model Params WER (%) CER (%)
Falcon-ASR 1.6B 22.73 10.19
Qwen3-Omni-30B-A3B-Instruct 30.0B (3.0B active) 26.80 12.72
Audar-ASR-V1-Turbo 2.35B 27.89 13.75
Cohere Transcribe Arabic (07-2026) 2.0B 31.05 18.07
Qwen3-ASR-1.7B-hf 2.0B 31.52 13.35
Audar-ASR-V1-Flash 0.78B 32.87 15.36

TII reports a 4.07-point WER advantage over Qwen3-Omni, the next-best system in this comparison.

English performance

On the seven public test sets used by the Hugging Face Open ASR Leaderboard, TII reports a mean WER of 5.74%:

  • LibriSpeech clean: 1.75%
  • LibriSpeech other: 4.21%
  • SPGISpeech: 2.02%
  • VoxPopuli: 3.87%
  • GigaSpeech: 8.15%
  • AMI: 8.33%
  • Earnings-22: 11.86%

TII did not publish comparison numbers for other models on English in the announcement. It also did not publish per-language results for French, Spanish or Portuguese.

Training data and conditions

TII says it trained on Emirati, MSA, other Gulf and Arabic dialects, and English. It augmented data with background noise, overlapping speech, music, room reverberation, telephony effects, and speed and pitch variation. The same treatment was applied to Emirati recordings. The architecture and training approach for Falcon3-Audio are described in TII's paper, "Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data." The announcement does not state the size or composition of the Falcon-ASR training set.

What this means

A 1.6B model that beats larger Arabic systems on a public leaderboard protocol, if the numbers hold, shows that dialect-focused data can matter more than parameter count. The result is strongest on the public Arabic benchmark, where the protocol is fixed and the baselines are published. The Emirati lead rests on an internal test set that outside parties cannot check. Treat it as directional until independent evaluations appear.

Three things are missing for production teams: pricing, a stated license, and confirmed weight availability. The announcement points only to a demo, with API access "planned." Until those details arrive, Falcon-ASR is something to evaluate, not yet something to deploy.

Related Articles

model release

TII's Falcon-Emirati-7B scores 84.83% on Alyah, a new Emirati-dialect Arabic benchmark

The Technology Innovation Institute (TII) released Falcon-Emirati-7B, a 7B-parameter model specialized for Emirati Arabic and built on Falcon-H1-Arabic. TII claims it scores 84.83% on the Alyah benchmark, ahead of every Arabic and multilingual model it compared against.

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

model release

Musubi Releases Open-Weights PolicyLM-1.7B, a Sub-50ms Content Moderation Model

Musubi released PolicyLM-1.7B on Tuesday, a 1.7-billion-parameter open-weights decision model built for real-time content moderation. The company claims it applies plain-English policies to messages in under 50 milliseconds and needs no retraining when policies change.

model release

Google releases EmbeddingGemma 2, a 740M-parameter open multimodal embedding model scoring 78.68 on MTEB Code

Google released EmbeddingGemma 2, an open 740-million-parameter model that embeds text, images, video, audio, and code. It scores 78.68 on MTEB (Code), up from 68.76 for its predecessor, and Google claims it beats models up to twice its size on multimodal embedding benchmarks.

Comments

Loading...