model releaseTiiuae

TII's Falcon-Emirati-7B scores 84.83% on Alyah, a new Emirati-dialect Arabic benchmark

TL;DR

The Technology Innovation Institute (TII) released Falcon-Emirati-7B, a 7B-parameter model specialized for Emirati Arabic and built on Falcon-H1-Arabic. TII claims it scores 84.83% on the Alyah benchmark, ahead of every Arabic and multilingual model it compared against.

3 min read
0

The Technology Innovation Institute (TII) has released Falcon-Emirati-7B, a 7-billion-parameter model specialized for Emirati Arabic. According to TII, it scores 84.83% on Alyah, a native Emirati-dialect benchmark, ahead of every other Arabic and multilingual model in its comparison set. The scores are TII's own and have not been independently verified.

What was released

Falcon-Emirati-7B is a dialect-specialized chat model built on the 7B variant of Falcon-H1-Arabic, TII's Arabic model family. TII published the announcement on Hugging Face on October 6, 2026.

The base family uses the Falcon-H1 hybrid architecture. Mamba-based State Space Model layers and Transformer attention run in parallel inside every block, and their outputs are fused before each block's projection. According to TII, the family:

  • spans 3B, 7B, and 34B parameters
  • supports context windows up to 128K and 256K tokens
  • was trained on Modern Standard Arabic (MSA) and on Gulf, Levantine, Egyptian, and Maghrebi dialects, alongside English and multilingual data

TII does not state the context window specific to Falcon-Emirati-7B in the material reviewed. It chose the 7B size as the balance between quality and cost. The 34B model would "likely push quality a bit further" at a cost TII judged unjustified, and the 3B model lacked headroom for cultural and linguistic depth.

Training data

TII built a dedicated Emirati data pipeline on top of Falcon-H1-Arabic's pretraining, using three sources:

  1. Native dialect web data. Content crawled from Emirati websites and forums, written in the dialect rather than translated or transliterated from MSA.
  2. MSA material on Emirati culture. Articles and references on customs, heritage, values, and social norms, including how Emiratis are perceived and stereotyped. This supplies cultural grounding rather than dialect production.
  3. Constrained synthetic data. Emirati-dialect text generated under strict rules and glossaries or dictionaries built for Emirati vocabulary and grammar, to cover topics the authentic text did not.

TII says no standard recipe exists for MSA-to-dialect adaptation. It ran ablations on how much dialectal data to inject and at which training stage, how to balance crawled against synthetic data without overfitting to synthetic patterns, and how much MSA cultural context was needed. The team combined automatic scoring with native-speaker review throughout.

Evaluation

TII used two methods: manual review by Emirati native speakers, who judged naturalness, tone, and cultural appropriateness, and automatic scoring on Alyah (الياه, "North Star").

Alyah is a fully native multiple-choice benchmark of 1,173 samples, collected manually from native Emirati speakers. Its categories include everyday greetings and etiquette, figurative language, heritage knowledge, and Emirati poetry. TII released it with the community.

Falcon-Emirati-7B reaches 84.83% on Alyah, according to TII. The announcement text available for this article cuts off before the full comparison table, so the specific competitor models and their scores are not confirmed here. Pricing, license terms, and hosted-API availability were not disclosed in the reviewed material.

What this means

The release reflects a broader shift from general multilingual Arabic models toward dialect-specific ones. MSA-trained models can translate Emirati sentences word for word and still miss idiom, proverb, and nabati poetry, which is the failure Alyah is designed to measure.

Two caveats apply. First, Alyah was co-created by the same organization reporting the lead, and the headline number comes from TII's own comparison. Independent replication on Alyah or on other dialect benchmarks is needed. Second, a multiple-choice format measures comprehension of dialect and culture, not open-ended generation quality. That is why TII paired it with native-speaker review.

For builders, a 7B hybrid Mamba-Transformer model is cheap to serve compared with the 34B sibling. That makes it a plausible option for UAE-focused assistants, customer service, and cultural-content applications where tone matters as much as correctness. TII's disclosure that the synthetic data was glossary-constrained is also a usable template for other low-resource dialects.

Related Articles

model release

Reflection AI unveils Beam: 501B-parameter open-weight MoE with 1M-token context

Reflection AI has unveiled Beam, a text-only mixture-of-experts model with 501 billion total parameters, 23 billion active, and a 1 million token context window. The company claims it matches Z.ai's GLM-5.2 on advanced reasoning benchmarks while using 3-4x less inference compute. Weights and the full technical report are due later this month.

model release

Reka AI releases Rho-1, a 19B-parameter omni-model for text, image, video and robot control

Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions in a single network. Reka says it uses no tool calls or external models. Context window, pricing, and benchmark scores have not been disclosed.

model release

China Telecom's Xing4.0-29B-A4B: 29B MoE, 4B Active, 256K Context, Trained Fully on Ascend NPUs

China Telecom AI's Xing4.0-29B-A4B (formerly the TeleChat line) is a mixture-of-experts model with 29B total and 4B active parameters and a native 256K context window, extensible to 512K. The company claims it is the first model of this scale trained entirely on Ascend NPUs with MindSpore. Community GGUF quantizations from Venastine-Research are already available.

model release

Cloudflare releases Clef decision models, claims 39 ms median latency vs. 524 ms for TypeSafe's Jev

Cloudflare has released Clef and Clef-flash, two open-weight decision models that return probabilities over predefined answer options instead of generating text. The company claims median latencies of 39 ms and 209 ms, against just over 524 ms for TypeSafe AI's Jev. Both support text and images and are API-compatible with Jev.

Comments

Loading...