model releaseCohere

Cohere releases tiny-aya-global, multilingual text model covering 100+ languages

TL;DR

Cohere Labs has released tiny-aya-global, a lightweight text generation model trained to support conversational tasks across 100+ languages. The model is available on Hugging Face under a CC-BY-NC-4.0 license and builds on the tiny-aya-base architecture.

1 min read
0

Cohere Labs has released tiny-aya-global, a lightweight multilingual text generation model designed for conversational tasks across 100+ languages.

Model Details

The tiny-aya-global model is available on Hugging Face and supports text generation and conversational applications. It's built as a fine-tuned variant of the tiny-aya-base model, optimized for multilingual performance across a broad language spectrum.

The model covers extensive language support spanning:

  • Major European languages: English, Dutch, French, Italian, Portuguese, Romanian, Spanish, Czech, Polish, Ukrainian, Russian, Greek, German, Danish, Swedish, Norwegian, Catalan, Galician, Welsh, Irish, Basque, Croatian, Latvian, Lithuanian, Slovak, Slovenian, Estonian, Finnish, Hungarian, Serbian, Bulgarian
  • Middle Eastern and South Asian languages: Arabic, Farsi, Urdu, Turkish, Maltese, Hebrew, Hindi, Marathi, Bengali, Gujarati, Punjabi, Tamil, Telugu, Nepali
  • Southeast Asian languages: Tagalog, Malay, Indonesian, Javanese, Khmer, Thai, Lao, Burmese
  • East Asian languages: Chinese, Japanese, Korean
  • African languages: Amharic, Hausa, Igbo, Malagasy, Shona, Swahili, Wolof, Xhosa, Yoruba, Zulu

Availability and Licensing

The model is distributed under the CC-BY-NC-4.0 license and can be accessed directly from Hugging Face. As of release, the model has generated 1,204 downloads and received 56 likes on the platform, indicating early adoption among developers building multilingual applications.

The model uses the transformers library and safetensors format for compatibility with standard ML tooling.

What This Means

Tiny-aya-global fills a gap in accessible multilingual models by providing a lightweight alternative for organizations needing broad language coverage without deploying massive foundation models. The focus on lower-resource languages alongside major ones suggests Cohere's intent to extend conversational AI capabilities beyond English-dominant markets. For teams building applications in non-English regions, the model's breadth of language support reduces the need to manage separate language-specific models or expensive API calls to large closed models.

Related Articles

model release

Moonshot AI and Alibaba release 2.8T and 2.4T parameter models, claim performance near GPT-5.6 and Claude Fable 5

Within days, Moonshot AI and Alibaba unveiled what they claim are frontier-class models. Moonshot's Kimi K3, at 2.8 trillion parameters, and Alibaba's Qwen3.8, at 2.4 trillion parameters, will both be released as open-weight models with full weights available for download.

model release

NVIDIA Releases Nemotron-3-Embed-1B-BF16: 1.14B Parameter Multilingual Embedding Model with 2048-Dimensional Vectors

NVIDIA has released Nemotron-3-Embed-1B-BF16, a 1.14 billion parameter text embedding model supporting 34 languages with a 32,768 token context window. The model generates 2048-dimensional embeddings and was derived from Ministral-3-3B-Instruct-2512 through two rounds of structured pruning and distillation, first to 2B then to 1.14B parameters.

model release

Alibaba releases Qwen 3.8, a 2.4 trillion parameter open-weight model claiming second place behind Fable 5

Alibaba has released Qwen 3.8, a 2.4 trillion parameter open-weight model that the company claims trails only Fable 5. The multimodal model processes images, videos, and documents, with a preview available through Alibaba's platforms at 10 percent of standard pricing.

model release

Moonshot AI Releases Kimi K3: 2.8T Parameter Open Model at $3/$15 Per Million Tokens

Moonshot AI has released Kimi K3, a 2.8 trillion parameter model with 1 million token context window and native multimodal input. The model ranks #1 in Frontend Code Arena and #9 in Text Arena, with pricing at $3 per million input tokens and $15 per million output tokens—comparable to Claude Sonnet 5 pricing while delivering performance the company claims is near Claude Opus 4.8 and GPT-5.5.

Comments

Loading...