model releaseCohere

Cohere Releases Command A+ Open Source Model with 25B Active Parameters, 128K Context

TL;DR

Cohere has released Command A+ as an open source model under Apache 2.0 license. The sparse mixture-of-experts architecture features 25 billion active parameters out of 218B total parameters, supports 128K input context length, and includes vision capabilities alongside tool use and reasoning features.

2 min read
0

Command A+ — Quick Specs

Context window128K tokens
Input$2.5/1M tokens
Output$10/1M tokens

Cohere Releases Command A+ Open Source Model with 25B Active Parameters, 128K Context

Cohere has released Command A+ (command-a-plus-05-2026) as an open source model under Apache 2.0 license. The sparse mixture-of-experts architecture features 25 billion active parameters out of 218B total parameters, supports 128K input context length with 64K output length, and includes multimodal vision capabilities.

Architecture and Specifications

Command A+ uses a decoder-only Sparse Mixture-of-Experts Transformer architecture with 128 experts, activating 8 per token plus one shared expert applied to all tokens. According to Cohere, the model employs a 3:1 ratio of sliding-window attention layers with Rotational Positional Embeddings to global attention layers without positional embeddings, a design first introduced in the earlier Command A model.

The sparse MoE layer is trained in a "fully dropless manner" using a token-choice router, with additive-bias-based load balancing to distribute token load across experts. The architecture replaces the standard softmax router activation function with a normalized sigmoid over the topk expert logits per token.

Deployment and Hardware Requirements

Cohere provides three quantization options with minimal quality differences:

  • BF16 (16-bit): Requires 4x B200 or 8x H100 GPUs
  • FP8 (8-bit): Requires 2x B200 or 4x H100 GPUs
  • W4A4 (4-bit): Requires 1x B200 or 2x H100 GPUs

Cohere recommends the W4A4 quantization for most use cases, claiming "superior speed and latency characteristics alongside a smaller hardware footprint."

Capabilities

The model supports 48 languages including English, Chinese, Japanese, Arabic, Spanish, and various European and Asian languages. It includes native tool use capabilities trained for conversational API interactions, with support for JSON schema tool descriptions and citation generation to ground responses in specific tool results.

Command A+ includes a reasoning mode that generates explicit thinking steps between <START_THINKING> and <END_THINKING> tags before producing final outputs. The model also accepts image inputs for multimodal processing.

Integration

The model requires transformers installation from source and is compatible with vLLM 0.21.0 or higher. Tool calling and reasoning parsing require Cohere's melody library (version 0.9.0+). The model is available on Hugging Face with a hosted demo space for testing before deployment.

What This Means

Command A+ enters the competitive open source model space with a sparse MoE design similar to Mixtral and DeepSeek's architectures, but with significantly more total parameters (218B vs Mixtral 8x22B's 141B). The 128K context window matches GPT-4 Turbo and Claude 3 capabilities, while the Apache 2.0 license allows unrestricted commercial use. The model's combination of vision, reasoning, and tool use in a single open source package targets enterprise deployments that previously required closed-source API providers.

Related Articles

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

model release

Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta

Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.

model release

inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter

inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.

model release

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

Comments

Loading...