Mistral Releases Mistral 3 Family: 675B-Parameter Large 3 MoE and Three Edge Models Under Apache 2.0
Mistral has released Mistral 3, including Mistral Large 3—a sparse mixture-of-experts model with 41B active and 675B total parameters—and three Ministral 3 edge models (3B, 8B, 14B). All models are released under Apache 2.0 license with multimodal capabilities and are available today on multiple platforms.
Mistral Large 3 — Quick Specs
Mistral Releases Mistral 3 Family: 675B-Parameter Large 3 MoE and Three Edge Models Under Apache 2.0
Mistral has released Mistral 3, a model family spanning from 3B to 675B parameters, all under the Apache 2.0 license. The release includes Mistral Large 3, a sparse mixture-of-experts architecture with 41B active parameters and 675B total parameters, alongside three Ministral 3 edge models at 3B, 8B, and 14B sizes.
Mistral Large 3 Specifications
Mistral Large 3 was trained from scratch on 3,000 NVIDIA H200 GPUs. According to Mistral, the model ranks #2 among open-source non-reasoning models on LMArena and #6 among all open-source models. The company claims the model achieves "parity with the best instruction-tuned open-weight models" on general prompts.
Pricing for Mistral Large 3:
- Input: $0.50 per 1M tokens
- Output: $1.50 per 1M tokens
Both base and instruction-tuned versions are available. A reasoning version is planned for future release.
Ministral 3 Edge Models
The Ministral 3 series includes three parameter sizes: 3B, 8B, and 14B. Each size offers base, instruct, and reasoning variants with image understanding capabilities. Mistral claims the 14B reasoning variant achieves 85% on AIME 2025.
Pricing for Ministral 3 8B (pricing for other sizes not disclosed):
- Input: $0.15 per 1M tokens
- Output: $0.15 per 1M tokens
According to Mistral, the instruct models generate "an order of magnitude fewer tokens" than comparable models while matching or exceeding performance.
Technical Implementation
Mistral partnered with NVIDIA, vLLM, and Red Hat for deployment optimization. The company released a checkpoint in NVFP4 format using llm-compressor, enabling Mistral Large 3 to run on a single 8×A100 or 8×H100 node via vLLM. NVIDIA integrated Blackwell attention and MoE kernels for efficient inference on GB200 NVL72 systems.
For edge deployment, NVIDIA delivers optimized deployments on DGX Spark, RTX PCs, and Jetson devices.
Availability
All Mistral 3 models are available today on Mistral AI Studio, Amazon Bedrock, Azure Foundry, Hugging Face, Modal, IBM WatsonX, OpenRouter, Fireworks, Unsloth AI, and Together AI. NVIDIA NIM and AWS SageMaker availability is coming soon.
What This Means
Mistral's Apache 2.0 licensing decision for a 675B-parameter model represents the largest permissively-licensed model release to date, potentially accelerating enterprise adoption of open-weight alternatives to proprietary models. The sparse MoE architecture with 41B active parameters positions Large 3 as computationally efficient compared to dense models of similar capability, though real-world cost-effectiveness will depend on actual serving infrastructure requirements and the efficiency gains from the optimized NVFP4 format.
Related Articles
Google DeepMind Launches Gemini 3.8 Live, Claims #1 Spot on Speech-to-Speech Benchmark
Google DeepMind has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice-dialogue models that reason and execute background tasks without interrupting conversation. Google claims the Extended Thinking model ranks #1 on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6.
Tencent Open-Sources AuK, a 1.5B-Parameter Speech Generation and Editing Model
Tencent has open-sourced AuK, a 1.5B-parameter foundation model for speech generation and editing that handles TTS, content editing, and audio enhancement through natural-language instructions. The release includes a distilled AuK-Flash variant for 4-step fast inference, both under MIT license.
Unverified 'GPT Astra' Model Appears on OpenRouter With 1.05M Token Context, No OpenAI Confirmation
OpenRouter is listing a model called 'OpenAI GPT Astra Latest' with a 1.05 million token context window and $10/$50 per-million-token pricing. OpenAI has made no public announcement, and the listing's own description says it is an auto-redirecting alias rather than a fixed model.
OpenRouter Lists 'GPT Sol Latest' — An Alias Pointer to OpenAI's Newest Sol-Family Model, Not a Standalone Release
OpenRouter has added a listing called '~openai/gpt-sol-latest,' described as an alias that always points to the newest model in an undisclosed 'GPT Sol' family from OpenAI. The listing shows a 1050K token context window and pricing of $2.00 per million input tokens and $10.00 per million output tokens, but OpenAI has not publicly confirmed a model line by this name.
Comments
Loading...