Mistral AI Releases Magistral Reasoning Models: 24B Open-Source and Enterprise Versions Score 70.7% and 73.6% on AIME202
Mistral AI has released Magistral, its first reasoning model line, in two versions: Magistral Small (24B parameters, Apache 2.0) and Magistral Medium (enterprise). Magistral Medium scored 73.6% on AIME2024 (90% with majority voting at 64 samples), while the open-source Small version achieved 70.7% (83.3% with voting).
Mistral AI Releases Magistral Reasoning Models: 24B Open-Source and Enterprise Versions Score 70.7% and 73.6% on AIME2024
Mistral AI has released Magistral, its first reasoning model line, in two versions: Magistral Small (24B parameters, Apache 2.0) and Magistral Medium (enterprise). According to Mistral AI, Magistral Medium scored 73.6% on AIME2024 (90% with majority voting at 64 samples), while the open-source Small version achieved 70.7% (83.3% with voting).
Technical Specifications
Magistral Small:
- 24 billion parameters
- Apache 2.0 license (open-source)
- Available on Hugging Face for self-deployment
- AIME2024: 70.7% (single attempt), 83.3% (majority voting @64)
Magistral Medium:
- Parameter count not disclosed
- Enterprise-only version
- AIME2024: 73.6% (single attempt), 90% (majority voting @64)
- Available via La Plateforme API, Le Chat, Amazon SageMaker
- Coming soon to IBM WatsonX, Azure AI, Google Cloud Marketplace
Core Capabilities
Both models feature chain-of-thought reasoning that operates natively across multiple languages and alphabets, including English, French, Spanish, German, Italian, Arabic, Russian, and Simplified Chinese. Mistral AI emphasizes the models are designed for multi-step logic with transparent, traceable thought processes.
The company claims Magistral Medium achieves up to 10x faster token throughput than most competitors when using "Flash Answers" mode in Le Chat, though specific token-per-second figures were not provided.
Target Applications
Mistral AI positions Magistral for:
- Legal research and compliance (with auditable reasoning chains)
- Financial forecasting and risk modeling
- Software development and system architecture
- Strategic planning and operational optimization
- Creative writing and content generation
The models are designed for use cases requiring "longer thought processing and better accuracy than with non-reasoning LLMs," according to the company.
Pricing and Access
Pricing for Magistral Medium API access was not disclosed. Magistral Small is free to download and deploy under Apache 2.0 license. Enterprise customers can request on-premises deployments through Mistral's sales team.
Research Publication
Mistral AI released an accompanying research paper covering training infrastructure, reinforcement learning algorithms, and evaluations. The company stated it aims to "iterate the model quickly" with constant improvements expected.
What This Means
Mistral's entry into reasoning models with an open-source 24B parameter option creates a new benchmark for accessible chain-of-thought models. The AIME2024 scores place both versions competitively within the current reasoning model landscape, though direct comparisons require knowing context windows and pricing structures still undisclosed for the Medium version. The multilingual chain-of-thought capability addresses a genuine gap in existing reasoning models, which typically perform best in English. However, the 10x speed claim lacks the specific comparative data needed for verification.
Related Articles
Z.ai Launches GLM-5.3-Flash: 1M-Token Context, Image Support, Claimed 10x Cost Cut Over GLM-5.2
Z.ai has released GLM-5.3-Flash, a 320-billion-parameter Mixture-of-Experts model with 18 billion active parameters, a 1-million-token context window, and image input support. The model launched on LM Studio's Bionic platform hours after its official unveiling, with LM Studio claiming it is 9-10x cheaper to run than GLM-5.2.
Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context
Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.
Google Launches Gemini 3.5 Transcribe, a Speech-to-Text Model That Cleans Up Rambling Speech
Google has released Gemini 3.5 Transcribe, a new speech-to-text model that automatically detects over 85 languages, removes filler words, and structures unstructured speech into clean text. The model powers Android's Rambler feature and is rolling out to Chrome, Docs, Gmail, and other Google products.
Google DeepMind Launches Gemini 3.5 Transcribe, Claims 2.6% Word Error Rate in Testing
Google DeepMind has released Gemini 3.5 Transcribe, a speech-to-text model available via two APIs for real-time streaming and pre-recorded audio. According to Artificial Analysis benchmarks cited by Google, the model achieves a 2.6% word error rate for non-streaming transcription and 4.0% for streaming.
Comments
Loading...