model release

Mistral Large 4 enters public preview: 1T-parameter open-weight multimodal model, weights due by end of October

TL;DR

Mistral AI has launched a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters. The preview API is live on Mistral Studio, and open weights are promised by the end of October 2026. Pricing and context window have not been disclosed.

3 min read
0

Mistral AI launched a public preview of Mistral Large 4 (ML4) on October 6, 2026. It is a 1-trillion-parameter, natively multimodal model with 49 billion active parameters. The preview API is available now on Mistral Studio, and the company says open weights will follow by the end of October.

Key specifications

  • Parameters: 1T total, 49B active per token
  • Modality: Natively multimodal (text and images)
  • Training hardware: 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters
  • Training data: Multilingual, spanning 160+ languages, including every official EU language, according to Mistral
  • Context window: Not disclosed
  • Pricing: Not yet disclosed
  • License and training cutoff: Not disclosed

Mistral says the preview is served from the same European infrastructure used for training. It also says the model was built with the same training, customization and RL environment offered to customers through Mistral Forge.

Benchmark claims

Mistral claims ML4 is competitive with the strongest open models globally and significantly outperforms any open-weight model developed in the US or Europe. These figures come from Mistral's announcement. Several rely on third-party indices (Artificial Analysis, vals.ai) as reported by Mistral. We have not independently verified them.

Coding and agents

  • DeepSWE v1.1: 61.7%
  • SWE-Atlas-QnA: 59.4%
  • Terminal-Bench 4: 28.3%
  • Coding Agent Index: 49.8%, which Mistral says is ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max
  • AutomationBench (657 business workflows): 59.9%, ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro
  • AA-Briefcase (long-horizon knowledge work): 1,393 Elo
  • Surge AI blind human evaluation of coding quality: 3.74 out of 5, ranking second of five models. Claude Opus 5 led at 4.22. Kimi K3 scored 3.59, GLM-5.2 scored 3.40 and GLM-5.3 scored 3.60.

Cybersecurity

  • Artificial Analysis Cyber Index: top five globally, per Mistral
  • Vulnerability reproduce-and-patch test within that index: 82%, the highest of any model according to Mistral
  • Cybench (40 competition-derived challenges): 93%

Mistral says Claude Opus 5.5 and GPT-6 Astra score near zero on the reproduce-and-patch test because they refuse the task. The model's score therefore partly reflects refusal behavior, not only raw capability.

Vision

  • Dense 200 visual grounding: 42%, versus 41% for GPT-6 Astra, per Mistral

Mistral also cites document, chart and geospatial use cases. The source text provided to us cuts off before the science and math section.

Release approach

Until the weights ship, Mistral is red-teaming ML4 with cybersecurity leaders, vetted partners and state authorities. These groups access a version with reduced moderation and expanded cyber capabilities. Mistral says further details on architecture, additional benchmarks and post-training methodology will follow. It also says ML4 will serve as the base for a new generation of specialized Mistral models.

The model will be offered in multiple regions, including a European deployment that Mistral says it operates end-to-end under European law.

What this means

A 1T-parameter model with 49B active parameters puts Mistral in the same size class as the largest open-weight releases from Chinese labs. If the weights ship on schedule, it would be the largest open-weight model from a US or European developer. Whether the weights are usable at that size depends on a license that has not been published.

The cyber positioning is the strategic bet. Mistral is selling the absence of provider-side refusals, plus self-hosting, as a feature for defenders. That claim is testable once weights are public, and it will draw scrutiny over misuse risk.

The coding results are mixed. The model places second in a human evaluation but sits behind Claude Opus 5. A 28.3% Terminal-Bench 4 score suggests long-horizon terminal work remains hard. Without a stated context window or pricing, teams cannot yet calculate deployment cost. Those two numbers, plus the license, are the main open questions before the weights drop.

Related Articles

model release

Mistral releases Large 4, a 1-trillion-parameter multimodal model, with open weights due in three weeks

Mistral AI released Mistral Large 4 (ML4), a multimodal model with one trillion parameters, on Tuesday. It is currently available only through a public guardrail endpoint, and Mistral plans to publish the weights in about three weeks after safety testing. Benchmark results, pricing and context window have not been disclosed.

model release

Reflection unveils 501B-parameter Beam, Mistral previews 1T-parameter Large 4, both open-weight

Reflection introduced Beam, a 501B-parameter mixture-of-experts model with 23B active parameters. Mistral said it is finishing Mistral Large 4, a 1T-parameter multimodal model with 49B active parameters. Both companies plan open-weight releases in October, and both are positioning the models against Chinese open-weight leaders.

model release

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

model release

Reflection announces Beam, a 501B-parameter open-weight coding model with 23B active parameters

Reflection has announced Beam, its first open-weight model, a 501B-parameter mixture-of-experts system with 23B active parameters per token, built for coding, reasoning and agentic tasks. The company claims it matches GLM 5.2 on demanding reasoning tasks with three to four times less compute. Weights are due under Apache 2.0 later this month.

Comments

Loading...