This model is 16 months old. The same company has since released something newer: Muse Glimmer 30B (2026-08-10). The details below are still accurate for this model — it just is not what we would recommend today.

Llama 4 Scout

Meta AI🇺🇸 United States
active

First Llama 4 model with 10M token context and native multimodal support.

Context window10000K tokens
Input / 1M tokens$0.17
Output / 1M tokens$0.66

Version History

Llama-4-Scout-17B-16E-0921patch

Llama 4 Scout patch fixing multimodal tokenization issues and improving throughput for long-context document tasks.

Llama-4-Scout-17B-16Emajor

First MoE Llama with 10M token context and native image and video understanding.

Benchmark Scores

Full leaderboard →
94.4%
DocVQA
79.4%
Hallucination Rate
88.7%
MMLU
160.0 tokens_per_sec
Speed (tok/s)