Moonshot AI Releases Kimi K3: Open-Weight 2.8T-Parameter Model With 1M-Token Context and Native Multimodality
Moonshot AI has released Kimi K3, an open-weight 2.8-trillion-parameter mixture-of-experts model with 104B activated parameters, a 1,048,576-token context window, and native multimodal support. The company describes it as the world's first open 3T-class model, built on a new Kimi Delta Attention architecture.
Moonshot AI has released Kimi K3, an open-weight mixture-of-experts (MoE) model with 2.8 trillion total parameters and 104 billion activated parameters per token. The model supports a 1,048,576-token (1M) context window and processes text, images, and video natively within a single architecture.
According to Moonshot AI, Kimi K3 is the "world's first open 3T-class model," positioned as the company's most capable release to date.
Architecture
Kimi K3 introduces two new architectural components: Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). The model has 93 layers total (92 MoE, 1 dense), split into 69 KDA attention layers and 24 Gated Multi-head Latent Attention (MLA) layers. It uses a Stable LatentMoE framework with 896 total experts, activating 16 per token plus 2 shared experts. Moonshot AI claims this design delivers roughly a 2.5x improvement in scaling efficiency over its predecessor, Kimi K2.
Other specifications include a 7168-dimension attention hidden size, 96 attention heads, a 160K-token vocabulary, and a SiTU-GLU activation function. The vision component, MoonViT-V2, adds 401 million parameters for native image and video understanding. The model was trained with quantization-aware training using MXFP4 weights and MXFP8 activations.
Capabilities
Moonshot AI positions Kimi K3 for long-horizon agentic work: sustained coding sessions across large repositories, GPU kernel optimization, compiler development, CAD, and chip design, alongside knowledge work such as deep research, dashboard generation, and video editing — all described as operating with minimal human oversight.
Benchmark Results
Moonshot AI reports the following scores for Kimi K3 at maximum reasoning effort, benchmarked against comparison models including Claude Opus 4.8, GPT-5.5, and GLM-5.2:
- GPQA Diamond: 93.5
- HLE-Full: 43.5 (without tools) / 56.0 (with tools)
- Terminal-Bench 2.1: 88.3
- BrowseComp: 91.2
- OSWorld-Verified: 84.8
- DeepSWE: 67.5
These figures come directly from Moonshot AI's release materials and have not been independently verified. The company notes that comparison scores were pulled from a mix of official leaderboards, Artificial Analysis, and vendor-published benchmarks as of July 2026, and that harness choice (e.g., Kimi Code harness vs. Codex vs. Claude Code) varies across models, which can materially affect results.
Availability
Model weights are released under the Kimi K3 License, Moonshot AI's open-weight license permitting research, deployment, and derivative work. Pricing for hosted access has not yet been disclosed.
What This Means
A 2.8T-parameter open-weight model with a 1M-token context window and native multimodality is a significant technical claim, and if the reported 104B active-parameter efficiency holds up under independent testing, it could meaningfully lower the cost of running frontier-scale inference compared to dense models of similar total size. However, the benchmark comparisons rely heavily on Moonshot AI's own harness selections and self-reported figures against models and benchmark suites (including several referenced as being current in mid-2026) that have not been independently corroborated. Buyers and researchers should treat the specific benchmark deltas as directional rather than definitive until third-party evaluations — and confirmed pricing — are available.
Related Articles
Alibaba Releases Qwen3.8-Max, a 2.4 Trillion-Parameter Model Built for Multi-Day Autonomous Tasks
Alibaba has released Qwen3.8-Max, a 2.4-trillion-parameter model with 95 billion active parameters per query, designed to run autonomous tasks over multiple days. The company claims it hits 93 on PaperBench and rivals Claude Opus 4.8 and GPT-5.6 Sol on internal benchmarks, with open weights arriving next week.
LG AI Research Releases K-EXAONE 2.0, a 750B-Parameter Open-Weight MoE Model with 262K Context
LG AI Research has released K-EXAONE 2.0, a 750-billion-parameter mixture-of-experts language model with 37B active parameters, a 262,144-token context window, and support for 10 languages. The model is open-weighted under Apache 2.0 and claims competitive results against Qwen3.5, GLM-5.1, and DeepSeek-V4 Pro on reasoning, coding, and long-context benchmarks.
MiniMax Releases H3, a 33B-Parameter Omni-Modal Model That Generates 2K Video With Native Stereo Audio
MiniMax has published MiniMax-H3, a 33-billion-parameter omni-modal generative model capable of producing up to 15 seconds of 2K video with native stereo audio. The model accepts text, image, video, and audio inputs, though its full 2K pipeline depends on a hosted preprocessing component not included in the open-source release.
Liquid AI Releases LFM2.5-2.6B, a 2.6B-Parameter Agentic Model with 128K Context for On-Device Use
Liquid AI has released LFM2.5-2.6B, a 2.6B-parameter model trained on 34 trillion tokens with a 128K context window, built for on-device agentic workloads. The company claims it is competitive with models four times its size on tool use and instruction following.
Comments
Loading...