Tencent Releases Hy3-Preview: 295B-Parameter MoE Model with 21B Active Parameters
Tencent has released Hy3-preview, a 295-billion-parameter Mixture-of-Experts model with 21 billion active parameters and a 256K context window. The model scores 76.28% on MATH and 34.86% on LiveCodeBench-v6, with particularly strong performance on coding agent tasks.
Hy3 Preview — Quick Specs
Tencent Releases Hy3-Preview: 295B-Parameter MoE Model with 21B Active Parameters
Tencent has released Hy3-preview, a 295-billion-parameter Mixture-of-Experts (MoE) model with 21 billion active parameters. The model features a 256K context window and includes 3.8 billion parameters in its MTP (Multi-Token Prediction) layer. Both base and instruct versions are now available on Hugging Face, ModelScope, and GitCode under the Tencent Hy Community License.
Architecture and Specifications
Hy3-preview uses a sparse MoE architecture with 192 experts and top-8 activation, meaning 8 experts are activated per token. The model has 80 standard transformer layers plus 1 MTP layer, with 64 attention heads using grouped query attention (8 KV heads). It requires approximately 8 H20 GPUs or equivalent hardware for inference.
The model supports BF16 precision and has a vocabulary size of 120,832 tokens. According to Tencent, this is "the first model trained on our rebuilt infrastructure, and the strongest we've shipped so far."
Benchmark Performance
On base model benchmarks, Hy3-preview scores:
- MATH: 76.28% (4-shot)
- GSM8K: 95.37% (4-shot)
- MMLU: 87.42% (5-shot)
- MMLU-Pro: 65.76% (5-shot)
- LiveCodeBench-v6: 34.86% (1-shot)
- CRUXEval-I: 71.19% (3-shot)
- C-Eval: 89.80% (5-shot)
- MMMLU: 80.15% (5-shot)
The company claims particularly strong results on coding agent benchmarks including SWE-bench Verified and Terminal-Bench 2.0, as well as search agent benchmarks like BrowseComp and WideSearch. Tencent also reports "excellent results" on the Tsinghua Qiuzhen College Math PhD qualifying exam (Spring 2026) and China High School Biology Olympiad 2025, though specific scores were not disclosed.
Context Learning and Reasoning Modes
Tencent built two proprietary benchmarks—CL-bench and CL-bench-Life—to measure context learning ability in real business scenarios. The model supports three reasoning modes via the reasoning_effort parameter: "no_think" for direct responses, "low" for moderate reasoning, and "high" for deep chain-of-thought processing.
Recommended inference parameters are temperature=0.9 and top_p=1.0. The model can be deployed using vLLM or SGLang with MTP-enabled speculative decoding.
Training and Quantization
The release includes a complete training pipeline supporting both full fine-tuning and LoRA, with DeepSpeed ZeRO configurations and LLaMA-Factory integration. Tencent provides AngelSlim, a compression toolkit supporting quantization algorithms, low-bit quantization, and speculative sampling.
Pricing information has not been disclosed. The model is available for commercial use under Tencent's community license.
What This Means
Hy3-preview represents Tencent's entry into the 200B+ parameter MoE space, competing directly with models like DeepSeek-V3 (671B total, 37B active) and Kimi-K2 (1043B total, 32B active). With 21B active parameters versus competitors' 32-37B, Hy3-preview achieves competitive performance while potentially reducing inference costs. The strong LiveCodeBench score (34.86% vs DeepSeek-V3's 29.31%) and emphasis on agent capabilities suggests Tencent is positioning this model for practical coding and agentic applications rather than pure benchmark optimization. The 256K context window and custom context-learning benchmarks indicate a focus on real-world enterprise use cases.
Related Articles
Meituan launches LongCat 2.0: 1.6T parameter MoE model with 1M+ context window at $0.30 per 1M input tokens
Meituan has released LongCat 2.0, a sparse mixture-of-experts language model with 48 billion active parameters out of 1.6 trillion total. The model features a 1,049,000 token context window and costs $0.30 per 1M input tokens and $1.20 per 1M output tokens.
Moonshot AI Releases Kimi K3: 2.8T Parameter Open Model at $3/$15 Per Million Tokens
Moonshot AI has released Kimi K3, a 2.8 trillion parameter model with 1 million token context window and native multimodal input. The model ranks #1 in Frontend Code Arena and #9 in Text Arena, with pricing at $3 per million input tokens and $15 per million output tokens—comparable to Claude Sonnet 5 pricing while delivering performance the company claims is near Claude Opus 4.8 and GPT-5.5.
Thinking Machines releases Inkling: 975B-parameter MoE model with Apache 2.0 license, first major US open-weight multimo
Thinking Machines Lab released Inkling, a mixture-of-experts model with 975B total parameters and 41B active parameters, trained on 45 trillion tokens across text, images, audio, and video. The Apache 2.0-licensed model supports up to 1M context and debuts alongside Inkling-Small (276B-A12B), marking what observers call the strongest US-based open-weight release to date.
Moonshot AI's Kimi K3 ranks #2 globally, will release 2.8T parameter weights July 27
Moonshot AI released Kimi K3 on July 16, 2026, a 2.8 trillion parameter mixture-of-experts model that ranks #2 on the Vals AI index and #3 on Artificial Analysis's Intelligence Index. The company will release the model's weights on July 27, making it the strongest open-weight model to date, surpassing all previous open releases including DeepSeek R1.
Comments
Loading...