model release

China Telecom's Xing4.0-29B-A4B: 29B MoE, 4B Active, 256K Context, Trained Fully on Ascend NPUs

TL;DR

China Telecom AI's Xing4.0-29B-A4B (formerly the TeleChat line) is a mixture-of-experts model with 29B total and 4B active parameters and a native 256K context window, extensible to 512K. The company claims it is the first model of this scale trained entirely on Ascend NPUs with MindSpore. Community GGUF quantizations from Venastine-Research are already available.

3 min read
0

China Telecom Artificial Intelligence Technology Co., Ltd. has released Xing4.0-29B-A4B, a mixture-of-experts language model with 29B total parameters and 4B activated per token, a native 256K context window (extensible to 512K), and training done entirely on Huawei Ascend NPUs. The company claims it is the first model of this scale trained fully on the Ascend platform with the MindSpore framework. This coverage is based on a GGUF quantization repository published by Venastine-Research; the base weights are listed under XingChen-AGI/Xing4.0-29B-A4B.

Xing is the successor branding to China Telecom's TeleChat series.

Architecture

According to the model card:

  • Parameters: 29B total, 4B active per token
  • Layers: 40, hidden size 3584
  • Attention: Multi-head Latent Attention (MLA)
  • Experts: 64 routed, 4 active per token, plus 1 shared; expert intermediate size 1024, dense FFN intermediate size 9216
  • Context: 256K native, extensible to 512K
  • Design: mHC + MLA + MTP (multi-token prediction), which the company says targets multi-step planning, tool calling and long-context coherence

Training ran on Ascend 910C clusters using MindSpore/MindFormers. China Telecom claims fine-grained MoE communication optimization, selective recomputation, DVM graph-operator fusion and Ascend C mHC fused operators raised training throughput by about 96% over out-of-the-box performance.

Benchmarks (self-reported)

The model card compares Xing4.0-29B-A4B with Gemma4-26B-A4B and Qwen3.6-35B-A3B. All figures come from China Telecom and have not been independently verified.

Benchmark Xing4.0-29B-A4B Gemma4-26B-A4B Qwen3.6-35B-A3B
SWE-bench Verified 75.00 53.00 76.00
SWE-bench Multilingual 66.00 51.00 67.20
Terminal-Bench 2.1 57.50 30.00 51.50
Claw-Eval 76.55 71.49 74.54
Tau3-Bench 64.63 58.90 67.20
DeepresearchBII 60.80 39.30 59.70
AIME2026 90.00 88.30 92.70
IFBench 69.67 72.67 65.50
AA.LCR 61.00 66.00 62.00

The model leads Qwen3.6-35B-A3B on Terminal-Bench 2.1, Claw-Eval, DeepresearchBII and IFBench, and trails it on the other five. It trails Gemma4-26B-A4B only on IFBench and AA.LCR. The SWE-bench runs use the SWE-agent harness with a 210K context window. Several scores are averages over 3 to 5 runs, and Terminal-Bench 2.1 uses a 24-hour timeout. Harness differences can shift agentic scores materially, so treat cross-model comparisons with caution.

Deployment and quantizations

The weights are in Hugging Face Transformers format and work with vLLM, SGLang and KTransformers. Fine-tuning is supported through LLaMA-Factory and MindFormers. The card says it has format alignment for agent frameworks including OpenCode, Claude Code, OpenClaw and Hermes. A thinking mode is toggled with enable_thinking in the chat template. Recommended sampling is temperature 1.0 for general and reasoning tasks and 0.8 for coding and agent tasks, with top_p 0.95 and repetition_penalty 1.05.

The Venastine-Research GGUF repo lists file sizes by quantization:

  • IQ2_M: 9.9 GB
  • IQ3_XXS: 11.6 GB (MTP variant also 11.6 GB)
  • IQ3_M: 13.9 GB (MTP variant 13.1 GB)
  • Q4_K_M: 19 GB
  • Q5_K_M: 22.2 GB
  • Q6_K: 25.7 GB
  • Q8_0: 33.2 GB
  • F16: 62.5 GB

The repo showed 11,013 downloads in the last month. No Inference Provider hosts it, and API pricing is not yet disclosed. The license, training cutoff date and training token count are not stated in the available material.

What this means

The most notable claim is the training stack: a competitive agentic-coding model, by its own numbers, trained entirely on Ascend 910C hardware without NVIDIA GPUs. If the results hold up independently, it shows that domestic Chinese hardware can support frontier-adjacent MoE training, though the 96% throughput gain is measured against an unoptimized baseline, not against a GPU cluster.

For builders, the practical appeal is the 4B active footprint. The Q4_K_M file at 19 GB puts a 256K-context agentic model within reach of a single high-memory consumer GPU or a unified-memory machine. Its self-reported edge over Qwen3.6-35B-A3B is narrow, and it is behind on SWE-bench Verified and Tau3-Bench, so independent evaluation on matching harnesses should come before any switch.

Related Articles

model release

inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter

inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

model release

Cloudflare releases Clef decision models, claims 39 ms median latency vs. 524 ms for TypeSafe's Jev

Cloudflare has released Clef and Clef-flash, two open-weight decision models that return probabilities over predefined answer options instead of generating text. The company claims median latencies of 39 ms and 209 ms, against just over 524 ms for TypeSafe AI's Jev. Both support text and images and are API-compatible with Jev.

model release

Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta

Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.

Comments

Loading...