analysisByteDance

Report: ByteDance Training 10-Trillion-Parameter Model, Largest in China

TL;DR

The Financial Times reports ByteDance is pretraining an AI model with up to 10 trillion parameters, which would make it three times larger than Moonshot's Kimi K3, currently China's largest model. The claim comes from anonymous insiders and has not been confirmed by ByteDance.

2 min read
0

The Claim

ByteDance is training an AI model with up to ten trillion parameters, according to a Financial Times report citing three unnamed insiders. If accurate, this would make it three times larger than Moonshot's Kimi K3, currently regarded as China's largest AI model.

Neither ByteDance nor the FT's sources have publicly confirmed a parameter count, model name, or release timeline. Parameter counts for frontier models are rarely disclosed by their developers and are typically estimated by industry analysts or leaked through internal sources.

Scale Comparison

At ten trillion parameters, the model would place ByteDance in a similar range to Anthropic's flagship system, which industry estimates place at roughly eight trillion parameters. Anthropic has not disclosed parameter counts for any of its models, so this figure remains an unverified estimate rather than a confirmed number.

According to the FT's sources, the model is currently in pretraining, a phase that typically lasts three to six months before moving into fine-tuning and evaluation. That timeline suggests any public release, if one follows, is unlikely before late 2026 at the earliest.

Training Approach

One source told the FT that ByteDance has avoided distillation, training on outputs generated by other companies' models, for more than a year. This claim, if accurate, would distinguish the effort from allegations previously made against other Chinese AI labs accused of distilling outputs from Western models to accelerate training.

Parameter count alone does not determine model quality. Data quality, training methodology, and architecture choices all factor heavily into final performance, and none of these have been detailed in the report.

Internal Direction

ByteDance founder Zhang Yiming reportedly told the company's roughly 2,000-person Seed research team to target world-leading model capabilities over the long term. This internal directive was described by sources but has not been independently verified through official ByteDance statements.

Broader Context

ByteDance is not alone in pursuing models at this scale. xAI is reportedly training Grok variants with six and ten trillion parameters on its Colossus 2 supercomputing cluster, according to statements from Elon Musk. Neither xAI's nor ByteDance's figures have been independently confirmed through technical documentation or benchmark disclosures.

What This Means

This report describes a model still in early training, not a finished product. No benchmark scores, pricing, context window, or release date exist yet, because none have been announced. The parameter counts cited come from anonymous sources rather than company disclosures, and historical estimates of "largest model" claims have often proven difficult to verify after release.

What is notable is the trajectory: multiple labs, including ByteDance and xAI, appear to be pursuing trillion-scale parameter counts well beyond current publicly confirmed frontier models. Whether raw parameter count translates into meaningfully better real-world performance remains an open question that only becomes answerable once benchmark results and independent evaluations are available. Until ByteDance confirms details or ships a product, this remains a claim about work in progress, not a verified model release.

Related Articles

analysis

Chinese Models Kimi K3 and GLM-5.3 Close In on GPT-5.5 and Claude Opus 5, New Analysis Finds

A new industry analysis argues the performance gap between Chinese and Western AI models has narrowed to single-digit differences on broad benchmarks. Moonshot's Kimi K3 and Zhipu's GLM-5.3 now trail OpenAI and Anthropic's top models by only a few points on the Artificial Analysis Intelligence Index, with a clear Western edge remaining only in abstract reasoning, output reliability, and offensive cybersecurity capability.

analysis

Xiaomi Lists Three New MiMo-V2.6 Models on OpenRouter: Pro, Flash, and Pro-UltraSpeed

Xiaomi has added three new entries to its MiMo model family on OpenRouter: MiMo-V2.6-Pro, MiMo-V2.6-Flash, and MiMo-V2.6-Pro-UltraSpeed. Full specifications, pricing, and benchmark results have not yet been disclosed.

analysis

Chinese Open-Weight Models Now Lead US Rivals by 2-6 Months, Congressional Briefing Shows

AI researcher Nathan Lambert's prepared testimony to Congress details how Chinese open-weight models have overtaken American ones on both downloads and capability benchmarks since mid-2025. The gap has widened to roughly 1.6 billion additional Hugging Face downloads and a near-double-digit lead on the Artificial Analysis Intelligence Index.

analysis

Chinese Open-Weight Models Now Lead US Rivals by Wide Margin, Interconnects Analysis Finds

A briefing prepared for Congress by AI researcher Nathan Lambert details how Chinese open-weight models have overtaken American counterparts since mid-2025, with a nearly 2x lead in Hugging Face downloads and a 19-22 point gap on the Artificial Analysis Intelligence Index.

Comments

Loading...