ByteDance is developing a new large language model that could reach up to 10 trillion parameters, according to a Financial Times report. The final size of the model is not yet fixed.

The new model is estimated to be similar in scale to Anthropic's Mythos 5 system. Industry sources estimate that Anthropic's Mythos 5 has approximately 8 trillion parameters.

10 trillion parameters

Potential maximum scale of ByteDance's new large language model.

If completed at this scale, the model would be more than three times larger than Moonshot AI's Kimi K3, which has 2.8 trillion parameters. For comparison, Alibaba's Qwen3.8-Max model has 2.4 trillion parameters, while Meituan's LongCat-2.0 and DeepSeek V4-Pro each have approximately 1.6 trillion parameters.

Direct comparisons with U.S. rivals remain difficult because Anthropic and OpenAI have not publicly disclosed the parameter counts for advanced models such as Mythos, Fable, or GPT-5.5. Industry sources estimate Anthropic's Fable 5 model has approximately 5 trillion parameters.

ByteDance's new model is currently in the pre-training phase, a process that typically lasts between three and six months. The company plans to fine-tune the model before deployment.

ByteDance AI expansion and infrastructure

ByteDance has been investing more in AI over past years than other major Chinese tech companies. The company is expanding its data center infrastructure and planning to develop its own AI chips.

The team behind ByteDance's Seed models includes approximately 2,000 employees in China and abroad, led by former Google DeepMind research vice president Wu Yonghui.

ByteDance's recent focus on international AI was primarily on media generation models like Seedance. The Seedance 2.5 video model, released on July 31, capable of generating videos up to 30 seconds long using image, video, audio, and text inputs.

Market position and model accessibility

In China, ByteDance operates Doubao, an AI app used monthly by over 300 million people.

While several Chinese companies have recently unveiled models exceeding 1 trillion parameters, industry analysts note that parameter count alone does not determine performance, as architecture efficiency, training data, and methods are also critical factors.


Reuters was unable to independently verify the Financial Times' report, and ByteDance has not responded to requests for comment.