The Chinese tech giant ByteDance is training a multi-trillion parameter AI model to compete with the scale of Anthropic’s Claude Mythos, according to the Financial Times.
Citing people familiar with the matter, the publication reported that ByteDance is in the pre-training stage for a 10 trillion parameter large language model (LLM). Model size is not a direct indicator of model performance but does indicate the scale and depth of data used to train it.
Insiders told the FT that ByteDance intends to build a model on the scale of Claude Mythos 5, which is believed to have eight to 10 trillion parameters. If would make it more than three times the size of Moonshot AI’s Kimi K3, Alibaba’s Qwen 3.8-Max, and almost 35 times the size of DeepSeek V4 Flash.
The parent company behind Tiktok and its Chinese counterpart Douyin, ByteDance is one of China’s biggest tech companies and has vast cloud resources at its disposal.
In May, Bloomberg reported the firm is mulling capital expenditures of up to $70 billion in 2026, as it seeks to build more AI infrastructure including data centres. Sources close to the matter told the publication that ByteDance could spend as much as $100 billion next year pending favourable economic conditions.
To date ByteDance has not released LLMs on the level of Western labs such as Anthropic, Google, or OpenAI nor those of other Chinese firms including Alibaba, DeepSeek, MiniMax, Moonshot AI, and Z.ai.
In the AI community it is best known for Seedream and Seedance, its text-to-image and text-to-video models considered cutting edge for their model category, as well as the chatbot Doubao.
But the tech giant’s AI team, led by former Google DeepMind Wu Yonghui and which the FT reported stands at around 2,000 employees, is working towards establishing itself as a competitive player
On Thursday, Reuters shared reports by Chinese publication The Paper stating that ByteDance founder Zhang Yiming had told staff to avoid model distillation, a practice in which smaller LLMs are trained on the outputs of more capable models.
The US government and Anthropic have formally accused firms such as Moonshot AI, as well as DeepSeek and Minimax, of distilling Claude to boost their own models in post-training runs.
Citing sources inside the firm, The Paper said Zhang’s memo urged staff to pursue “long-termism and delayed gratification, rather than using others’ output to achieve short-term leaderboard rankings”.


