← Home

ByteDance trains a massive AI model in its bid to rival Anthropic

In a move that redraws the global artificial intelligence landscape, ByteDance, the parent company of TikTok, is training a model of up to 10 trillion parameters — a scale that puts it squarely within reach of Anthropic and its most advanced system, Mythos. The report, first published by the Financial Times and picked up by Ars Technica and other outlets, is not merely a milestone of size: it is a statement of ambition, a signal that China's largest tech group has decided to chase the industry's absolute frontier instead of settling for cheaper, competitive models.

The figures help put the bet in perspective. According to three people familiar with the matter, the model is currently in pre-training, a phase that typically takes three to six months, with the exact size to be determined later, before fine-tuning and a potential release. Ten trillion parameters would be more than three times the size of Moonshot's Kimi K3, the largest Chinese model released to date, which sits at around 2.8 trillion. Anthropic does not disclose the dimensions of its systems, but industry estimates place the Mythos 5 near 8 trillion parameters and the Fable 5 at roughly 5 trillion. Even so, parameter count is not an automatic proxy for capability: data quality and training methodology weigh as heavily as raw scale.

The geopolitical backdrop makes the decision even more loaded. For years, Chinese labs built a reputation on efficiency: DeepSeek proved that competing with smaller, cheaper models was possible, while Alibaba's Qwen and Moonshot's Kimi perform well on benchmarks, trailing only Fable 5 in certain tasks. ByteDance, however, is choosing the opposite path: maximum scale, at a moment when the American industry leads on compute cost and access to advanced chips. It is a wager that overtaking Anthropic will require going beyond efficiency and investing in sheer processing power.

The company has concrete advantages in this race. ByteDance has spent the past three years investing in AI more aggressively than any other Chinese tech giant, building out a network of data centers, hiring researchers, and harboring ambitions to develop its own chips. Its Seed team, led by former Google DeepMind scientist Wu Yonghui, numbers around 2,000 people. The Doubao assistant is the most popular AI product in China, with 324 million monthly active users, and SeeDance ranks among the world's most advanced video-generation models. What it lacks is precisely a frontier model that competes at the top with American labs.

There is also a crucial strategic dimension. ByteDance has pursued an independent development policy for more than a year, refusing to distill models from other labs — a common knowledge-compression method that many rivals used to accelerate progress. Founder Zhang Yiming reportedly told his team at a recent internal meeting that only independent development can produce a model that outperforms competitors, urging it to target "world-leading capabilities" without worrying about short-term lag. That choice helps explain why the company has been seen as slower than some peers, but it also positions it for potentially durable leadership.

The open question is whether compute cost — and export restrictions on chips — will allow the plan to materialize. With closed models, ByteDance trades transparency for room to maneuver, yet faces the same geopolitical ceiling as other Chinese labs. If it can train a 10-trillion-parameter model with quality comparable to Mythos, the balance of power in global AI shifts. If it fails, it will be the most expensive demonstration yet that scale alone is not destiny.

Sources: Ars Technica, The Next Web, Techloy, Mint

✓ Independent sources cross-checked and verified before publishing