The Information Machine
Concluded·following since 7 Aug 2026·Day 4·3 sources·updated 10 Aug 2026

ByteDance's frontier AI pre-training effort

The gist

ByteDance is pre-training a 10-trillion-parameter AI model, aiming to rival Anthropic

A successful run would show ByteDance can execute frontier-scale pretraining without relying on a rival model as a teacher. The model at up to 10 trillion parameters would be three times larger than the biggest Chinese model released to date.

The full picture

ByteDance is in the pre-training stage of an AI model with up to 10 trillion parameters, according to the Financial Times as reported by multiple outlets. The model is roughly three times larger than Moonshot's Kimi K3 at 2.8 trillion parameters, described as the biggest Chinese model released to date. The effort is aimed at approaching the scale of Anthropic's Mythos system. Pre-training typically takes three to six months before fine-tuning and potential release, and the exact final model size will be determined later. ByteDance has avoided distilling rival models for more than a year, preferring independent model development. The model uses a Mixture-of-Experts architecture; given collapsing sparsity ratios across recent Chinese models, active parameters are estimated at 200 to 500 billion, making the 10 trillion figure primarily a memory cost rather than a compute cost. ByteDance has a contract for approximately 36,000 Blackwell GPUs worth $2.5 billion through a Malaysian cloud operator, described as compliant with US export controls.

How it developed
10 August 2026

The Neuron Daily on August 9 corroborated that ByteDance's pre-training effort approaches the scale of Anthropic's Mythos system, adding no new details.

The established account holds: ByteDance is independently pre-training a model with up to 10 trillion parameters, roughly three times the size of Kimi K3, backed by a contract for approximately 36,000 Blackwell GPUs worth $2.5 billion through a Malaysian cloud operator.

First citedThe Neuron
9 August 2026

The Neuron Daily corroborates the reported 10-trillion-parameter pre-training effort at a scale near Anthropic's Mythos

Technical analysis published August 8 added a compute breakdown to the Financial Times report that ByteDance is pre-training a model with up to 10 trillion parameters: given the Mixture-of-Experts architecture, active parameters are estimated at 200 to 500 billion, making the headline count primarily a memory cost. The analysis also noted ByteDance holds a contract for roughly 36,000 Blackwell GPUs through a Malaysian cloud operator; a quoted source said inference is the larger challenge, with distilled smaller models the likely serving path.

8 August 2026

Compute analysis estimates active parameters at 200-500 billion and notes ByteDance's 36,000 Blackwell GPU contract through a Malaysian cloud operator

The Financial Times reported August 7 that ByteDance is pre-training a model with up to 10 trillion parameters, roughly three times larger than Kimi K3 and aimed at approaching Anthropic's top models. Technical analysis published August 8 found that the Mixture-of-Experts architecture puts active parameters at 200 to 500 billion, making the headline figure primarily a memory cost; ByteDance holds a contract for about 36,000 Blackwell GPUs through a Malaysian cloud operator, and a source said inference is the larger challenge.

7 August 2026

Financial Times reports ByteDance is pre-training a model with up to 10 trillion parameters, aiming to rival Anthropic's Mythos

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free