ByteDance 10T AI Model 2026: Why It Bans Distillation
Category: Tech Deep Dives
This analysis was written by the aifreetool Editorial Team — a group of full-time AI-industry researchers and writers who verify every claim against primary sources. Last updated August 26, 2026. We keep no affiliate relationship with the companies covered here.
Quick answer: ByteDance is pre-training an AI model with as many as 10 trillion parameters, roughly 3.6 times larger than the biggest Chinese model released so far. More importantly, founder Zhang Yiming has directed the Seed team to stop using model distillation from rivals, choosing a slower, more defensible path to the frontier as Washington turns distillation into a sanctions issue.
On August 7, 2026, the Financial Times reported that ByteDance had begun pre-training the largest model any Chinese company has publicly been associated with. The figure — 10 trillion parameters — grabbed the headlines. But the more consequential detail came a day earlier, when Reuters reported Zhang Yiming had told ByteDance's Seed research team to abandon distillation, the technique of training a smaller model on the outputs of a stronger one, even if that means falling behind domestic competitors in the short term.
The Scale of ByteDance's Bet

If the 10 trillion figure holds, the model would be larger than Moonshot AI's Kimi K3, currently the largest open-weight model ever released at roughly 2.8 trillion parameters, and larger than industry estimates for Alibaba's Qwen3.8-Max at 2.4 trillion. It would also rival estimates for Anthropic's Mythos 5 at around 8 trillion parameters, though Anthropic does not disclose official counts.
The model is reportedly using a Mixture-of-Experts architecture, which means only a fraction of the total parameters activate for any given token. That makes the parameter budget physically trainable in a way a dense 10-trillion-parameter model would not be. The project is being run by ByteDance's Seed team, which now has around 2,000 staff and is led by Wu Yonghui, a former Google DeepMind scientist who joined in February 2025.
The hardware commitment is equally large. South China Morning Post reported ByteDance planned to spend $14 billion on Nvidia AI chips in 2026 as part of a total AI capital expenditure plan exceeding $23 billion, with some internal figures as high as $70 billion. The Financial Times said the training run could require roughly 30,000 GPUs and three to six months of continuous pre-training.
Why the Distillation Ban Matters More Than the Parameter Count

Distillation is the fastest known shortcut for closing capability gaps. It is also the exact practice Washington is now treating as an intellectual-property violation. On July 22, White House science and technology policy director Michael Kratsios accused Moonshot AI of using Anthropic's Fable model to help build Kimi K3. Treasury Secretary Scott Bessent followed up on X: "Open source is not open season on American IP," warning that sanctions and Entity List designations could follow.
ByteDance has reasons to be cautious that Moonshot does not. TikTok's US business is still tied up in national-security politics, and ByteDance retains a 19.9% stake in the TikTok US joint venture. AI2Work's report notes that ByteDance's internal caution around distillation dates back several years, so the policy is not new — but the public posture is. In late July, Zhang reportedly told the Seed team not to worry about short-term leaderboard position and to aim for "world-class model capabilities" over the long run.
The strategic logic is defensive as much as technical. A Chinese lab that can credibly claim to train a frontier model from its own runs is harder to accuse of IP theft and harder to sanction. ByteDance was also notably absent from Anthropic's February 2026 disclosure naming DeepSeek, Moonshot, and MiniMax for industrial-scale distillation campaigns.
Model Scale Comparison

| Model | Total params | Active params | Status |
|---|---|---|---|
| ByteDance (reported ceiling) | Up to 10T | Not disclosed | Pre-training |
| Moonshot Kimi K3 | 2.8T | 104B | Released July 2026 |
| Alibaba Qwen3.8-Max | 2.4T | Not disclosed | Released August 2026 |
| DeepSeek V4-Pro | 1.6T | 49B | Released August 2026 |
| Anthropic Mythos 5 (estimate) | ~8T | Not disclosed | Limited release |
What 10 Trillion Parameters Actually Tells You
Parameter count is a noisy signal. Total parameters measure storage footprint and training cost; active parameters measure inference cost and loosely predict capability per token. DeepSeek V4-Pro has 1.6 trillion total parameters but only 49 billion active. Kimi K3 has 2.8 trillion total and 104 billion active. ByteDance has disclosed neither an active-parameter count nor a final architecture, and the Financial Times reported that 10 trillion is the ceiling under consideration, not a finalized spec.
What the number does reveal is compute-budget ambition. A training run of this scale signals that ByteDance believes the returns to size have not yet flattened and that raw pre-training can still produce emergent capabilities in multi-step reasoning, code generation, and cross-domain scientific inference. It is a bet against the efficiency-first direction many Western labs have taken. TechFastForward's analysis argues that ByteDance is betting raw scale still has room, a wager that resets the competition if it pays off.
FAQ
How large is ByteDance's new model?
The Financial Times reported an upper bound of 10 trillion total parameters, citing three people familiar with the project. The final size has not been fixed.
When will the model be released?
Pre-training typically takes three to six months, putting the earliest plausible release window in late 2026 or early 2027.
What is model distillation?
Distillation trains a smaller model to mimic the outputs of a larger one. It is faster and cheaper than training from scratch but is now being framed by US officials as a potential IP violation.
Why is ByteDance avoiding distillation?
Founder Zhang Yiming reportedly wants the Seed team to build original frontier capability, reducing legal and sanctions risk as Washington targets Chinese labs over the practice.
How does this compare to other Chinese models?
At 10 trillion parameters, the reported ceiling is roughly 3.6 times larger than Moonshot's Kimi K3 and Alibaba's Qwen3.8-Max. Capability will depend on architecture, data quality, and training efficiency, not just parameter count.
My Take / The Bottom Line
The 10 trillion number is attention-grabbing, but the distillation ban is the real strategic move. ByteDance is choosing a slower, more expensive, and legally cleaner route to the frontier at exactly the moment Washington is turning distillation into a trade-policy weapon. If the model works, it gives ByteDance a defensible claim to genuine frontier capability. If it fails, the company will have burned an enormous amount of GPU time and several months of leaderboard position. Either way, the decision tells you where AI competition is heading in late 2026: from a benchmark race into a contest over training provenance, compute sovereignty, and regulatory defensibility.
For a broader view of the models shaping this race, see our AI engine and model directory and our coverage of Doubao, ByteDance's consumer AI assistant.









