VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

ByteDance Training Model With Up to 10 Trillion Parameters to Rival Anthropic's Mythos

ByteDance is pre-training a model with as many as 10 trillion parameters, aiming to match Anthropic's Mythos system, per sources cited by Ars Technica.

Vector Wire — AI-assisted editorial illustration

ByteDance is training an AI model with as many as 10 trillion parameters, a scale that would approach Anthropic's most advanced Mythos system, according to three people with knowledge of the matter cited by Ars Technica2. The model would be three times larger than Moonshot's Kimi K3, described as the biggest Chinese model released to date.

The effort is at an early stage of pre-training — a phase that typically takes three to six months — before fine-tuning and potential release, one of the people said. The exact model size would only be determined at a later stage.

Chinese-language reporting offers a range of parameter targets. One report describes ByteDance as discussing training a model with over 5 trillion parameters, a scale that may exceed the largest existing model in China4. A separate report frames the target at 5 trillion parameters, with Doubao's intelligence "expected to reach its peak" at the cost of "million-level GPU computing power"1. A third Chinese-language source describes a "50 trillion parameter" target that would exceed both Kimi K3 and Alibaba's Qwen 3.8-Max, and reports that ByteDance founder Zhang Yiming ordered an internal ban on distillation5. The Financial Times separately reported that ByteDance is targeting a model nearing Anthropic's Mythos3.

ANALYSIS The wide spread in reported parameter counts — from 5 trillion to 10 trillion to 50 trillion — suggests either that ByteDance is exploring multiple configurations or that sources are referencing different metrics (dense parameters versus total mixture-of-experts parameters, for instance). The discrepancy warrants caution in treating any single figure as definitive.

Regardless of the final parameter count, the project signals a substantial compute commitment by ByteDance. A model at even the lower end of the reported range would represent a significant step up in scale for Chinese AI labs.

The effort comes shortly after ByteDance released Seedance 2.5 on July 31, which doubled the single-generation video duration of its video-creation model from 15 seconds to 30 seconds ctx. That release demonstrated ByteDance's continued investment across multiple AI modalities.

ANALYSIS The reported distillation ban attributed to Zhang Yiming, if accurate, would mark a notable internal policy shift. Distillation — training smaller models on the outputs of larger ones — has been a common technique across the industry, and restricting it internally could indicate a strategic emphasis on proprietary model development at full scale rather than efficiency shortcuts.

CORRECTIONS: none for this article · this piece updates automatically as the story develops · corrections policy & trail →