DeepSeek plans to train a model with 80 trillion parameters

Sep 22, 2026 09:51:56

DeepSeek CEO Liang Wenfeng revealed to investors that the company is training a 20 trillion parameter model and plans to develop an 80 trillion parameter model afterward. The current flagship V4-Pro has a total of 1.6 trillion parameters.

Among the publicly disclosed ultra-large models, Moonshot AI's Kimi K3 reaches 2.8 trillion parameters. Kimi K3 uses the MoE architecture, with each token activating 104 billion parameters out of the total 2.8 trillion. The planned 80 trillion parameter model by DeepSeek is nearly three times that size.

DeepSeek completed over 50 billion yuan in financing in July, with a valuation exceeding 50 billion dollars. V4-Pro is open-sourced with a total of 1.6 trillion parameters and 49 billion active parameters.

Recent Fundraising

More
$6MSep 22
$100MSep 22
Sep 22

New Tokens

More
Oct 8
Sep 21
FFociFOCI
Sep 21

Latest Updates on 𝕏

More
Sep 21
Sep 21
Sep 21