DeepSeek plans to train a model with 80 trillion parameters

Sep 22, 2026 09:51:56

DeepSeek CEO Liang Wenfeng revealed to investors that the company is training a 20 trillion parameter model and plans to develop an 80 trillion parameter model afterward. The current flagship V4-Pro has a total of 1.6 trillion parameters.

Among the publicly disclosed ultra-large models, Moonshot AI's Kimi K3 reaches 2.8 trillion parameters. Kimi K3 uses the MoE architecture, with each token activating 104 billion parameters out of the total 2.8 trillion. The planned 80 trillion parameter model by DeepSeek is nearly three times that size.

DeepSeek completed over 50 billion yuan in financing in July, with a valuation exceeding 50 billion dollars. V4-Pro is open-sourced with a total of 1.6 trillion parameters and 49 billion active parameters.

Recent Fundraising

More
Sep 25
$1.7MSep 24
$37MSep 24

New Tokens

More
Oct 8
Sep 28
PPurserfiPURSER
Sep 23

Latest Updates on 𝕏

More
Sep 24
Sep 24