DeepSeek plans to train a model with 80 trillion parameters
Sep 22, 2026 09:51:56
DeepSeek CEO Liang Wenfeng revealed to investors that the company is training a 20 trillion parameter model and plans to develop an 80 trillion parameter model afterward. The current flagship V4-Pro has a total of 1.6 trillion parameters.
Among the publicly disclosed ultra-large models, Moonshot AI's Kimi K3 reaches 2.8 trillion parameters. Kimi K3 uses the MoE architecture, with each token activating 104 billion parameters out of the total 2.8 trillion. The planned 80 trillion parameter model by DeepSeek is nearly three times that size.
DeepSeek completed over 50 billion yuan in financing in July, with a valuation exceeding 50 billion dollars. V4-Pro is open-sourced with a total of 1.6 trillion parameters and 49 billion active parameters.