Skip to content
RCreddit.com·
Not on the current live radar

Deepseek training 2T and plans 8T model

AI summary

DeepSeek is currently training a 2T-parameter model and has future plans to develop an 8T-parameter model. Their existing models include Flash, with a parameter count of 552 billion, and Pro, which features 1.6T total parameters and 49B activated weights per token. These developments are notable in the context of models like Mythos / Fable, which is estimated to have a 10T parameter count.

Why this one

This report details DeepSeek's current 2T model training and future 8T model plans, unlike previous reports that focused on their 552B Flash or 1.6T Pro models.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 21, 2026, 16:02 UTC

Ingested
Sep 21, 2026, 16:02
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com