跳到正文
RCreddit.com·
暂不在当前实时榜单

Deepseek training 2T and plans 8T model

AI 摘要

DeepSeek is currently training a 2T-parameter model and has future plans to develop an 8T-parameter model. Their existing models include Flash, with a parameter count of 552 billion, and Pro, which features 1.6T total parameters and 49B activated weights per token. These developments are notable in the context of models like Mythos / Fable, which is estimated to have a 10T parameter count.

为什么是这条

This report details DeepSeek's current 2T model training and future 8T model plans, unlike previous reports that focused on their 552B Flash or 1.6T Pro models.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月21日 16:02 UTC

收录
2026年9月21日 16:02
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com