跳到正文
RCreddit.com·
暂不在当前实时榜单

I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

AI 摘要

A developer trained Apex-2, a 3.87B MoE model (1.45B active) from scratch using only 86.5B tokens. The model, with a Decoder-only MoE architecture and 32 layers, achieved a HumanEval+ score of 41.5, matching Qwen2.5-1.5B despite significantly less pretrain data. However, it showed limitations in multilingual ability, knowledge, and math, and DPO training negatively impacted its performance.

为什么是这条

This report highlights that Apex-2 matched Qwen2.5-1.5B's HumanEval+ score using only 86.5B pretrain tokens, unlike Qwen2.5-1.5B's 18T tokens.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月4日 21:00 UTC

收录
2026年10月4日 21:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com