Skip to content
RCreddit.com·
Not on the current live radar

I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

AI summary

A developer trained Apex-2, a 3.87B MoE model (1.45B active) from scratch using only 86.5B tokens. The model, with a Decoder-only MoE architecture and 32 layers, achieved a HumanEval+ score of 41.5, matching Qwen2.5-1.5B despite significantly less pretrain data. However, it showed limitations in multilingual ability, knowledge, and math, and DPO training negatively impacted its performance.

Why this one

This report highlights that Apex-2 matched Qwen2.5-1.5B's HumanEval+ score using only 86.5B pretrain tokens, unlike Qwen2.5-1.5B's 18T tokens.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 4, 2026, 21:00 UTC

Ingested
Oct 4, 2026, 21:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com