跳到正文
RCreddit.com·
暂不在当前实时榜单

Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine

AI 摘要

Researchers developed Kyojin, an engine built on ExLlamaV3 for AMD Strix Halo (gfx1151, ROCm), enabling two 300B-class Mixture-of-Experts (MoE) models to run on a single 128 GB mini PC. The GLM-5.3-Flash model achieves approximately 580 tok/s prefill, while the MiMo-V2.6-Flash model reaches up to 44 tok/s decode. These models utilize EXL3 weights, with MiMo being a custom quantization and GLM a mix of public turboderp tensors and custom tuning.

为什么是这条

This is the first time two 300B-class MoE models have been shown running on a single 128 GB mini PC, unlike previous demonstrations on more powerful hardware.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月4日 00:00 UTC

收录
2026年10月4日 00:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com