跳到正文
RCreddit.com·

LoRA over GGUF: Train Qwen3.8-Flash-Next in 40G VRAM

AI 摘要

A new method allows training Qwen3.8-Flash-Next using LoRA over GGUF, requiring 40G VRAM. This approach achieves a training speed of 9.5 seconds per iteration for a context chunk size of 2048, translating to 200 tokens per second on Strix Halo. While there is potential for optimization compared to previously achieved speeds of over 1600 tokens per second with PP, this method offers a viable way to train large models with reduced VRAM.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年10月9日 03:54 UTC

收录当时偏移:UTC+02026年10月9日 16:00 UTC

发布
2026年10月9日 03:54
收录
2026年10月9日 16:00
来源类型
开发者社区
档位
社区
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

https://github.com/woct0rdho/transformers5-qwen3.5-recipe

An update to my LoRA over GGUF series: Now we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM, with no CPU offloading, with engram on disk that does not reduce training speed.

On Strix Halo it trains context chunk size 2048 at 9.5 s/it. That's 200 token/s. There is still room to optimize, compared to > 1600 token/s PP we've achieved, and the common sense that LoRA training takes 2-3x work of PP.

Since transformers 5.18, initial support for modern GGUF has been merged, and we can expect more work in this direction.

Spoiler: In the torch-ggml-ops repo there is something called GGTensile. Basically it's Tensile-like asm-level optimization on MMQ kernels. We already see it's faster than HIP in many cases. I'll make a new post when I have something to show on this.

I guess I'll skip DeepSeek-V4.1, unless someone can quantize or prune it to

来源·reddit.com