返回
RCreddit.com
12
·21小时前·开发者社区 · RSS

ExLlamav3 Recent Updates: CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

查看原文
LlamaQwen

热度趋势

新上榜
最近 24 小时与此前 24 小时对比 · 7 天曲线

百分比基于当前可用热度信号,而非评论数或独立用户人数。

AI 摘要

ExLlamav3 近期获得了大量更新,其中包括 MoE 专家的 CPU 卸载功能,以及 Qwen-3.8-Flash-Next ngram 磁盘卸载。其他改进包括 GLM-5.3-Flash 和一种新的自校准优化技术,同时还有无数其他优化和改进。用户可以在 ExLlama Discord 和 subreddit 上获取更频繁的最新消息。

- CPU offload of MoE experts
 -  Qwen-3.8-Flash-Next  ngram disk offload
 -  GLM-5.3-Flash
 - New  self-calibrated optimization  technique
 - Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt: Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord More frequent news on the exllama sub

ExLlamav3 Recent Updates: CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++ · BuzzRadr