返回
RCreddit.com
12
·2天前·RSS
暂不在当前实时榜单

Qwen3.8-Flash-Next optimised for Macs

查看原文
订阅权益

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

AI 摘要

Qwen3.8-Flash-Next has been optimized for Macs, with further optimizations for devices with limited RAM and cache. While disabling MTP (Multi-Tensor Processing) can achieve higher prefill speeds (180-190 tps at 4K context), enabling MTP significantly boosts decode performance by 70%, reaching 22 btps, despite slightly reducing prefill to 170 tps (at 4K) or 150 tps (at 256K context) due to increased RAM usage for KV cache. MTP is generally recommended for its overall performance benefits.