RCreddit.com
17
·1天前·RSS
暂不在当前实时榜单
Qwen3.8-Flash-Next at 170K context on a single 96 GB card. ~110 tok/s.
模型发布
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
Qwen3.8-Flash-Next, utilizing a quantized n-gram to INT4, has demonstrated impressive performance with a 170K context on a single 96 GB GPU, achieving approximately 110 tokens/second. The model, configured with vLLM, successfully handled complex HTML games and improved its gameplay autonomously. Key configurations include --max-model-len of "173400", --gpu-memory-utilization of "0.95", and --enable-chunked-prefill, showcasing its capability for long-horizon tasks and efficient memory management.