RCreddit.com
17
·1 days ago·RSS
Not on the current live radar
Qwen3.8-Flash-Next at 170K context on a single 96 GB card. ~110 tok/s.
Model release
Heat trend
Collecting trend data
The percentage is based on available heat signal, not comment count or independent people.
Qwen3.8-Flash-Next, utilizing a quantized n-gram to INT4, has demonstrated impressive performance with a 170K context on a single 96 GB GPU, achieving approximately 110 tokens/second. The model, configured with vLLM, successfully handled complex HTML games and improved its gameplay autonomously. Key configurations include --max-model-len of "173400", --gpu-memory-utilization of "0.95", and --enable-chunked-prefill, showcasing its capability for long-horizon tasks and efficient memory management.