Back
RCreddit.com
17
·1 days ago·RSS
Not on the current live radar

Qwen3.8-Flash-Next at 170K context on a single 96 GB card. ~110 tok/s.

View original
Model release

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.

AI summary

Qwen3.8-Flash-Next, utilizing a quantized n-gram to INT4, has demonstrated impressive performance with a 170K context on a single 96 GB GPU, achieving approximately 110 tokens/second. The model, configured with vLLM, successfully handled complex HTML games and improved its gameplay autonomously. Key configurations include --max-model-len of "173400", --gpu-memory-utilization of "0.95", and --enable-chunked-prefill, showcasing its capability for long-horizon tasks and efficient memory management.