RCreddit.com·
暂不在当前实时榜单
Qwen3.8-27B-NVFP4 1M context. So far so good.
A user successfully configured and ran the Qwen3.8-27B-NVFP4 model with a 1M context using vLLM. The setup involved specific environment variables and vLLM serve parameters, including --tensor-parallel-size 4, --gpu-memory-utilization 0.91, --kv-cache-dtype fp8, and --max-model-len 1000000. The configuration also utilized a speculative decoding setup with "method": "mtp" and "num_speculative_tokens": 3, alongside custom hf-overrides for rope_parameters to achieve the extended context.
This report details the specific vLLM parameters and hf-overrides used to achieve a 1M context with Qwen3.8-27B-NVFP4, unlike general discussions that often omit such configurations.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月16日 01:00 UTC
- 收录
- 2026年9月16日 01:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →