Qwen3.8-27B-NVFP4 1M context. So far so good.
A user successfully configured and ran the Qwen3.8-27B-NVFP4 model with a 1M context using vLLM. The setup involved specific environment variables and vLLM serve parameters, including --tensor-parallel-size 4, --gpu-memory-utilization 0.91, --kv-cache-dtype fp8, and --max-model-len 1000000. The configuration also utilized a speculative decoding setup with "method": "mtp" and "num_speculative_tokens": 3, alongside custom hf-overrides for rope_parameters to achieve the extended context.
This report details the specific vLLM parameters and hf-overrides used to achieve a 1M context with Qwen3.8-27B-NVFP4, unlike general discussions that often omit such configurations.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 16, 2026, 01:00 UTC
- Ingested
- Sep 16, 2026, 01:00
- Source type
- Dev community
Full text isn't available here.
Read at source →