Dear 24G owners, try VLLM you might be able to run Qwen3.8 27B INT4, 144K FP8 KV on RTX 3090 with better speed. (TLDR VLLM AOT)
Users with 24GB VRAM, such as an RTX 3090, are advised to try VLLM to run the Qwen3.8 27B INT4 model with 144K FP8 KV for potentially faster speeds. Benchmarks show an average prefill and prompt processing speed of 871.93 tok/s, with a specific run achieving 1000.26 tok/s for 10K prompts. Even with low inference intensity and INT4 weights, the system scored 71/75 (94.7%) overall across various tasks. Performance might further improve on bare metal systems compared to WSL2.
This report uniquely details how VLLM can enable Qwen3.8 27B INT4 to run on an RTX 3090 with 24GB VRAM, unlike other benchmarks that typically require more powerful hardware.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 14, 2026, 05:01 UTC
- Ingested
- Sep 14, 2026, 05:01
- Source type
- Dev community
Full text isn't available here.
Read at source →