Skip to content
RCreddit.com·
Not on the current live radar

Dear 24G owners, try VLLM you might be able to run Qwen3.8 27B INT4, 144K FP8 KV on RTX 3090 with better speed. (TLDR VLLM AOT)

AI summary

Users with 24GB VRAM, such as an RTX 3090, are advised to try VLLM to run the Qwen3.8 27B INT4 model with 144K FP8 KV for potentially faster speeds. Benchmarks show an average prefill and prompt processing speed of 871.93 tok/s, with a specific run achieving 1000.26 tok/s for 10K prompts. Even with low inference intensity and INT4 weights, the system scored 71/75 (94.7%) overall across various tasks. Performance might further improve on bare metal systems compared to WSL2.

Why this one

This report uniquely details how VLLM can enable Qwen3.8 27B INT4 to run on an RTX 3090 with 24GB VRAM, unlike other benchmarks that typically require more powerful hardware.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 14, 2026, 05:01 UTC

Ingested
Sep 14, 2026, 05:01
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com