跳到正文
RCreddit.com·
暂不在当前实时榜单

Dear 24G owners, try VLLM you might be able to run Qwen3.8 27B INT4, 144K FP8 KV on RTX 3090 with better speed. (TLDR VLLM AOT)

AI 摘要

Users with 24GB VRAM, such as an RTX 3090, are advised to try VLLM to run the Qwen3.8 27B INT4 model with 144K FP8 KV for potentially faster speeds. Benchmarks show an average prefill and prompt processing speed of 871.93 tok/s, with a specific run achieving 1000.26 tok/s for 10K prompts. Even with low inference intensity and INT4 weights, the system scored 71/75 (94.7%) overall across various tasks. Performance might further improve on bare metal systems compared to WSL2.

为什么是这条

This report uniquely details how VLLM can enable Qwen3.8 27B INT4 to run on an RTX 3090 with 24GB VRAM, unlike other benchmarks that typically require more powerful hardware.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月14日 05:01 UTC

收录
2026年9月14日 05:01
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com