Skip to content
RCreddit.com·

Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s

AI summary

The uncensored Qwen3.8-27B (HauhauCS) model, originally GGUF-only, has been adapted to run on a 4090 with NInfer. This adaptation achieves 262K context, which is the model's full native window, and approximately 130 tok/s decode speed with MTP3 and 70.8% acceptance. It also boasts a 3,591 tok/s prefill on a 9K prompt, with perplexity within 1.3% of the official artifact. Vision and tool calls remain functional, utilizing E8 4Bit quantization.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Oct 9, 2026, 05:27 UTC

IngestedOffset at this time: UTC+0Oct 9, 2026, 16:00 UTC

Published
Oct 9, 2026, 05:27
Ingested
Oct 9, 2026, 16:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

HauhauCS ships their uncensored Qwen3.8-27B as GGUF only. NInfer, a C++/CUDA wanted its own format.

Now the same model that ran at 91.6 tok/s / 131K under llama.cpp does:

Source·reddit.com