RCreddit.com·
Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s
The uncensored Qwen3.8-27B (HauhauCS) model, originally GGUF-only, has been adapted to run on a 4090 with NInfer. This adaptation achieves 262K context, which is the model's full native window, and approximately 130 tok/s decode speed with MTP3 and 70.8% acceptance. It also boasts a 3,591 tok/s prefill on a 9K prompt, with perplexity within 1.3% of the official artifact. Vision and tool calls remain functional, utilizing E8 4Bit quantization.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年10月9日 05:27 UTC
收录当时偏移:UTC+02026年10月9日 16:00 UTC
- 发布
- 2026年10月9日 05:27
- 收录
- 2026年10月9日 16:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
HauhauCS ships their uncensored Qwen3.8-27B as GGUF only. NInfer, a C++/CUDA wanted its own format.
Now the same model that ran at 91.6 tok/s / 131K under llama.cpp does: