Skip to content
RCreddit.com·
Not on the current live radar

The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks

AI summary

A local agent loop ran for approximately 21 days on a single RTX 3090, tasked with building a CUDA inference engine optimized for its GPU architecture. This process involved 180 subagents, processing around 230M tokens in and out, and 1.7B cache-reads. Compaction consumed about 83 hours across 699 instances, representing 17% of the total calendar time, with typical compactions lasting around 7 minutes for 160k+ token prompts. The project resulted in working kernels and benchmarks, though it did not outperform llama.cpp. The backend used was HyperQwen.

Why this one

This report details the token and cache activity of an agent loop, unlike most benchmarks that focus solely on inference speed.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 21, 2026, 03:01 UTC

Ingested
Sep 21, 2026, 03:01
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com