跳到正文
RCreddit.com·
暂不在当前实时榜单

The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks

AI 摘要

A local agent loop ran for approximately 21 days on a single RTX 3090, tasked with building a CUDA inference engine optimized for its GPU architecture. This process involved 180 subagents, processing around 230M tokens in and out, and 1.7B cache-reads. Compaction consumed about 83 hours across 699 instances, representing 17% of the total calendar time, with typical compactions lasting around 7 minutes for 160k+ token prompts. The project resulted in working kernels and benchmarks, though it did not outperform llama.cpp. The backend used was HyperQwen.

为什么是这条

This report details the token and cache activity of an agent loop, unlike most benchmarks that focus solely on inference speed.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月21日 03:01 UTC

收录
2026年9月21日 03:01
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com