Skip to content
RCreddit.com·
Not on the current live radar

Qwen3.5 0.8B on CPU

AI summary

The Qwen3.5 0.8B model's performance on CPUs was evaluated for local dictation cleanup. A specialized engine, qwen35-cpu, demonstrated significant improvements over llama.cpp running Unsloth Q4_0. Specifically, qwen35-cpu achieved approximately 2.9x prefill, 1.3x single-request decode, and 1.7x batch-16 throughput. This makes the Qwen3.5 0.8B model a viable option for CPU-based tasks, especially when GPUs are occupied with other processes.

Time & source

Ingested
09/08, 08:00 UTC+0
Source type
Dev community
Article

Full text isn't available here.

Read at source →
Source·reddit.com