Back
Ssignal
17
·7 days ago·1 signals
Archived topic · source no longer tracked

3 days benchmarking most llama.cpp flags on my weird 40gb vram laptop + tb4 egpu setup. Got +70% generation, +40% prefill, 60k more context, and filed a bug in llama around MTP. What I learned.

LlamaNVIDIAOn-device

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.