Back
RCreddit.com
12
·17 hr ago·Dev community · RSS

Given how common RTX 3090 use is for LLMs, why don't we see more INT8 W8A8 models?

View original
NVIDIAModel releaseOn-device

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.

Why it matters

NVIDIA model activity is surfacing — worth tracking for capability changes, ecosystem impact, and availability.

AI summary

Despite the RTX 3090 being a popular GPU among LLM enthusiasts, and its native INT8 tensor cores offering better performance with INT8 W8A8 models, there's a noticeable lack of such models. Users often opt for FP8 or smaller quantization methods instead. This raises questions about why the potential performance benefits of INT8 W8A8 on the RTX 3090 are not being fully utilized by the community.

Based on https://huggingface.co/hardware, the RTX 3090 is the second most used GPU by LLM enthusiasts.

Because RTX 3090 has native INT8 tensors cores, it can provide better performance with INT8 W8A8.

However people seems to default to FP8 or smaller quants anyway.

I suppose I am missing information that explains why?

Given how common RTX 3090 use is for LLMs, why don't we see more INT8 W8A8 models? · BuzzRadr