RCreddit.com
12
·18小时前·开发者社区 · RSS
Given how common RTX 3090 use is for LLMs, why don't we see more INT8 W8A8 models?
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
NVIDIA 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。
尽管RTX 3090是LLM爱好者中第二常用的GPU,并且其原生INT8张量核心可以为INT8 W8A8模型提供更好的性能,但目前这类模型的使用率并不高。用户似乎更倾向于使用FP8或更小的量化方法。这引发了一个疑问,即为什么RTX 3090在INT8 W8A8方面所能提供的潜在性能优势没有得到更广泛的利用。
Based on https://huggingface.co/hardware, the RTX 3090 is the second most used GPU by LLM enthusiasts.
Because RTX 3090 has native INT8 tensors cores, it can provide better performance with INT8 W8A8.
However people seems to default to FP8 or smaller quants anyway.
I suppose I am missing information that explains why?