RCreddit.com
13
·1天前·开发者社区 · RSS
Are models with N-Gram tables going to completely change the AI race?
Qwen模型发布
热度趋势
↓ 降温 11%
百分比基于当前可用热度信号,而非评论数或独立用户人数。
Qwen 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。
Qwen 3.8 Flash Next 模型中引入的 N-Gram 表技术,可能会彻底改变人工智能模型的部署方式。这项技术有望使万亿参数模型在单个服务器上运行,仅需适度的 GPU 和大量的系统内存,而不再需要通过 NVlink 连接的庞大 GPU 服务器机架。这一进展引发了人们的思考,即它是否能比预期更快地缩小自托管模型与旗舰模型之间的能力差距。
The news about Qwen 3.8 Flash Next is the first I'm reading about n-gram tables. I may be completely misunderstanding how they work but it seems they could open the door for 1T+ parameter models to be run on a single server with modest GPUs and a ton of system RAM rather than needing a rack of GPU servers connected with something like NVlink.
Could we be looking at shrinking the capability gap between self hosted and flagship models faster than we thought, or am I way off base?