返回
RCreddit.com
18
·14小时前·RSS
暂不在当前实时榜单

model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support by YanissAmz · Pull Request #25444 · ggml-org/llama.cpp

查看原文
LlamaNVIDIA模型发布端侧推理

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

推荐理由

Llama 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。

AI 摘要

A pull request to ggml-org/llama.cpp, #25444, adds support for NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle). This 75B MoE model features a hybrid architecture with interleaved Mamba, MoE, and Attention layers. It supports Multi-Token Prediction (MTP) for faster text generation, similar to Nemotron-3-Super. The Puzzle-75B-A9B reduces parameters from 120.7B total / 12.8B active to 75.3B total / 9.3B active compared to its parent model.