RCreddit.com
18
·13 hr ago·RSS
Not on the current live radar
model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support by YanissAmz · Pull Request #25444 · ggml-org/llama.cpp
Heat trend
Collecting trend data
The percentage is based on available heat signal, not comment count or independent people.
Llama model activity is surfacing — worth tracking for capability changes, ecosystem impact, and availability.
A pull request to ggml-org/llama.cpp, #25444, adds support for NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle). This 75B MoE model features a hybrid architecture with interleaved Mamba, MoE, and Attention layers. It supports Multi-Token Prediction (MTP) for faster text generation, similar to Nemotron-3-Super. The Puzzle-75B-A9B reduces parameters from 120.7B total / 12.8B active to 75.3B total / 9.3B active compared to its parent model.