跳到正文
HChuggingface.co·

Transformers now runs llama.cpp quants

AI 摘要

Hugging Face Transformers now supports running GGUF models efficiently, allowing users to load checkpoints sized for their laptop's memory using familiar Transformers APIs. This integration enables generating text on personal machines by picking a GGUF from the Hub and loading it with from_pretrained. The underlying kernels can also be integrated into other Transformers models and loading workflows, potentially extending to computer vision, audio, and multimodal models, reusing compatible attention, normalization, and matrix multiplication kernels.

为什么是这条

This integration marks the first time Hugging Face Transformers directly supports GGUF models, unlike previous methods that required external tools or conversions for quantized models.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年9月22日 00:00 UTC

收录当时偏移:UTC+02026年9月22日 12:02 UTC

发布
2026年9月22日 00:00
收录
2026年9月22日 12:02
来源类型
官方发布
档位
当事方
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

讨论趋势

→ 平稳
最近 24 小时与此前 24 小时的快照均值对比 · 7 天曲线

百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。

We're adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop's memory through the familiar transformers APIs. Pick a GGUF from the Hub, load it with from_pretrained, and start generating on your own machine.

Running AI models on your laptop has become much easier, and llama.cpp has been a big part of that. Its inference engine powers local AI tools such as Ollama, LM Studio, and Jan. Alongside projects like MLX , it has helped make local inference a practical option for everyday use.

A recent example of what local AI can feel like:

This is where we are right now. And i’m not gonna lie it feels pretty magical 🧙‍♀️

Qwen3.6 27B running inside of Pi coding agent via Llama.cpp on the MacBook Pro

For non-trivial tasks on the @huggingface codebases, this feels very, very close to hitting the latest Opus in Claude… pic.twitter.com/lsIxLoUneU

— Julien Chaumond (@julien_c) April 24, 2026

GGUF, developed by the llama.cpp team, is a widely used format for local inference. The team also shares quantized checkpoints under ggml-org on the Hub . Publishers such as Unsloth , LM Studio Community , and bartowski also provide ready-to-use GGUF checkpoints in a range of quantizations, so users can pick the version that fits their machine. GGUF models have been downloaded millions of times.

We want to make it easier to run these models locally with transformers, too. Compatibility is only useful if the model is pleasant to run. To bring performance close to llama.cpp, we're reusing its underlying ggml kernels through the kernels library, and reducing overhead in generate. Our initial focus is local inference on Apple Silicon, starting with the Qwen3.5 architecture.

What is the GGUF file format?

GGUF packages model weights and metadata, including tokenizer information and an optional chat template, in one file. It supports different quantization levels, letting you trade some precision for a smaller memory footprint. Variants such as Q4_K_M mix tensor precisions, using mostly 4-bit weights while keeping sensitive tensors at higher precision.

Here's how quantization changes the file size of Unsloth's Qwen3.5-4B :

GGUF variant File size Tradeoff

BF16 8.42 GB Unquantized reference

Q6_K 3.53 GB More precision than the smaller variants

Q5_K_M 3.14 GB A middle ground between size and precision

Q4_K_M 2.74 GB A practical starting point for local inference

We suggest starting with Q4_K_M, then trying Q5_K_M or Q6_K if you have more memory available. More aggressive quantization can help larger models fit, but the quality tradeoff depends on the model and the task. Evaluate it on the work you actually want the model to do. The Hub's GGUF documentation describes the available quantization types.

Load GGUF with transformers

To get started, you need:

- An Apple Silicon Mac.

- A PyTorch version supported by the published ggml-quantization kernel builds, usually the two latest PyTorch releases.

- The latest version of transformers (main for now, until the next release) and a compatible version of kernels.

pip install -U "git+https://github.com/huggingface/transformers.git" kernels