RCreddit.com·
暂不在当前实时榜单
GGUFs in transformers natively!
Hugging Face's transformers library now natively supports GGUF (llama.cpp quants) models, allowing users to load them with AutoModelForCausalLM.from_pretrained using a gguf_file parameter. This integration aims to provide smaller, quantized models that fit on laptops, offer PyTorch tooling for debugging, and simplify custom generation. While not replacing llama.cpp for maximum local inference performance, it offers a more flexible environment, with performance on Apple Silicon M2 Max showing competitive token generation rates for models like Qwen3.5-4B Q4_K_M.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月23日 09:01 UTC
- 收录
- 2026年9月23日 09:01
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →