跳到正文
RCreddit.com·
暂不在当前实时榜单

k_llama.cpp MoE Optimizations: Expert Residency, Hybrid CPU/GPU Execution, Q2_0 Support

AI 摘要

A developer has released a fork of k_llama.cpp with MoE optimizations, including expert residency, hybrid CPU/GPU execution, and Q2_0 support. The fork, available on GitHub, aims to improve performance, especially on systems with limited VRAM. The developer is seeking feedback to potentially merge these optimizations upstream, highlighting features like --moe-resident auto and --moe-resident-mib for managing expert residency.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月6日 08:00 UTC

收录
2026年10月6日 08:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com