返回
RCreddit.com
16
·13小时前·开发者社区 · RSS

IFM/K2-Horizon-MoVA-36B-A4B-GGUF · Hugging Face

查看原文
Hugging Face模型发布模型访问

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

推荐理由

Hugging Face 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。

AI 摘要

Hugging Face 上的 IFM/K2-Horizon-MoVA-36B-A4B-GGUF 模型在 4B 活跃参数下展现了“前沿级结果”。在代理和推理基准测试中,它超越了开放权重密集模型(大约 30B 模型大小)以及高达其 15 倍大小的 MoE 模型。此外,它还能与封闭式前沿模型有效竞争,具体表现可参见其基准测试结果。该模型预计将提供更多尺寸版本。

more sizes (probably still uploading):

https://huggingface.co/IFM/K2-Horizon-32B-GGUF

https://huggingface.co/IFM/K2-Horizon-7B-GGUF

https://huggingface.co/IFM/K2-Horizon-3.7B-GGUF

https://huggingface.co/IFM/K2-Horizon-0.9B-GGUF

from IFM:

K2-Horizon-MoVA-36B-A4B is the sparse member of the K2-Horizon family: a Mixture-of-Experts model with Mixture-of-Values attention (MoVA) that stores 36B parameters and runs 4B per token. We have released the final checkpoint; intermediate checkpoints, along with the data and the training code, will be released.

K2-Horizon-MoVA-36B-A4B Highlights

- Frontier-class results at 4B active parameters. On agentic and reasoning benchmarks it outscores open weight dense (approximately 30B model size) and MoE models up to 15× its size; and also performs competitively against closed frontier models (see Benchmark Results ).

- 512K context. Native 524,288-token context from the midtraining stages onward.

- Intermediate checkpoints. Intermediate checkpoints will be released so capability changes can be studied across training rather than at a single checkpoint.

- Fully open. Training data/recipe and the training code will be made public.

collection: https://huggingface.co/collections/IFM/k2-horizon