H3-World: Turning Language Understanding into World Control
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
Hugging Face 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。
H3-World 系统通过将角色和摄像机动作组合成文本指令,并通过 MiniMax-H3 的预训练文本路径注入,从而将语言理解转化为世界控制。该系统展现了高效性和泛化能力,仅使用 8,000 个游戏样本、10,000 个 LoRA 步骤和 0.199% 的可训练参数,就实现了可控的角色和摄像机运动,包括未曾见过的动作组合和视觉场景。相关资源包括 ArXiv 论文、GitHub 上的代码、项目页面以及 Hugging Face 模型。
- Language-Native Control: Composes character and camera actions into textual instructions and injects them through MiniMax-H3’s pretrained text pathway.
- Temporally Grounded: Assigns one action prompt to each video latent interval, enabling precise control when actions change over time.
- Efficient & Generalizable: Uses only 8,000 gameplay samples, 10,000 LoRA steps, and 0.199% trainable parameters to achieve controllable character and camera motion, including unseen action compositions and visual scenarios.
📄 ArXiv: https://arxiv.org/abs/2609.01560 💻 Code: https://github.com/Danzer1xxxxChan/H3-World 🏠 Project: https://danzer1xxxxchan.github.io/H3-World/ 🤗 Model: https://huggingface.co/DANNY621/H3-World