RCreddit.com·
暂不在当前实时榜单
Microsoft trained a 4B coding agent almost entirely with Reinforcement Learning, without a bigger teacher
Microsoft Research Montréal, in collaboration with Mila and UC San Diego, developed FrogNano, a 4-billion-parameter coding agent. This model, based on Qwen3.5-4B, was trained almost entirely using reinforcement learning without human-labeled data or a larger "teacher" model. Its learning involved synthetic software engineering tasks generated and refined across approximately 1,500 environments through repeated rounds of task creation and RL, as detailed in arXiv:2609.07925 [cs.AI].
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月10日 17:00 UTC
- 收录
- 2026年9月10日 17:00
- 来源类型
- 开发者社区
正文
本站未收录正文。
前往源站阅读 →