RCreddit.com·
暂不在当前实时榜单
Browser demo of our Clash Royale RL environment: a 5.6k-parameter REINFORCE policy learns defensive placement against a brute-force optimum [P]
A browser-based demo showcases a 5.6k-parameter REINFORCE policy learning defensive placements in a Clash Royale RL environment. The developers observed that a strong local optimum exists for 'Giant vs Cannon,' with 5 out of 6 runs getting stuck there using a constant entropy coefficient of 0.01. A linear anneal from 0.1 to 0.005 over 10k tries improved this to 1 out of 6 runs. However, the 'Battle Ram vs Valkyrie' pairing proved challenging, with no setting achieving more than 55% of the optimum.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月28日 18:00 UTC
- 收录
- 2026年9月28日 18:00
- 来源类型
- 开发者社区
讨论趋势
→ 平稳
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
本站未收录正文。
前往源站阅读 →