Browser demo of our Clash Royale RL environment: a 5.6k-parameter REINFORCE policy learns defensive placement against a brute-force optimum [P]
A browser-based demo showcases a 5.6k-parameter REINFORCE policy learning defensive placements in a Clash Royale RL environment. The developers observed that a strong local optimum exists for 'Giant vs Cannon,' with 5 out of 6 runs getting stuck there using a constant entropy coefficient of 0.01. A linear anneal from 0.1 to 0.005 over 10k tries improved this to 1 out of 6 runs. However, the 'Battle Ram vs Valkyrie' pairing proved challenging, with no setting achieving more than 55% of the optimum.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 28, 2026, 18:00 UTC
- Ingested
- Sep 28, 2026, 18:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →