Skip to content
RCreddit.com·
Not on the current live radar

Browser demo of our Clash Royale RL environment: a 5.6k-parameter REINFORCE policy learns defensive placement against a brute-force optimum [P]

AI summary

A browser-based demo showcases a 5.6k-parameter REINFORCE policy learning defensive placements in a Clash Royale RL environment. The developers observed that a strong local optimum exists for 'Giant vs Cannon,' with 5 out of 6 runs getting stuck there using a constant entropy coefficient of 0.01. A linear anneal from 0.1 to 0.005 over 10k tries improved this to 1 out of 6 runs. However, the 'Battle Ram vs Valkyrie' pairing proved challenging, with no setting achieving more than 55% of the optimum.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 28, 2026, 18:00 UTC

Ingested
Sep 28, 2026, 18:00
Source type
Dev community

Discussion trend

→ Steady
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Full text isn't available here.

Read at source →
Source·reddit.com