Mercury released Mercury Decide, I benchmarked it.
A benchmark of Mercury Decide, Jev, Solar Decide, and Kev was conducted using a Korean-focused dataset for Roblox ToS violations. Jev achieved 83.3% accuracy, outperforming Mercury Decide's 66.7%, Solar Decide's 57.8%, and Kev's 72.2%. The benchmark highlighted the importance of minimizing false negatives, as invalid reports lead to reporter bans. Jev had 3 false negatives out of 90, while Mercury Decide had 28, indicating it often rejects reports. Jev was deemed the most stable API decision model.
This benchmark uniquely uses an uncontaminated, Korean-focused dataset for Roblox ToS violations, unlike typical benchmarks, and highlights Jev's superior stability with only 3 false negatives compared to Mercury Decide's 28.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月1日 11:00 UTC
- 收录
- 2026年10月1日 11:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →