Mercury released Mercury Decide, I benchmarked it.
A benchmark of Mercury Decide, Jev, Solar Decide, and Kev was conducted using a Korean-focused dataset for Roblox ToS violations. Jev achieved 83.3% accuracy, outperforming Mercury Decide's 66.7%, Solar Decide's 57.8%, and Kev's 72.2%. The benchmark highlighted the importance of minimizing false negatives, as invalid reports lead to reporter bans. Jev had 3 false negatives out of 90, while Mercury Decide had 28, indicating it often rejects reports. Jev was deemed the most stable API decision model.
This benchmark uniquely uses an uncontaminated, Korean-focused dataset for Roblox ToS violations, unlike typical benchmarks, and highlights Jev's superior stability with only 3 false negatives compared to Mercury Decide's 28.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 1, 2026, 11:00 UTC
- Ingested
- Oct 1, 2026, 11:00
- Source type
- Dev community
Full text isn't available here.
Read at source →