RCreddit.com·
Not on the current live radar
building a humor benchmark for LLMs: someone told me my benchmark's best result was just memory, so i ran his test
A developer created a humor benchmark for LLMs, which revealed that models excel at explaining why successful jokes work (95%+) but struggle with explaining joke failures (81-92%). This particular tier of the benchmark, consisting of 25 items, was noted as the most challenging component built. The developer is open to discussing the evaluation system further.
This benchmark uniquely highlights a performance gap, showing LLMs consistently ace explaining successful jokes but struggle with failed ones, unlike other humor evaluations.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 28, 2026, 06:00 UTC
- Ingested
- Sep 28, 2026, 06:00
- Source type
- Dev community
Full text isn't available here.
Read at source →