RCreddit.com·
暂不在当前实时榜单
building a humor benchmark for LLMs: someone told me my benchmark's best result was just memory, so i ran his test
A developer created a humor benchmark for LLMs, which revealed that models excel at explaining why successful jokes work (95%+) but struggle with explaining joke failures (81-92%). This particular tier of the benchmark, consisting of 25 items, was noted as the most challenging component built. The developer is open to discussing the evaluation system further.
This benchmark uniquely highlights a performance gap, showing LLMs consistently ace explaining successful jokes but struggle with failed ones, unlike other humor evaluations.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月28日 06:00 UTC
- 收录
- 2026年9月28日 06:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →