跳到正文
RCreddit.com·
暂不在当前实时榜单

building a humor benchmark for LLMs: someone told me my benchmark's best result was just memory, so i ran his test

AI 摘要

A developer created a humor benchmark for LLMs, which revealed that models excel at explaining why successful jokes work (95%+) but struggle with explaining joke failures (81-92%). This particular tier of the benchmark, consisting of 25 items, was noted as the most challenging component built. The developer is open to discussing the evaluation system further.

为什么是这条

This benchmark uniquely highlights a performance gap, showing LLMs consistently ace explaining successful jokes but struggle with failed ones, unlike other humor evaluations.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月28日 06:00 UTC

收录
2026年9月28日 06:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com