Skip to content
RCreddit.com·
Not on the current live radar

building a humor benchmark for LLMs: someone told me my benchmark's best result was just memory, so i ran his test

AI summary

A developer created a humor benchmark for LLMs, which revealed that models excel at explaining why successful jokes work (95%+) but struggle with explaining joke failures (81-92%). This particular tier of the benchmark, consisting of 25 items, was noted as the most challenging component built. The developer is open to discussing the evaluation system further.

Why this one

This benchmark uniquely highlights a performance gap, showing LLMs consistently ace explaining successful jokes but struggle with failed ones, unlike other humor evaluations.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 28, 2026, 06:00 UTC

Ingested
Sep 28, 2026, 06:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com