跳到正文
HNHacker News·
暂不在当前实时榜单

When LLM judges agree, should we believe them?

AI 摘要

A study explored the reliability of LLM judges, particularly when evaluating retrieval-augmented-generation systems. It found that when multiple judge models assess the same passage, their agreement can be leveraged. Using all 10 judge models and maximum training data, dependence-aware results showed 0.912 accuracy on relevance, outperforming weighted majority vote (0.820) and uniform majority vote (0.804). Similar improvements were observed for toxicity (0.792 vs. 0.694/0.695) and summarization (0.806 vs. 0.737/0.561) tasks.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月14日 18:00 UTC

收录
2026年9月14日 18:00
来源类型
未分类

本站未收录正文。

前往源站阅读 →
来源·Hacker News·amazon.science