HNHacker News·
暂不在当前实时榜单
When LLM judges agree, should we believe them?
A study explored the reliability of LLM judges, particularly when evaluating retrieval-augmented-generation systems. It found that when multiple judge models assess the same passage, their agreement can be leveraged. Using all 10 judge models and maximum training data, dependence-aware results showed 0.912 accuracy on relevance, outperforming weighted majority vote (0.820) and uniform majority vote (0.804). Similar improvements were observed for toxicity (0.792 vs. 0.694/0.695) and summarization (0.806 vs. 0.737/0.561) tasks.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月14日 18:00 UTC
- 收录
- 2026年9月14日 18:00
- 来源类型
- 未分类
本站未收录正文。
前往源站阅读 →