HNHacker News·
Not on the current live radar
When LLM judges agree, should we believe them?
A study explored the reliability of LLM judges, particularly when evaluating retrieval-augmented-generation systems. It found that when multiple judge models assess the same passage, their agreement can be leveraged. Using all 10 judge models and maximum training data, dependence-aware results showed 0.912 accuracy on relevance, outperforming weighted majority vote (0.820) and uniform majority vote (0.804). Similar improvements were observed for toxicity (0.792 vs. 0.694/0.695) and summarization (0.806 vs. 0.737/0.561) tasks.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 14, 2026, 18:00 UTC
- Ingested
- Sep 14, 2026, 18:00
- Source type
- Unclassified
Full text isn't available here.
Read at source →