Skip to content
HNHacker News·
Not on the current live radar

When LLM judges agree, should we believe them?

AI summary

A study explored the reliability of LLM judges, particularly when evaluating retrieval-augmented-generation systems. It found that when multiple judge models assess the same passage, their agreement can be leveraged. Using all 10 judge models and maximum training data, dependence-aware results showed 0.912 accuracy on relevance, outperforming weighted majority vote (0.820) and uniform majority vote (0.804). Similar improvements were observed for toxicity (0.792 vs. 0.694/0.695) and summarization (0.806 vs. 0.737/0.561) tasks.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 14, 2026, 18:00 UTC

Ingested
Sep 14, 2026, 18:00
Source type
Unclassified

Full text isn't available here.

Read at source →
Source·Hacker News·amazon.science