Back
Hhackernews·sbulaev
52
·4 hr ago·Official API
Not on the current live radar

LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes

View original

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.

AI summary

This research investigates the effectiveness of LLM judges in detecting omissions in AI-generated clinical notes. While these judges perform well in identifying added or altered content (0.79-0.94), their performance significantly drops for omissions (0.50-0.63). Standard designs fail to reliably flag omissions. However, restructuring the task to list facts from the transcript and then check the note for each improves detection. A per-fact pipeline and a GEPA-evolved prompt both achieve this, with the single-call method detecting more omissions at a lower false alarm rate and cost.