返回
Hhackernews·sbulaev
52
·4小时前·官方 API
暂不在当前实时榜单

LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes

查看原文

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

AI 摘要

This research investigates the effectiveness of LLM judges in detecting omissions in AI-generated clinical notes. While these judges perform well in identifying added or altered content (0.79-0.94), their performance significantly drops for omissions (0.50-0.63). Standard designs fail to reliably flag omissions. However, restructuring the task to list facts from the transcript and then check the note for each improves detection. A per-fact pipeline and a GEPA-evolved prompt both achieve this, with the single-call method detecting more omissions at a lower false alarm rate and cost.