What Sante's 83.83 on DiagnosisArena-MCQ actually measures [D]
Sante在DiagnosisArena-MCQ上的83.83分衡量了其在提供备选项的多项选择题上的表现。这个分数,结合MedXpertQA-Text(53.88)和HealthBench Professional(45.73)的结果,为Sante在医学文本任务中提供了更广泛的评估概况。HealthBench Professional使用医生编写的评分标准评估开放式临床聊天,其分数并非百分比准确率。Sante图表缺乏足够细节来确定HealthBench Professional的值是否经过长度调整。
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月9日 13:01 UTC
收录当时偏移:UTC+02026年9月9日 19:00 UTC
- 发布
- 2026年9月9日 13:01
- 收录
- 2026年9月9日 19:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
https://preview.redd.it/xv2epabu6ioh1.png?width=1171&format=png&auto=webp&s=4e22c69855ec509bd63a038a24f765caeaafa59c
Ant Ling reports 83.83 on DiagnosisArena-MCQ for Ling-3.0-flash-Sante, its new medical reasoning model. The suffix matters: the task provides case information, examinations and tests, then asks the model to choose from four diagnoses.
That result tells us about selecting an answer when the candidate set and case evidence are supplied. It does not establish how the same model would generate an unrestricted differential, decide what history is missing, or choose which investigation to request next. Those would require different evaluations.
Evaluation Sante result What the task adds MedXpertQA-Text 53.88 Challenging medical questions in a text subset. HealthBench Professional 45.73 Open-ended professional clinical chat, assessed with physician-written rubrics. The published HealthBench Professional definition includes care consultation, writing/documentation and medical research. Its score is not percentage accuracy. The Sante chart does not provide enough scoring detail to identify the reported value as length-adjusted or unadjusted, so a comparison with another published HBP result would need that checked first.
This is why the three results are useful together. They give Sante a broader medical-text evaluation profile than an exam score alone, while leaving specific questions open. For a case-answering application, the first decision is whether users supply the alternatives or expect the model to construct them. The release supports including Sante in that evaluation; the 83.83 figure applies to the supplied-options version.