What Sante's 83.83 on DiagnosisArena-MCQ actually measures [D]
Sante's 83.83 score on DiagnosisArena-MCQ measures its performance on multiple-choice questions where alternatives are supplied. This score, combined with results from MedXpertQA-Text (53.88) and HealthBench Professional (45.73), provides a broader evaluation profile for Sante in medical text tasks. HealthBench Professional assesses open-ended clinical chat using physician-written rubrics, and its score is not a percentage accuracy. The Sante chart lacks detail to determine if the HealthBench Professional value is length-adjusted.
Time & source
- Published
- 09/09, 13:01 UTC+0
- Ingested
- 09/09, 19:00 UTC+0
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
https://preview.redd.it/xv2epabu6ioh1.png?width=1171&format=png&auto=webp&s=4e22c69855ec509bd63a038a24f765caeaafa59c
Ant Ling reports 83.83 on DiagnosisArena-MCQ for Ling-3.0-flash-Sante, its new medical reasoning model. The suffix matters: the task provides case information, examinations and tests, then asks the model to choose from four diagnoses.
That result tells us about selecting an answer when the candidate set and case evidence are supplied. It does not establish how the same model would generate an unrestricted differential, decide what history is missing, or choose which investigation to request next. Those would require different evaluations.
Evaluation Sante result What the task adds MedXpertQA-Text 53.88 Challenging medical questions in a text subset. HealthBench Professional 45.73 Open-ended professional clinical chat, assessed with physician-written rubrics. The published HealthBench Professional definition includes care consultation, writing/documentation and medical research. Its score is not percentage accuracy. The Sante chart does not provide enough scoring detail to identify the reported value as length-adjusted or unadjusted, so a comparison with another published HBP result would need that checked first.
This is why the three results are useful together. They give Sante a broader medical-text evaluation profile than an exam score alone, while leaving specific questions open. For a case-answering application, the first decision is whether users supply the alternatives or expect the model to construct them. The release supports including Sante in that evaluation; the 83.83 figure applies to the supplied-options version.