§1Understanding this review
This review looks at how medical AI benchmarks and clinical studies are designed, ranking them by the quality of their evidence rather than model performance.
Published clinical AI metrics rely on a chosen set of tasks, a reference standard, a grader, and a comparator. These choices shape what the resulting scores can, and cannot, show. This review evaluates 14 benchmarks and clinical studies using a fixed rubric, making it easier to assess the strength and limitations of their evidence.
Nothing here scores the models themselves. The grades reflect how rigorously each benchmark or study evaluated them. A benchmark can earn an A even if every model performs poorly. Likewise, a model can rank first on a benchmark graded F, indicating that result is supported by weaker evidence.
§2Benchmarks and studies reviewed
We selected 14 benchmarks and clinical studies that have been cited to support claims about clinical AI capability or safety. All 14 were evaluated using the same rubric and published together. Each entry also identifies the model cohort studied, indicating how current the evaluated systems were.
Scores range from 47 to 94 out of 100, with a median score of 77. Of the 14 benchmarks and studies reviewed, one was classified as tier 1, five as tier 2, three as tier 3, and five as tier 4.
The review sample is not exhaustive, nor is it intended to be. We focused on evaluations cited in product claims, procurement conversations, and press coverage, since those are where a methodological weakness does the most damage downstream. A benchmark or study’s absence from this review should not be interpreted as a judgment of its quality.
Of the 14, 9 are independent of the systems they evaluate. 4 are published by the developer of a model under test and 1 carries a funding or affiliation caveat; each is flagged on its row and explained in its panel.
§3Why include studies and benchmarks
This review considers both benchmarks and studies, and the differences between them affect how their grades should be interpreted.
In the cohort, there are live benchmarks, static benchmarks, and studies and trials. Live benchmarks collect new cases on a schedule and freeze versioned snapshots of the test set, helping ensure that the models have not seen the cases in their training data. Static benchmarks are fixed at publication, making them cheap and easy to reproduce but more likely to degrade over time. Studies and trials involve real cases and real clinicians and cannot be rerun without recruiting clinicians again. Thus, they produce a single set of results rather than an ongoing leaderboard.
The goal here is to determine whether the published evidence supports claims about clinical AI, rather than to identify the best benchmark. Capability benchmarks can establish that models score well, but they may say nothing about whether those results hold when a clinician is involved, the patient is real, or the workflow changes. Trials answer questions that a leaderboard cannot, while benchmarks provide breadth and repeatability that a single trial cannot. Live benchmarks can also show whether results hold on cases the model could not have seen, which static benchmarks cannot. Including all three provides a more complete view of the evidence behind clinical AI claims.
The same six-domain rubric is used to assess every study and benchmark because the same questions apply to all of them: Is the task realistic? Are the labels sound? Is the model used fairly? Is the scoring defensible? Does it hold up under stress? Does it mean anything clinically? When the type of study or benchmark affects how a criterion applies, we clarify that in the relevant write-up rather than adjusting the scale.
Grades are most useful when comparing similar types of evaluations. The same letter grade does not make a clinical trial equivalent to a benchmark with thousands of test questions. It only shows how well each one meets the rubric’s criteria for its particular design.
§4Rubric
Each of the six evaluated domains contains subcriteria with its own point ceiling. A reviewer scores each subcriterion, and the maximum score for each domain is the sum of those ceilings. All domain weights add up to 1.
Task Design
20%Does the eval task genuinely reflect clinical work? Are prompts ecologically valid, are input/output and metric aligned, and is coverage adequate to the claims made?
Data, Labels & Leakage
20%Clinical data realism and provenance, expert annotation and label quality, leakage and contamination controls, transparency and auditability, and scale of coverage.
Model-Use Fidelity
15%Exact model, version and configuration reporting; prompting and harness fidelity; fairness of tool, RAG, multimodal and agentic access; appropriateness of baselines.
Scoring Rigor
15%Judge quality and rubric specificity, reliability and calibration evidence, uncertainty quantification and stochastic stability, and error and failure-mode analysis.
Robustness
15%Robustness checks and sensitivity analyses, empirical generalizability across populations, sites and languages, and explicit edge-case or adversarial probing.
Clinical Validity
15%Clinician comparator quality and clinical anchor, harm, safety, uncertainty and escalation handling, and honesty of the deployment framing.
| Domain | Weight | Subcriteria and point ceilings | Max |
|---|---|---|---|
| TD Task Design | 20% |
|
8 |
| DL Data, Labels & Leakage | 20% |
|
14 |
| MF Model-Use Fidelity | 15% |
|
10 |
| SS Scoring Rigor | 15% |
|
10 |
| RG Robustness | 15% |
|
8 |
| CV Clinical Validity | 15% |
|
9 |
§5How a grade is produced
A domain score normalises its subcriteria to a 10-point scale. The overall score is the weighted sum of the six domain scores. Both are read off a fixed ladder.
Worked example — MedHELM
Task Design carries 3 subcriteria over a domain maximum of 8, giving a domain score of 9.4. Applying the weights across all six domains:
Domain grade ladder
Overall grade ladder
The overall scoring scale uses narrower bands because it combines results from all six domains, allowing it to distinguish between studies or benchmarks that would otherwise receive the same grade. Individual domain scores use broader bands because differences as small as 7.4 versus 7.6 would suggest more precision than a reviewer can reasonably support.
§6Grade caps
A study’s grade is capped if it has any of four major methodological flaws. These flaws call all of its results into question, so points earned elsewhere cannot cancel them out.
A cap is a ceiling rather than a penalty. Those that already score below it are unaffected, which is why even if a cap is recorded, the published grade is unaffected. Every capped entry shows its uncapped score beside the published one.
-
C1
Model identity not reported. Exact model identifiers or versions are not reported for the evaluated systems.Overall capped at B- (8.2)3 of 14 entries
-
C2
Uncalibrated LLM judge. An LLM-as-judge is used without calibration, cross-judge validation, or reliability evidence generated for this study.Scoring Rigor capped at C (6.9)1 of 14 entries
-
C3
No leakage control. Public or static benchmark with no leakage, contamination, or temporal-separation control.Data, Labels & Leakage capped at C (6.9)3 of 14 entries
-
C4
Unsupported clinical claim. Clinical superiority, safety, or deployment readiness is claimed without a fair comparator, outcome evidence, or validated reference standard. Includes deployment claims from synthetic or simulation-only data.Clinical Validity capped at C (6.9)0 of 14 entries
Seven of the 14 entries received a cap, of which three were C1, one was C2, and three were C3. C4 did not apply to any entry. The only cap that changed a reported score was applied to HealthBench Professional’s Scoring Rigor, which fell from 7.0 to 6.9. The other six entries were already below the maximum grade allowed by their caps.
§7Evidence tiers and model cohorts
Each entry also has two labels that do not affect its grade.
The evidence tier describes the type of evidence being evaluated rather than its quality. A tier 1 exam set can be designed well, while a tier 4 clinical trial can be designed poorly. The label helps readers compare similar benchmarks and studies within the appropriate context.
| Tier | Definition | Entries |
|---|---|---|
| 1 |
Exam or public QA benchmark
Questions with a fixed key, drawn from exams or public sources. Cheap to run, but least representative of clinical work. |
1 |
| 2 |
Synthetic, simulated or authored clinical benchmark
Cases written or simulated for the benchmark, usually by clinicians. Realistic in theory, but are not real cases. |
5 |
| 3 |
Retrospective real clinical data or cases
Built on real notes, EHR extracts or consultation records. Real material, no prospective clinician baseline to compare against. |
3 |
| 4 |
Real clinical cases with a clinician comparator
Real cases put in front of clinicians as a reference or control. Strongest evidence, but minimal reproducibility. |
5 |
Model cohort describes how current the evaluated systems are. Capability moves faster than evaluation, so a study or benchmark can be methodologically excellent and still tell you nothing about the model about to be deployed.
§8Reviewer process
Each study and benchmark follows the same six steps. Scores were assigned before the write-ups were drafted to prevent the written explanation from influencing the grade.
-
Independent scoringTwo reviewers read the artifact in full and score every subcriterion on their own, without seeing the other's figures.
-
Joint reviewThe two reviewers then go through the artifact together, comparing scores subcriterion by subcriterion and recording where they differ.
-
ReconciliationConsensus defaults to the mean of the two independent scores. Divergence of 1.5 or more on a domain is flagged and talked through, and reviewers can lock an agreed value over the default. One domain in this corpus crossed that threshold; the agreed value matched the mean.
-
Independent narrativesEach of the two reviewers writes their own prose assessment of every domain, separately.
-
CompilationThe same two reviewers merge their narratives into the single write-up that gets published.
-
Accuracy reviewTwo further people check the compiled narrative against the source manuscript for factual accuracy before it is published.
All grades shown on the dashboard were reached through this consensus process. The final scores are calculated directly from the subcriteria rather than entered manually.
§9What this cannot tell you
The rubric measures how well an evaluation is constructed and reported. The following should be kept in mind when interpreting the grades:
- The rubric does not directly measure model capability. Rather, it measures the robustness of the tools used to review models.
- The benchmark or study is graded as published. If work is sound, scores may still be poor due to a lack of verification means.
- The rubric still depends on reviewer judgment, and two careful reviewers may score the same subcriterion differently. Each domain therefore shows both reviewers’ scores, with larger disagreements flagged rather than hidden by an average.
- Grades are comparable within an artifact type and only loosely across types.
- Scores are a snapshot. Living benchmarks change and preprints become papers, so a grade reflects the version read on the date recorded in the entry.
- The 14 studies and benchmarks were selected, not sampled.
§10Our own conflicts
This review flags conflicts of interest on every entry it grades, so it should declare its own. Members of this author group are authors on 3 artifacts in the corpus.
-
NOHARM
Ethan Goh is a senior author; Anastasia Perez is a co-author.
-
From tool to teammate
Ethan Goh is a co-author.
-
MedHELM
Michael Wornow is a co-author.