State of Healthcare AI Benchmarks

Exams are saturated. Usage has exploded. The evaluation that answers "does this help?" happens in private, at product-analytics cadence.
Compiled August 2026 · from published benchmarks, peer-reviewed gap studies, and vendor disclosures · interactive — best viewed in a browser
01 · TL;DR

Five things true at once

Part I

The record

What public evaluation actually is right now: the flagship instruments, and the five eras of question they were built to answer.
02 · The landscape

The flagship evaluations, mid-2026

BenchmarkStewardWhat it gradesDataCadence
USMLE / MedQA eraMedical knowledge recall, multiple choicethe 2023–24 generation; now table stakesExam itemsStatic, saturated
HealthBench May 2025OpenAIMulti-turn health conversations against physician-written rubrics5,000 conversations · 262 physicians · 48,562 criteriaSimulatedStatic + "Hard" subset
MedHELM 2025Stanford + partners121 clinician-validated tasks across 5 categories, 35 benchmarksdecision support · documentation · communication · research · adminReal EHR data (incl. private/gated)Quarterly leaderboard
ARISE · MAST 2026Stanford Medicine networkComposite "living benchmark" of clinical benchmarksNOHARM, MedAgentBench, CPC-Bench, PhysicianBench …Mixed clinicalRolling refresh
HealthBench Professional Apr 2026OpenAIThe tasks clinicians actually bring to a chatbot at workcurated from 15,079 real clinician conversationsReal usage, frozenStatic
RCTs rareAcademiaWhole-workflow effect on real decisions and outcomese.g. Abaluck et al. 2026: assistance changed deliberation, not test appropriatenessReal deploymentsYears per verdict
03 · The eras

Five eras, each asking a more real question

Eras stack; they don't replace each other. Each asks a more real question: item → conversation → task → usage → outcome. The last bar is dashed: its public instruments barely exist.
20192020202120222023202420252026
TODAY
The exam era — can it recall medicine?multiple-choice knowledge · saturated at 84–90%, at/above physician level
MedQA '20
MedMCQA '22
Med-PaLM passes '22
GPT-4 aces USMLE '23
ends as marketing:
"100% on USMLE"
The rubric era — can it converse safely?physician-written rubrics over multi-turn chats
Med-PaLM long-form '23
HealthBench May '25
length-adjusted scoring '26
The task era — can it do clinical work?real EHR data · agentic workflows
MedHELM '25 · 121 tasks
MedAgentBench '25 · 70→92% in 6mo
ARISE · MAST '26
The usage erawhat do clinicians actually ask? benchmarks curated from real use
OpenEvidence 1M/day '26
HealthBench Pro Apr '26 · 15k real chats
usage ran ahead of evaluation:
~⅔ of US physicians before any
benchmark measured their questions
The outcome eradid it change what happened to the patient?
Abaluck et al. · ARISE calls for post-deployment eval '26
the era the field is asking for — its instruments
(funnels, decision streams, reconciliation)
already run daily, on the private side of the wall
Part II

The ceiling

Why those instruments stopped discriminating — and why the field's response, moving up into more realism, is not the same as moving right into a faster loop.
04 · Saturation

Exams stopped discriminating

84–90%
frontier accuracy on USMLE-style exams — at or above average physician performance
JMIR systematic review, Dec 2025
"100%"
a perfect USMLE score, used as consumer marketing by a clinical AI vendor
capability score as product claim
70% → 92%
best-model success on MedAgentBench in ~6 months
ARISE / MAST tracking
length-adjusted
HealthBench scoring revised after verbose answers were found to game rubric coverage
the benchmark patching its own incentive bug
05 · The gap

Knowledge ≠ practice, now quantified

−39 to −45 pts
drop from exam-style to practice-based performance across 39 medical benchmarks
JMIR systematic review, Dec 2025
95% → 34%
one system's fall from evaluation to deployment conditions
Bean et al., via CMU 2026
40–50%
accuracy on safety-critical scenarios — the worst tier of the degradation ladder
factual 85–93 · reasoning 50–60 · diagnosis 45–55 · safety 40–50
Δ = assumptions
the gap traced to implicit protocol assumptions violated at deployment: task structure, who interacts, how outputs become decisions
CMU, 2026
06 · Meanwhile

Usage didn't wait for the evaluations

1M / day
physician consultations on OpenEvidence in a single 24h period (Mar 2026)
verified-NPI clinicians
~27M / mo
clinical encounters touched by one platform (Apr 2026) — ~⅔ of US physicians
the largest evaluative dataset in medicine, unpublished
47% → 63%
physicians using AI daily, early 2025 → early 2026
Doximity, n = 3,151 across 15 specialties
~60%
of queries are patient-specific decision questions — the unit of work, not the exam item
OpenEvidence disclosure
07 · Direction of travel

Benchmarks are drifting up, not right

Part III

Behind the wall

The evaluation that answers does this help? already exists and runs daily. It sits on the other side of a property line, and nothing it learns comes back.
08 · The chasm

Evaluation is split by a public/private wall

Public — published & citable · verdicts · frozen · publication cadence Private — proprietary telemetry · process · adaptive · live denominators
Left of the wall Benchmarks, studies, surveys. Optimized for generality and citation. Can only use data that can be published — so the patient is a curated vignette and the verdict is "useful or not," delivered on publication cadence.
Right of the wall Decision streams, funnels, workload-adjusted utilization, outcome reconciliation. Sits on the real data, updates daily, and improves the product — but optimizes for one deployment's next sprint and never publishes.
09 · The visit clock

Put every effort on one timeline and the coverage gap is visible

One clock, anchored on the visit. Private band: the whole arc, continuous. Public band: frozen slices. Amber notes: the blind spot at each layer.
Usage analytics — observed liveproduct analytics + app DB
Post-usage — reconciledEHR evidence + reconciliation, batch
−7d
−5d
−3d
−1d
VISIT
+1d
+3d
+7d
+14d
+30d
Private — product telemetry · continuous
1Delivery
generated → publishedapp DB — the tracker never sees it
no consumer analogue: the product pushes — generation can stop for one clinic while usage charts still look alive
2Exposure
queue → viewed pre-visit or at the visittracker views ∪ app-DB opens
"delivered" isn't reach — one deployment's accept rate went 12%→44% when the denominator became viewed
3Engagement
the session itselfopen time · chat turns · patients touched
DAU measures the clinic schedule — normalize per appointment-provider-day
4Action
accept / rejectapp DB, deduplicated · some land days later
the decisive click is where the SDK isn't — and vendor staff use the same UI: know whose click it was
5Trust
thumbs → rejection comments in reviewtraces + feedback + reviewer notes
stars say nothing — the rejection comment ("note doesn't exist — hallucination?") is the clinician's bug report
6Outcome
EHR evidence lands → reconciled to what the AI saidchart review + LLM attribution + manual citations
conversion happens in someone else's system weeks later — it doesn't exist unless deliberately captured
Public — benchmarks on the same clock · frozena curated vignette stands in for the patient
Knowledge · safety
the answer at the CDS momentUSMLE · safety evals · HealthBench
a synthetic vignette hands the model perfect exposure — layers 1–2 assumed away
Task benchmarks
documentation & admin tasksMedHELM slices — the visit plus a day
graded on output quality, never on adoption — layer 4 invisible
Outcome studies
RCTs — rare, years from question to answerthe only public instrument reaching layer 6
gold standard, wrong cadence: one verdict per multi-year study vs a funnel that updates daily
Uncovered
no public benchmark sees this side of the visitsupply, delivery, exposure on real patients
and the workload denominator — time saved per appointment-day — is measured nowhere public
10 · The lenses

Same artifacts, three audiences

X = which side of the wall; it never changes. Y = value under the chosen lens. Grey dots fall into the chasm; the gold one crosses and becomes a benchmark.
SCOPE — exam item → the whole outpatient visit & its outcome
the chasm · a public / private wall
Public — published & citable
Private — proprietary telemetry
Field studies & surveys
Exam hall
Ops dashboards
Evaluation as product analytics
USMLE / MedQA-style exams84–90% — saturated, at/above physician level
"100% on USMLE" marketingcapability score used as a product claim
HealthBench (2025)5k simulated conversations · 262 physicians' rubrics
ARISE · MAST"living benchmark" — faster refresh, still task-scoped
MedHELM121 clinician-validated tasks · real EHR data
HealthBench Professional (2026)curated from 15k real clinician chats — frozen into a benchmark
Doximity annual survey69% *say* AI improved care — self-report outcome
JMIR systematic reviewthe 39–45 pt knowledge–practice gap, quantified
Abaluck et al.gold-standard outcome · years from question to answer
ARISE State of Clinical AI '26annual synthesis — itself calls for post-deployment eval
CI eval regressionsLLM-judge suites on every model / prompt change
Per-turn feedback streamthumbs + free-text rejection reasons
Latency & cost telemetryper turn, per org
Delivery → exposure → action funnelviewed-denominator accept rates, per clinic
Workload-adjusted utilizationminutes per appointment-provider-day, not DAU
Decision streams (app DB)who accepted, deduplicated, internal actors excluded
Post-visit outcome reconciliationEHR evidence attributed back to what the AI said
OpenEvidence query logs~27M encounters/mo — exists, unpublished as evaluation
VERDICT · publication cadence⟵ the chasm ⟶PROCESS · adaptive cadence
11 · Blind spots

What no benchmark measures

Everything before the answer Supply, delivery, exposure. Benchmarks hand the model a fully-specified vignette — perfect exposure by construction. In deployment, whether output was generated, delivered, and opened is the first half of the funnel, and it lives only in the vendor's database.
Adoption Task benchmarks grade output quality, never whether anyone accepted it. The decisive click happens where analytics SDKs aren't, and vendor staff pass through the same UI as customers — attribution has to be earned.
The workload denominator Time saved per appointment-day — the number clinicians actually care about — is measured nowhere in the public literature. DAU-style metrics measure the clinic schedule, not the tool.
Case: MAST's "Do NOHARM" demo The strongest version of output grading: expert rubrics per case, multi-judge autograder, scores for safety, precision, completeness, restraint. And the evaluation ends at the sentence — it checks whether the recommendation was correct, never whether the action was executed. Recommended ≠ accepted ≠ ordered ≠ administered ≠ outcome; every arrow lives in product telemetry. "Safety: Excellent" is a property of the text, not the visit.
Part IV

The instrument

The evaluation that fits the shape of the problem — and what each audience does about it on Monday.
12 · The thesis

Most model evaluation should be product analytics

The benchmark is the unit test; the deployment funnel is the exam. Six layers, two epochs:
Observed live (usage analytics) 1 · Delivery — did the AI produce and ship anything?
2 · Exposure — delivered ≠ viewed; who never opens it?
3 · Engagement — utilization per unit of real work
4 · Action — accept/reject from the system of record, deduplicated, correctly attributed
Reconciled after the fact (post-usage) 5 · Trust — rejection reasons as expert-labeled error analysis
6 · Outcome — evidence in someone else's system, attributed back to what the AI said

Foundation — data hygiene: exclude internal and impersonated sessions, one clock, identity resolution, one event = one decision. Uncorrected, these bias every number with a flattering sign.
13 · Implications

What should happen next

Builders Instrument decisions as first-class events: log who decided, where the decision happens, and mark your own staff. Report viewed-denominator accept rates and workload-adjusted utilization. Build the reconciliation layer before quoting outcome claims.
Researchers Publish reconciliation methods, not just frozen test sets. Partner for query-log access under governance — the largest evaluative datasets in medicine are sitting unpublished. Treat deployment assumptions as first-class objects of study (per CMU).
Health systems & buyers Stop asking for benchmark scores; every serious vendor clears them. Ask for the funnel: delivery coverage, viewed-denominator adoption, decision attribution, and what the vendor's own rejection log says. A vendor who can't produce these isn't measuring their product.
14 · Sources

Sources