the chasm · a public / private wall
Public — published & citable
Private — proprietary telemetry
Field studies & surveys
Exam hall
Ops dashboards
Evaluation as product analytics
USMLE / MedQA-style exams84–90% — saturated, at/above physician level
"100% on USMLE" marketingcapability score used as a product claim

HealthBench (2025)
5k simulated conversations · 262 physicians' rubrics

ARISE · MAST
"living benchmark" — faster refresh, still task-scoped

MedHELM
121 clinician-validated tasks · real EHR data

HealthBench Professional (2026)
curated from 15k real clinician chats — frozen into a benchmark

Doximity annual survey
69% *say* AI improved care — self-report outcome

JMIR systematic review
the 39–45 pt knowledge–practice gap, quantified
Abaluck et al.gold-standard outcome · years from question to answer

ARISE State of Clinical AI '26
annual synthesis — itself calls for post-deployment eval
CI eval regressionsLLM-judge suites on every model / prompt change
Per-turn feedback streamthumbs + free-text rejection reasons
Latency & cost telemetryper turn, per org
Delivery → exposure → action funnelviewed-denominator accept rates, per clinic
Workload-adjusted utilizationminutes per appointment-provider-day, not DAU
Decision streams (app DB)who accepted, deduplicated, internal actors excluded
Post-visit outcome reconciliationEHR evidence attributed back to what the AI said

OpenEvidence query logs
~27M encounters/mo — exists, unpublished as evaluation