One chart, then the argument

The encounter

Start with the thing being measured. Pivot the timeline on the note and both sides go dark — nobody hydrates the context coming in, nobody tracks the inbox going out. The note itself is 12.5% of the arc.

Anatomy of an encounter

The clinical process runs from booking to closure. Every tool sits in one narrow band of it — and three bands have nothing in them at all.
The encounterclinical process
Schedulingwhy they booked
Prechartingread the whole chart
Historythe interview
Examhands on the patient
CDSthe decision
Documentationthe note
Follow-updid it happen?
booksdays beforeIN THE ROOMafterweeks → months
Answer enginesOE · DoxGPT
the question, answeredyou gather and paste the context by hand
ScribesAbridge · Ambience
hears the room → writes the notenever reads the chart — no precharting
Nobodyno product, no eval
?
?
?
The three empty bands are not mysteries — they are where the questions already go. Classified and counted, the stream is the workflow eval those ? boxes are waiting for.
the ? boxes, answered · praxis — question stream → intent → subintent · de-identified · counts land here as the taxonomy matures
lands before the noteContext questions
“what changed since her last visit?”
“summarize the outside cardiology notes”
“which meds still need reconciling?”
interval history · chart synthesis · med rec
lands at the notePoint-of-care questions
“max metoprolol dose in CKD?”
“differential for this presentation?”
“what does the guideline say here?”
med dosing · differential · guideline lookup
lands after the noteFollow-through questions
“draft a message explaining this result”
“handle this refill request”
“write the prior-auth appeal”
result interpretation · patient messaging · admin drafting
The compute belongs on the dark sides of the note — hydrate the chart coming in, track the inbox and closure going out. Every question above is a vote for where it goes; classified and counted, the stream becomes the workflow eval the ? boxes are waiting for. And learn how clinicians actually use it before trying to teach them. One recommendation, walked end-to-end →
The chart is the whole picture. Every tool hands you a lens instead — so sweep it around and see how little comes into focus, and how long it takes.
simulated questions
ask a question, or sweep the glass yourself
patient context 0
Empty. Whatever you hand the answer engine, this is all of it — and right now it is nothing.
0 of 0 items in context
You uncover only what you thought to look at, in whatever shape your sweep happened to make. No benchmark scores what never came into focus.
Run the same patient twice. Identical engine, identical question — the only difference is whether anybody swept the chart first.
the same patient, twice · illustrative
⏺ without the chart
T+0
“Cellulitis — what antibiotic?”the same question, both times
T+0
Context assembled: 2 itemsage, sex, the complaint
T+0
The chart is not sweptthe 2019 outside ED record is never opened
2019-04 · OUTSIDE ED · amoxicillin → ANAPHYLAXIS
T+2 min
Prescribed: amoxicillin–clavulanatea guideline-standard choice for cellulitis
T+90 min
outcome
Anaphylaxis
epinephrine · ED · admitted overnight. The allergy was in the chart the whole time.
⏺ with the chart
T+0
“Cellulitis — what antibiotic?”the same question, both times
T+0
Context assembled: 38 itemsincluding the scanned outside records
T+0
The chart is sweptthe 2019 outside ED record surfaces
2019-04 · OUTSIDE ED · amoxicillin → ANAPHYLAXIS
T+2 min
Prescribed: doxycyclinethe guideline’s choice when penicillin is out
T+6 d
outcome
Resolved
no reaction, no return visit. Nothing about the model was different.
Both prescriptions are correct answers to the question that was asked. Score either one against a rubric and it passes — the drug is guideline-appropriate for cellulitis, the reasoning is sound, the citation is real. The harm is upstream of everything a benchmark can see.

The question that never gets asked

The precharting stream, filtered by the doctor, classified by praxis — and the questions that never form.
no productNOTHING SERVES EXTRACTIONcherry-picked by the clinicianDATAPRECHARTINGthe doctorEXTRACTIONREASONINGOE · DoxGPT · ChatGPT · UpToDate …answer engines — counted as usageLOST
speaker notes
Every gray dot is a before-visit inquiry — unscheduled, uncounted, unpaid; most of the chart is never reviewed.
The doctor is the filter: attention, time, and what they know to ask — only a fraction enters, and color is assigned on the way out.
praxis classifies what emerges (intent → subintent): extraction = interval history · med rec · outside records · guideline lookup; reasoning = differential · risk synthesis · dose in context · goals of care.
Both lanes end in the answer-engine log, counted as product usage (“1M consultations a day”) — the same cognition, relocated; the tool takes credit.
Never formed — the red dots in the pile: “could this be ATTR amyloid?” exists only if you’ve read the 2026 guidance — an unasked question is indistinguishable from no need.

What we benchmark instead

Every instrument ever built for medical AI aims at that 12.5% — the answer at the moment of the note. Here they are, era by era, and here is the same chart recolored to ask who each one is for and who gets to see what it produces.

A benchmark score is an engine on a stand.
what we benchmarkA Mercedes-AMG Formula 1 turbo-hybrid power unit on a stand, disconnected from any car
The power unit, alone.
Mercedes-AMG F1 · Hullian111, CC BY-SA 4.0
what sets the timeA McLaren Formula 1 car on track at the 2026 Austrian Grand Prix
The same engine, in a package.
McLaren, 2026 Austrian GP · Lukas Raich, CC BY-SA 4.0
the engine is necessary · the package sets the time
One Mercedes power unit goes into four teams in 2026. They do not finish anywhere near each other. Every instrument on the chart grades the engine.

Four eras of measuring medical AI

2019–24Exam era“Are you booksmart?”

Multiple-choice recall — integrated retrieval, not practice. Saturates at 84–90%, at or above physician level; falls to 45–69% on clinical tasks. Ends as marketing.

2023–Rubric era“Can it converse safely?”

Physician-written rubrics over multi-turn conversations — a curated vignette stands in for the patient.

2024–Task era“Can it do clinical work?”

Real EHR data and agentic workflows — graded on output quality, never on adoption.

2025–Deployment era“What does deployment show — and did it change the patient?”

Real use and real outcomes, one era: benchmarks curated from live telemetry, while the outcome question runs daily as private telemetry.

patient-chat long tail, not charted: K Health · Ada · Babylon · HealthTap · Woebot · symptom checkers & wellness bots
JAN 2026
The instruments, as they arrived
January 2026 — where the exam era actually endsPatients had been asking for three years — 230 million health questions a week — before any instrument measured it. Then the labs stopped selling scores and built for that instead, in a single week: ChatGPT Health on the 7th, Claude for Healthcare on the 12th. The question stops being can it pass? and becomes what is it doing?

Two questions, one chart

The argument

Medicine already knows how to close a loop — it does it thousands of times a day, and chases the ones left open. The AI recommendation is the one event in the chart that never got a close. The reason is access, not rigour.

Medicine runs on tickets

Every workflow the chart trusts is a ticket. Something opens; something specific is allowed to close it. Open-without-close isn’t a gap — it’s an error the system chases until it dies.

The ticket nobody closes

Run the same schematic on an AI recommendation. The event happens. The action happens. And the ticket that should close it — the follow-up, the outcome — is never even opened. Not failed: unlogged.

Everyone grades the model. No one grades the outcome.

Every instrument and product on the chart lives on one side of a wall: they grade artifacts of the session — answers, notes, sandboxed orders. Whether the thing was done lives on the other side, inside the chart. The literature is lopsided in exactly this shape: of 4,609 clinical LLM studies, 1,048 touched real patient data and 19 were prospective randomised trials; an earlier review found 5% used real patient-care data at all. Governance has noticed — CHAI and the Joint Commission now require monitoring for “changes in outcomes” — but a mandate is not an instrument: none of it says what to measure, over what window, against what denominator. The wall is EHR access.
Grading the model
0 instruments & products — the full timeline. everyone.
THE WALL — EHR · PATIENT-DATA ACCESS
Grading the outcome
the same timeline — one entry here; 19 prospective trials in 4,609 studies
The follow-up lives in the EHR. Almost no evaluator can see the EHR. That’s the whole story.
Sources

References & prior art

Every source behind the chart, generated from the same cards the marks open · full bibliography in REFERENCES.md