What to monitor for an embedded AI assistant

One clock, anchored on the visit. Left of the dashed line is observable live; right of the epoch boundary must be reconciled after the fact. The same clock also shows the public/private chasm as two bands: the top band is the builder's proprietary telemetry — continuous, adaptive, process-oriented; the bottom band is where public benchmarks sit — frozen slices, publication cadence, verdict-oriented. Amber notes = the blind spot consumer-app thinking has at each layer.
Usage analytics — observed liveproduct analytics + app DB
Post-usage — reconciledEHR evidence + reconciliation, batch
−7d
−5d
−3d
−1d
VISIT
+1d
+3d
+7d
+14d
+30d
Private — product telemetry · continuous · improves the process
1Delivery
generated → publishedapp DB — the tracker never sees it
no consumer analogue: the product pushes — generation can stop for one clinic while every usage chart still looks alive
2Exposure
sits in the queue → viewed pre-visit or at the visittracker views ∪ app-DB opens
"delivered" isn't reach — one deployment's accept rate went 12%→44% when the denominator became viewed
3Engagement
the session itselfopen time · chat turns · patients touched
DAU measures the clinic schedule — normalize per appointment-provider-day, not per calendar day
4Action
accept / rejectapp DB, deduplicated · some land days later
the decisive click is where the SDK isn't — and know whose click: vendor staff use the same UI (one deployment: 76 of the first 96 "verdicts" were internal reviewers)
5Trust
thumbs at the turn → rejection comments in reviewtraces + feedback + reviewer notes
stars say nothing — the rejection comment ("note doesn't exist — hallucination?") is the clinician's bug report
6Outcome
EHR evidence lands → reconciled to what the AI saidchart review + LLM attribution + manual citations
conversion happens in someone else's system weeks later — it doesn't exist as data unless deliberately captured
Foundation
data hygiene — every row above is wrong without itinternal & impersonated sessions out · one clock · identity resolution · one event = one decision · name the unattributable
uncorrected, these don't add noise — they add bias with a flattering sign
Public — benchmarks on the same clock · frozen · verdicts publication cadence; a curated vignette stands in for the patient
Knowledge · safety
the answer at the CDS momentUSMLE · safety evals · HealthBench
a synthetic vignette hands the model perfect exposure — layers 1–2 assumed away, no real chart
Task benchmarks
documentation & admin tasksMedHELM slices — the visit plus a day
graded on output quality, never on adoption — layer 4 invisible
Outcome studies
RCTs — rare, years from question to answerthe only public instrument reaching layer 6
gold standard, wrong cadence: one verdict per multi-year study vs a funnel that updates daily
Uncovered
no public benchmark sees this side of the visitsupply, delivery, exposure on real patients
and the workload denominator — time saved per appointment-day — is measured nowhere on the public side
Sources per layer: 1–2 app database (invisible to trackers) · 3 product analytics · 4 both, deduplicated · 5 traces, feedback, human review · 6 deliberate outcome capture (reconciliation pipeline). Companions: the verdict-vs-process quadrant · the long-form essay.