I spent the last month preparing a lecture on ambient AI scribes for the California Telehealth Resource Center. I read the randomized trial. I read the big multisite study. I read the burnout cohorts, the note-quality audits, the qualitative interviews, the legal analyses. Forty-one references, each one checked against the primary source.
Here is what the map shows, and here is the hole in the middle of it.
What the evidence actually says
The technology is real, and the benefit is real. I want to be precise about its size. Here is what the three most important studies in the field tell us:I use one of these tools every day. I am not going back to typing. That is exactly why the next part matters.
The hole in the map
Line up the settings of every major study: UCLA. Mass General Brigham. Emory. UCSF. Yale. UC Davis. Kaiser Northern California. Penn. Stanford.
Every one an academic or large integrated system. Every one on the same enterprise EHR. Every one with implementation teams, IT depth, and slack in the schedule.
The number of published trials or large studies conducted in federally qualified health centers or community health centers: zero.
The number of peer-reviewed studies reporting scribe accuracy stratified by patient race, language, or accent in real clinical use: also zero.
Now hold that against what we already know. The landmark test of commercial speech recognition found error rates nearly twice as high for Black speakers as for white speakers, and more than twenty percent of Black speakers’ audio was degraded beyond usability, versus under two percent for white speakers. That was general-purpose technology in 2020, and vendors have invested heavily since.
But here is why the absence of FQHC studies is not just a gap in the literature. It is a clinical risk. We already know the underlying technology struggles with diverse populations. The very clinics that serve those populations are entirely missing from the evidence base. Independent, demographically stratified validation of today’s medical scribes is essentially absent from the literature. We are taking the improvement on faith.
Physicians in the richest qualitative study we have said the quiet part: the tool had limited functionality with non-English-speaking patients.
Why this creates ethical tension
Health centers serve more than 31 million people. The visits are long, multilingual, and heavy with social complexity, which means documentation burden per visit peaks exactly there. If ambient AI works anywhere, the human payoff should be biggest in the safety net.
And the risk concentrates in the same place. The transcription disparities land on these patients. Consent is hardest where trust is thinnest and languages are many. California safety-net leaders told the California Health Care Foundation the rest: pricing models built for enterprise margins, thin IT staffing, liability worries with no in-house counsel.
So the two-tier scenario writes itself. Well-resourced systems buy back their clinicians’ evenings. Safety-net patients get error-prone notes. Not because anyone chose it. Because the evidence, the pricing, and the implementation support were all built somewhere else.
The patients most likely to benefit are the least represented in the evidence and the most exposed to its failure modes.
The equity lens, operationally
Who carries the risk? Follow the pipeline. Speech recognition degrades on accented and code-switched speech. The patients affected are the least likely to notice, contest, or even access what the record says about them. The clinics serving them are the least resourced to audit it. And the studies that would surface the problem have not been run where these patients receive care.
None of this is an argument against the technology. It is an argument about who gets to generate the evidence.
How I now think about this
I have stopped waiting for the literature to come to us.
If you practice in the safety net and you deploy one of these tools, your QI program is the study. Decline rates by language. Note quality by population. Edit burden on interpreter-mediated visits. None of that requires a grant. It requires deciding, on day one, that equity gets measured instead of assumed.
If you hold a procurement pen, you hold the only leverage that reliably moves vendors: stratified accuracy data by language and demographic group, before contracting. What contracts require, roadmaps deliver.
And if you are a frontline clinician who controls neither procurement nor the QI dashboard, you still control the feedback loop. When the tool drops the cultural context or misrepresents an interpreter’s translation, do not just fix the note. Flag the error to your medical director. Create the internal paper trail. The data that does not exist in the literature can start accumulating in your clinic today.
A gap in evidence is not a verdict. It is a to-do list.
Digital health pearls
•Measured time savings are modest: about a minute per appointment at scale. The attention benefit is the real product.
•Platform choice changed the result inside a single randomized trial. Demand data from settings like yours.
•No FQHC has hosted a major study. Whoever runs the first one changes this field.
•Stratified accuracy data is a procurement demand, not a favor. Ask before you sign.
•What you do not measure, you cannot defend. To your board, or to anyone else.
An invitation to compare notes
I gave a lecture this week to an audience of people who support telehealth in exactly the settings the literature skipped. My working theory is that the first FQHC-based evidence in this field will come from someone in that room, or someone reading this.
If your clinic is piloting an ambient scribe outside a big academic system: what are you measuring? I mean that as a real question, and I would like to compare notes.
Disclaimer: All views expressed are my own and do not represent my employer or any institution I am affiliated with. Patient stories in this publication are composites drawn from multiple encounters, with identifying details changed; no single patient is described, and any resemblance to a specific individual is coincidental. Tools, products, and companies mentioned are discussed for educational purposes as commentary and opinion; I have no sponsorships or financial relationships with them unless explicitly disclosed. Nothing here is medical advice, and reading or corresponding with this publication does not create a doctor-patient relationship. Always consult your own physician.








Thank you for this thoughtful article. What stayed with me most was your reminder that “a gap in evidence is not a verdict—it is a to-do list.” That’s a principle I hope we carry into AI more broadly. As these systems become part of everyday life, I hope we continue asking not only how capable they are, but who has been included when we measure that capability. The people who stand to benefit the most should never become an afterthought. Thank you for writing this.