Edgar Negrete Jimenez was working as an ER tech in downtown Los Angeles when he first noticed something shift. Patients started arriving with AI-generated symptom assessments on their phones. Some were accurate. Some were dangerously wrong. Nobody had taught him how to tell the difference.
He is not in medical school yet. He is already facing the gap.
This is the second Alma First Roundtable: a monthly conversation between practitioners at different stages of their careers, asking what responsible AI use looks like when the people closest to patients are the ones figuring it out in real time. This month, we focus on pre-clinical AI literacy: what medical students are and are not learning before they ever touch a patient.
Medical students now engage with AI tools an average of five times per week, and over 90% report using more than two AI tools in their daily work. But only 14% first learned about AI through their medical schools. The rest figured it out on websites, social media, and peer conversation. The tools are here. The instruction is not. And the less clinical experience a student has, the harder it is to catch the errors the tools produce.
What is medical education still not teaching well enough about AI, and why does it matter?
Andrew O’Malley is a Senior Lecturer at the University of St Andrews Medical School in Scotland, where he leads the SimPatient project, building high-fidelity virtual patients to address capacity challenges in clinical training. His research sits at the center of a question most curricula have not caught up to: what does it mean to evaluate competence when the student had AI help?
Medical education curricula currently fail to adequately instruct trainees on the underlying technical architecture and inherent biases of generative AI. As our research on the SimPatient platform and associated studies into AI-generated medical imagery demonstrate, commercial large language models frequently exhibit demographic and behavioral biases. Trainees need to understand how these models are trained and where their data originates. When students lack this technical literacy, they risk accepting flawed clinical outputs as ground truth.
Current teaching rarely distinguishes between generic chatbot interfaces and fine-tuned tools designed for specific clinical contexts. The next generation of physicians will rely heavily on these systems for diagnostic support and patient interaction. If they cannot critically evaluate the algorithms they use, they compromise patient safety and perpetuate health inequalities. Students are familiar with generic tools like ChatGPT for summarization, but they require targeted instruction on the structural limitations and ethical deployment of clinical AI to ensure equitable patient care.
Edgar Negrete Jimenez is half Peruvian and Mexican, born and raised in Los Angeles, a first-generation college graduate with a degree in molecular and cellular biology now pursuing medical school. He has worked as an EMT, an ER tech in downtown LA, and now as a medical assistant in a neuropsychiatry center. He is also an Alma First fellow. His perspective matters here because he is the person this conversation is about: someone preparing to enter medicine right now, with no formal AI training and no shortage of AI tools.
Medical education still is not teaching enough about AI’s role in healthcare, and that matters to me because technology is advancing faster than our training. I want to have a voice in how AI tools are developed and used, and I worry about the solutions that come with paywalls, leaving patients who cannot afford them behind. It is critical that clinicians with strong medical backgrounds evaluate these systems through a medical lens to ensure they are equitable and safe.
I have seen tools like OpenEvidence, which clinicians access in a clinic with an NPI number, and I think they show that AI has the potential to empower physicians, PAs, and NPs working under supervision. I want to learn how to verify AI-generated information and integrate it responsibly into patient care, because without proper training, I risk relying on technology without understanding its limitations or biases. Equipping myself and my peers to guide AI use will make healthcare more accurate, accessible, and fair.
Gigi Magan is a family physician, researcher, and co-founder of Alma First. She studies AI implementation in safety-net primary care settings and co-leads the Alma First Roundtable series.
What concerns me is not that students use AI. It is that the less clinical experience you have, the more convincing an AI output looks. A third-year resident reads an AI-generated note and notices what is missing because she has written hundreds of notes by hand. A pre-clinical student reads the same note and sees a finished product. The tool’s polish masks its gaps.
Recent research confirms this pattern: when studies assess pre-clinical students on AI literacy, the lowest scores consistently appear in understanding AI terminology, logic, and data science fundamentals. Students know how to use the tools. They do not know how the tools think. And here is what makes this urgent: one study found that participation in AI curricula increased students’ knowledge but simultaneously decreased their enthusiasm for AI in medical education. The more they learned about how these systems work, the more concerned they became. That is not a problem. That is the beginning of the right kind of judgment. But most students never get far enough to reach that discomfort. They stay at the surface, where the tools feel helpful and the risks are invisible.
What struck me: Andrew sees a technical literacy problem. Students do not know how these models are built or where the data comes from. Edgar sees an access problem. The tools behind paywalls serve some patients and not others. I see an experience problem. The less you have seen, the less you question. Three different lenses on the same gap. And none of us mentioned the same tool or the same population. That tells you something about how wide this is.
What skills, mindsets, or habits do future physicians need to use AI well without losing clinical judgment, humanity, or trust?
Andrew:
Future physicians require a mindset grounded fundamentally in critical appraisal. They must develop the habit of treating AI outputs as hypotheses rather than definitive conclusions. In our development of the SimPatient platform, we observed that trainees must maintain robust clinical reasoning independently of the simulated environment.
The primary skill required is the ability to integrate computational insights without degrading interpersonal communication. Physicians need to learn to triangulate AI-generated differential diagnoses with active listening and nuanced patient observation. Trainees must also cultivate a deep understanding of cultural competence. Trust in clinical practice is built on human empathy; therefore, the habit of prioritizing the patient narrative over the digital output is essential. Physicians should use AI to augment their technical capacity, deliberately reserving their cognitive bandwidth for the complex, human elements of care that large language models cannot replicate.
Edgar:
As a future physician, I need to approach medicine with a mindset of constant learning and being part of innovation. I make it a habit to read new research, and I think the suggestions generated by AI should be approached the same way: with peer review and verification. Developing my own clinical judgment is essential, because firsthand patient care experiences cannot be replaced by technology. There are no shortcuts in medicine, and I take that to heart by being thorough in everything I do.
I also recognize that as tools like telehealth expand care to remote areas, I have to be culturally aware and question potential biases in the data, since each patient population has unique needs. I strive to build trust and cultural humility with my patients through respect and connection, and I believe that when I do this, they are more likely to engage with AI-powered digital health tools. I want to use AI to enhance my practice, not replace the human relationships at its core.
Gigi:
The habit I keep coming back to is productive discomfort. Not skepticism. Not resistance. The willingness to sit with the feeling that something looks right and still check.
Medical training already builds this reflex for other domains. We teach students to question a lab value that does not fit the clinical picture. We teach them to re-examine a physical exam finding when the history points in a different direction. AI outputs deserve the same scrutiny, but they arrive with a polish that lab values do not have. A differential diagnosis on a screen reads like an answer. A number on a lab slip reads like raw data. Students are trained to interpret raw data. They are not trained to interrogate polished outputs.
The second skill is knowing what the tool did not see. An AI system generating a differential has no access to the patient’s tone of voice, the hesitation before answering a screening question, the family member in the corner who has been quiet the whole visit. These are the inputs that change clinical decisions. If students build the habit of asking “what did the tool not have access to?” before accepting its output, they protect something no algorithm replicates: the judgment that forms between the data and the decision.
Andrew frames it as treating outputs like hypotheses. Edgar frames it as peer review. I keep thinking about what the tool never had access to in the first place: the pause, the tone, the thing the patient almost did not say. We are all describing the same instinct from different distances.
If you could change one thing about how we prepare trainees for an AI-enabled healthcare world, what would it be?
Andrew:
I would encourage the integration of controlled, high-fidelity AI simulations directly into the core communication and clinical skills curricula. Currently, trainees engage with AI in an ad hoc, unsupervised manner, often utilizing generic commercial models that lack pedagogical alignment. Based on our deployment of the SimPatient platform, replacing this unstructured usage with medically fine-tuned large language models provides a rigorous environment for experiential learning.
By embedding interactive, multimodal virtual patients into formal training, we safely expose students to complex, culturally diverse clinical scenarios that are otherwise restricted by the logistical limitations of human simulated patient programs. This shifts the educational focus from passive technological consumption to active, assessable clinical practice. It allows institutions to systematically evaluate both consultation competence and the student’s ability to navigate algorithmic bias within a risk-free setting. Standardizing AI simulation ensures every trainee develops a baseline proficiency in both clinical reasoning and AI literacy before they interact with actual patients.
Edgar:
If I had full control of the curriculum for one month, I would focus on three things. First, hands-on AI labs where I work with real patient data to assist in diagnoses, learning to understand what the AI gets right and where it goes wrong. Second, immersive telehealth simulations with patients from all kinds of backgrounds, practicing how to use AI insights while still connecting with people, building trust, and being aware of cultural differences. Third, an ethics and responsibility workshop where I reflect on how AI decisions affect patients, spot potential biases, and practice explaining AI-informed recommendations clearly. Combining these experiences would give me practical skills, ethical awareness, and the confidence to use AI in ways that protect my patients.
Gigi:
I would change what we assess, not what we teach. Right now, the conversation about AI in medical education focuses almost entirely on adding content: a new module, a new lecture, a new elective. But content without assessment is optional. Students prioritize what gets tested.
If I had one structural change, it would be this: make AI output evaluation part of how we test clinical competence. Put an AI-generated note into a standardized patient encounter and ask the student to identify what was missed. Include an AI-suggested differential on a clinical reasoning exam and ask the student to explain why they agree or disagree. Build it into OSCEs. Build it into shelf exams. When students know they will be assessed on their ability to question an AI output, the habit forms whether or not anyone teaches a formal AI course.
Andrew is building something like this with SimPatient. Edgar described wanting hands-on labs with real data. The through-line is the same: stop telling students to be careful with AI and start putting them in situations where careful is the only way through.
What Stayed With Me
Andrew builds simulated patients so students encounter complexity before the wards. Edgar wants labs with real data and real consequences. I want assessments that force the habit of questioning.
All three of us, independently, arrived at the same structural argument: the fix is not more information about AI. The fix is more situations where students have to use judgment in the presence of AI. That is a different kind of teaching. It is harder to build, harder to scale, and harder to assess. It is also the only version that works.
Edgar said something I will carry: “There are no shortcuts in medicine.” He said it from the position of someone who has not yet started medical school. That matters. He already knows what many curricula have not figured out: the goal is not to make students faster. The goal is to make sure speed does not replace the thing that makes a physician trustworthy.
Fourteen percent of students learn about AI from their medical schools. The other 86% are figuring it out alone. The question is not whether they will learn. The question is what they will learn without us.
If You Teach, Try This
Discussion questions for small groups or clerkship debriefs:
Where did you first learn about AI tools for clinical work? Was it from your school, a peer, or something you found on your own? What did you learn that turned out to be wrong?
Edgar describes AI as something that “should be approached the same way as new research: with peer review and verification.” What would a peer review process for an AI-generated clinical output look like in practice?
Andrew argues that students need to understand “how these models are trained and where their data originates.” How would knowing an AI model was trained primarily on data from academic medical centers change how you use its output in a community health center?
Classroom exercise:
Show students the same clinical vignette (a 58-year-old patient with diabetes, hypertension, and mild cognitive impairment presenting with fatigue). Have half the group generate a differential on their own. Have the other half use an AI tool. Compare the two lists. Ask: What did the AI include that you would not have? What did you include that the AI missed? What information about this patient was the AI unable to access?
Recommended reading:
O’Malley, A. “From Surgical Prep to Clinical Decisions: Is AI Ready for the Wards?” (andrewomalley.substack.com)
O'Malley, A. “Introducing SimPatient: AI-Powered Virtual Patients for Healthcare Training”
The Alma First Roundtable is a monthly series from Alma First, a nonprofit building leadership pathways for underrepresented communities in healthcare and technology. Each conversation brings together practitioners at different career stages to think through what responsible AI implementation looks like when the people closest to patients are the ones figuring it out.
What is the one thing you wish your training had prepared you for? Tell us in the comments.
Disclaimers: All views expressed are my own and do not represent my employer or any institution I am affiliated with. Any tools, products, or technologies mentioned are included for educational purposes only and are not sponsored or endorsed. Nothing in this piece should be interpreted as medical advice.







