Every medical school is now running an experiment it never designed. Students arrive already using AI to study, and they will practise medicine in a health system where AI runs through diagnosis, documentation, and decision support. Universities understand this. They know AI is part of the future of healthcare, that it will matter enormously across their students' careers, and that it has to be brought into the curriculum rather than held at the door. Many can also see that, used well, it may help relieve some of the otherwise intractable pressures on the traditional medical school model, from rising student numbers to the rising cost of delivering teaching at quality.
So this is not simply apprehension. Medical educators can see the opportunity clearly. The difficulty is the speed: it puts real pressure on a university's ability to absorb it comfortably, at a pace educators do not control, and the people who must make decisions about it are being asked to do so while the ground is not just moving but accelerating beneath them.
The honest position for most schools is that AI is already in the building and the policy is still being written. The question is no longer whether students will use these tools. They already do. HEPI's 2025 student survey found that 92% of students now use AI tools, up from 66% a year earlier, and that 88% have used them for assessments. Among medical students specifically, a multicentre survey found roughly four in five already using generative AI, and more than nine in ten wanting to learn how to use it in their future practice. Your students use these tools now, they will use them throughout their careers, and part of your job is to prepare them to do so well.
Given that, governance is not about saying no. It is about being able to say yes to the right things, on purpose, and fast enough to matter.
Why you cannot govern AI reactively
The instinct in any institution facing a new technology is to wait. Let the dust settle, see what good practice emerges, then write the policy. With AI, the dust does not settle. The rate of change is unprecedented, and that single fact is what makes AI unlike the technologies medical schools have absorbed before. A rule written for last year's tool is often meaningless against this year's, because the tool itself has changed underneath it. This should not be treated as just another new technology arriving on the usual timescale.
The exposure is also wider than it first appears. It is not only that students use these tools to study. The material people encounter day to day is increasingly AI-generated, and the path from a system's output back to a trustworthy source is often opaque: it can be hard to know what a model was drawing on, or whether it was drawing on anything real at all. That problem is not confined to students. Faculty and staff are increasingly leaning on the same tools to work faster, to draft, summarise, and prepare teaching material, sometimes with invented citations or unverifiable claims folded in unnoticed. So the governance question is not only about what your students do. It is about what your whole institution is now producing and consuming with AI in the loop.
For an institutional buyer, the deepest problem with a reactive stance is not caution, it is speed: reacting is structurally too slow. Acting reactively means waiting until a tool is on the market and then putting it through the assessment process you already have. But those processes were built for stable products with a fixed specification. By the time a procurement or curriculum process you designed even two years ago has run its course, the thing it assessed has often already been superseded, and the conclusion is obsolete before it is signed off.
That speed problem is also a safety problem, which is what makes it matter, because keeping students safe is the whole point of governing in the first place. By the time you have reached a considered position on one tool, something new has arrived from another direction carrying the same kind of risk, so a slow, case-by-case response never closes the gap it exists to close. Being reactive does not, in the end, keep anyone safe. The only workable answer is to be proactive: to decide in advance what you require of any tool, so that when something new appears you can judge it almost immediately against principles that outlast the particular product in front of you, rather than starting from scratch each time while the risk sits unmanaged.
What "AI in medical education" actually covers
Part of why this is hard is that "AI in medical education" is not one thing. The label is stretched across very different kinds of tool: consumer chatbots students use on their own, general-purpose models dressed up with a medical interface, tools that constrain a model to an approved set of sources, clinical tools built for practising doctors that students borrow, and systems built around structured knowledge designed specifically for learning. Each carries a different risk profile, and a policy that treats them as one category will get all of them wrong. We have written about these different kinds of tool in more detail across our other posts; for governance, the move that matters is to stop asking "should we allow AI" and start asking "which of these, for which job, and with what assurances."
The two jobs governance has to do
Governing AI in a medical school really means doing two jobs at once.
The first is procurement: deciding which tools the institution adopts, pays for, or recommends. This is where you carry the most responsibility, because a tool the school endorses carries the school's name.
The second is policy for student use: what students may and may not do with the tools they bring themselves. This is the harder half, because it runs straight into assessment. Where is the line between using AI to understand a topic and using it to complete an assignment? What does an exam mean when a plausible answer is thirty seconds away? Schools need a position students can actually understand and follow, not a blanket prohibition that everyone ignores.
Both jobs rest on the same underlying ability: to look at a tool and know what it is, where its content comes from, and whether you can stand behind it.
The questions to ask any vendor
Most procurement instincts were built for static products with a fixed specification. AI tools are not static, so the questions have to change, and they have to be asked explicitly, because the failure modes here are quieter than the ones procurement is used to catching.
- Provenance: Where does the knowledge come from, and can you verify it? If the honest answer is "the model knows," that is not a source. You should be able to trace a given answer back to something a clinician would recognise as authoritative.
- Architecture: Is this a raw model with an interface on top, or is it genuinely constrained, and if so, to what? What is the system actually permitted to draw on when it answers a student?
- Fit to the learner: Can it be sandboxed to our curriculum and to a student's stage, or does it answer everything at the level of a specialist? A tool that cannot be held to the level of the learner is hard to use reliably in a programme.
- Ownership and IP: This runs in two directions. First, what goes in: who owns the content and curriculum material that passes through the tool, and how is your faculty's intellectual property protected, given how much of it will move through it in the course of teaching? Students feeding course material or their own work into a consumer tool can hand the university's IP to a third party without anyone ever consenting to it. Second, what comes out: if a tool generates material that infringes someone else's IP, the liability does not stay neatly with the vendor. Depending on the jurisdiction, liability for infringing output may extend to the user of a tool, and not only its provider, so "the tool produced it" may not be a defence you can count on.
- Data and data sovereignty: What happens to student data, and how is patient data protected? Is any of it used to train the provider's models, or monetised beyond running the service for us? This matters more in medicine than in most settings, because students are placed in a privileged position: through the requirements of their curriculum, they operate in clinical environments where they are exposed to real patient information outside the protected silos of health IT systems. So a fair question to ask is what would happen if such information were entered into a tool, deliberately or not: where would it go, could it be written back into the product, and could it end up training a model? A school has to be careful about what tools it gives its students. There is also the question of where data physically sits: which country holds it, and under whose laws and jurisdiction does it then fall? Knowing where your users' data, and any health data that finds its way into a tool, ends up is not an optional extra. It is a core part of due diligence.
- Educational effect: How does this actually support learning, and what does the vendor do to avoid de-skilling and never-skilling rather than simply handing students answers? Does it fit what we know from educational science, or does it just deliver answers efficiently? This is an educational procurement, not only an IT one, and the pedagogy is a fair thing to interrogate.
- Representation and localisation: Can the tool represent the health needs of the specific populations your students will go on to treat, including minority groups and the local clinical context, or does it default to the majority case? A language model gives you the consensus, not necessarily the truth, and the consensus skews toward the majorityby the nature of the mathematical models underneath. The point is to look past what a vendor says it values and ask whether the tool can actually do this: it is a question of capability, not of intent. A tool that only generates will tend to flatten difference toward that majority, whereas one built on structured, localised knowledge can be made to hold the minority presentations and the local context deliberately. We have made the case for local clinical contextbefore, and the same logic applies to any population a general model has seen little of.
- Dependency: Easy to overlook, and central to the risk. What does the tool itself depend on? A product built as a wrapper around a single third-party model, or leaning on an avatar or components licensed from elsewhere, inherits every change, price rise, outage, and policy shift of the things beneath it. That fragility does not stay with the vendor; it becomes the institution's problem, and potentially the students' problem mid-course. It is a risk buyers consistently underrate: in one survey of executives with active AI vendor contracts, 81% were worried about this dependency, yet only 6% believed they could lose their primary AI vendor without disruption.
- Exposure: What is the security and reputational exposure if the tool produces something wrong, biased, or harmful while carrying your name?
None of these are exotic. They are the questions any board asks of any serious supplier: what are we relying on, what could go wrong, and how would we know in time to act.
Why "check the editorial board" is no longer sufficient
For a textbook or a static reference, there was a dependable shortcut: see that a credible editorial board had signed it off, and trust it accordingly. That breaks down with generative tools, because there is no fixed output to sign off — the tool produces a slightly different answer every time it is asked, and reviewing each one by hand is impossible by definition. A serious roster still tells you something, but it can no longer settle the question, which moves from "who put their name on this" to "where did the information in this answer come from, and can you show me." It is the shift we argued for when we looked at reference tools, and it is what a list of names, however credible, can no longer provide on its own.
This is an extension of how governance already works
None of this requires inventing a parallel bureaucracy. The Institute of Directors in New Zealand's guidance for boardsmakes a useful point: AI is best treated as a governance and performance issue rather than a technology project, and the instruments needed to govern it largely already exist. You adapt what you already use rather than building something new. Whatever you already use to set your institution's appetite for risk, the tone for what is and is not acceptable, can carry your position on AI: the legal, ethical, and reputational boundaries within which any tool, whether bought centrally or brought in by a student, has to sit. Drawn deliberately, those boundaries are not a brake on innovation; they are what let you say yes quickly, because the lines are already clear.
That guidance frames the work as three moves, and they translate cleanly from a boardroom to a faculty. First, take a stocktake: find out what AI is actually being used across the school today, by students and by staff, rather than assuming. Second, set the boundaries explicitly, covering privacy, data, intellectual property, attribution, and accuracy, so people know where the lines are. Third, revisit regularly, because this is not a set-and-forget exercise; what you settled last year will not hold this year, and increasingly what you settled last semester will not hold this one. The discipline underneath is to be clear about what you need to assure yourself of directly, what you can delegate to others to run, and what you keep under continuous review.
It is important to understand that governance best practice says you cannot abdicate your responsibilities to AI. A tool can inform a decision, but accountability for it stays with the people who chose to rely on it. For a medical school, that means the question is not whether a vendor's system looks impressive. It is whether your use of it is deliberate, whether you can explain it, and whether it is backed by enough evidence to justify the confidence you are placing in it.
What is at stake if you get it wrong
It is worth being blunt about the stakes, because what is being decided here is far larger than any single procurement decision. Get one tool wrong and you can unwind it. Get the governance wrong and the failures compound, because the same blind spot repeats across every tool the school adopts and every cohort that moves through it.
The governance risks particular to these products are quiet rather than loud, which is exactly what makes them hard to govern. The failures those vendor questions are built to catch — de-skilling students who skip building a differential and overtrust the answer, intellectual property and real patient data leaking into unvetted systems, dependence on a model that can be repriced or withdrawn under you mid-programme, bias baked in by a tool that defaults to the majority case — all share a shape: hard to spot in a demo, hard to reverse once they take hold, and often visible only long after the decision that caused them. That is precisely the profile of risk a reactive, tool-by-tool process is worst at catching.
And the consequences do not stay inside the institution. A medical school ultimately holds a social licence to operate: it exists because society trusts it to produce competent doctors in the AI era, doctors who are worth more than a language model on its own. What a school endorses is also, in large part, the basis for how its students will carry AI forward into their own careers, which means the norms of AI in medicine will be set in its medical schools. Getting this wrong at scale is not one bad purchase to be written off; it puts at risk the very thing the institution exists to protect.
This is why governance and learning are, in the end, the same conversation: the question is not whether to permit these tools, nor even whether to govern them, since opting out is not on the table, but how to govern them so that students emerge better prepared than any cohort before them. Get that right and the stakes invert: a stronger workforce, faculty who can teach with confidence, and an institution known for having handled this well.
What good looks like
This is not an argument for caution as an end in itself. The upside is enormous. Used well, these tools can extend students further than any cohort before them: more practice, more feedback, more reps at clinical reasoning and recall than a traditional curriculum could ever provide. The goal of governing them well is not to hold that back. It is to make it real, and to graduate the best-prepared doctors we have ever trained.
Good governance is what makes that upside safe to pursue. In practice it means a few things. Accept that your students already use these tools and will use them throughout their careers, so prepare them rather than pretend otherwise. Hold any tool you adopt to real questions about provenance, sandboxing, data, and dependency, and do not accept a list of names in place of a verifiable source. Favour tools that represent medicine as structured, traceable knowledge over those that merely generate or retrieve. And anchor on principles rather than the product features of the moment, because the principles are what will still be standing when the next model ships.
The responsible-use principles that bodies like the AAMC have set out point the same way: transparency, human oversight, and equity, applied deliberately rather than after the fact. None of this requires waiting for the field to settle. It requires deciding, now, what you need to be able to verify, and then asking every vendor, us included, to show you they can.