Blog
Clinical reasoning Dr. Alastair Dunne

Why recall isn't reasoning: what you actually need to learn at medical school

Question banks can feel like the centre of medical study. They are not. Here is what recall tools are genuinely for, what they leave out, and where to spend your scarcer hours.

A student completing a written exam.

For the past decade or two, medical study tools have leaned overwhelmingly toward one format: the question bank. Thousands of single best answer items, sorted by system and topic, each one resolving to a single letter. The format became so dominant that most students never stopped to ask why their preparation looks like this, or what it actually trains them to do.

A multiple-choice question is a test of recognition by design. The diagnosis, the investigation and the drug treatment are already printed on the page, and the task is to identify the best one. Clinical practice rarely offers that courtesy. The patient does not arrive with five options attached, and the work is to gather incomplete and sometimes misleading information, hold several hypotheses at once, and decide what to do next when no option is clean. Given all of that, it is fair for a student to ask why so much of what they are sold to study with tests only the first kind of task.

That distance, between recognising the right answer and constructing it from scratch, is the difference between recall and reasoning. It is the part of medical training under most pressure as student numbers grow, because it is the hardest thing to teach and to assess at scale. It is also the gap a student feels on a Monday morning, when the question bank percentage is climbing but the patient in bed four does not match any clean stem.

Why did question banks become the default?

Question banks dominate because recall is the cheapest competency to assess reliably and at scale, not because it is the most important one to learn.

Start with the commercial reality. A bank of single best answer questions is cheap to produce, simple to mark automatically, and easy to wrap in progress metrics and percentages. Under the economics of the past two decades, that made it the obvious thing to build and sell at volume. Assessing the richer skills, the reasoning, the communication and the judgement under uncertainty, is slow, expensive and labour-intensive by comparison. Faced with that trade, the study-tool market built what scaled, and the question bank became the default product.

Medical schools leaned the same way, for reasons most clinical educators will state plainly. Multiple-choice assessment is reliable, auditable and markable across very large cohorts, which matters more each year as student numbers grow, while writing and marking anything that tests reasoning directly is far more costly. Those same educators are usually the first to point out where the format falls short. They know an MCQ tests a baseline of recall, not the clinical reasoning that actually separates a safe doctor from an unsafe one, and many programmes are working hard to assess reasoning more directly. The tools were adopted because they let institutions operate at scale, not because anyone concluded that recall is what makes a good doctor.

What is the difference between recall and reasoning, exactly?

Recall answers a question that someone else has already framed. Clinical reasoning is the work of framing the question yourself, from a real and ambiguous patient encounter. Multiple-choice questions test the first. Clinical practice requires the second.

Recall still matters, because a doctor does need a reliable base of knowledge to work from. But in the digital era, and even more so in the AI era, knowledge recall is no longer the primary skill that defines a doctor. Reasoning is.

The clearest map of this distinction is Miller's pyramid (Miller, 1990Academic Medicine), still the reference framework in assessment science. At its base sits Knows, factual recall. Above it, Knows How, applying knowledge to a problem. Higher again, Shows How, demonstrating a skill in a controlled setting such as an OSCE. At the top, Does, performing with real patients in real practice. Question banks live almost entirely in the bottom layer. That layer matters, but it is not the core of the job, and there is a wide gap between answering at the base of the pyramid and practising at the top of it.

When you study with an MCQ, the hard part has already been done for you. Someone has decided what the question is, narrowed the world to a handful of options, and guaranteed that one of them is correct. On the ward none of that holds. The work is deciding what matters, what to ask next, and what not to be reassured by. Clinical reasoning is the iterative version of that work: you form an early impression, use it to direct what you ask and examine next, and revise as new information arrives. We went deeper on what clinical reasoning actually is in an earlier piece.

What does the evidence say about medical school assessment?

Most medical school assessment has traditionally leaned more toward recall than reasoning, and that shapes how students study far more than any single revision tool does. Written finals have relied heavily on single best answer questions, which are well suited to testing factual recall and much less suited to testing whether a student can reason through an unstructured problem, a point made directly in recent work in Frontiers in Medicine. None of this is a criticism of medical schools or the academics who teach in them. Assessment has to be marked fairly and consistently across very large cohorts, and single best answer questions do that job well. However, the direction of travel is clear. Many programmes are already working to assess clinical reasoning more directly, and purely recall-based study is increasingly at risk of becoming obsolete.

The contrast is easiest to see side by side:

Recall-style assessment itemReasoning task
"Which of the following is the first-line treatment for X?" 

The question is framed, the options are given, and one of them is correct.
"This patient looks unwell and the picture is ambiguous. What do you do next, and why?" 

Nothing is framed yet, and deciding what matters is the work.

The evidence on what builds the second column is reasonably settled. Case-based work, structured reasoning practice and reflective debriefing reliably outperform recall-only review for reasoning outcomes (Plackett et al., 2022BMC Medical Education). Recall practice builds the base. It was never designed to complete the pyramid.

What does AI change about all of this?

AI changes the picture in two ways. The first is about the job you are training for. The second is about the economics of the content you study from.

Start with the job. The part of being a doctor that is pure knowledge recall has been shrinking for a long time, well before generative AI, as guidelines, decision support and point-of-care references steadily moved the facts off the clinician's memory and onto a screen. AI accelerates that sharply. A model can already retrieve and recombine medical facts faster, and across far more of the literature, than any individual ever could. The recall layer is therefore exactly the part of the job that is becoming least defensible, while the work that stays human is the reasoning, the judgement under uncertainty, and the responsibility for what gets decided. If AI is going to shape your career, and it is, it will shape it away from being the best knowledge-recall machine in the room and toward being the person who can reason with, interrogate and take responsibility for what the machine produces. That is the same argument this blog series has made throughout, and it is why the next generation of doctors will need to be better than ever at clinical reasoning.

The second change is economic. AI has collapsed the cost of generating content, which moves the value to something harder: being sure the content is right. For most of the history of medical education, producing good explanatory material was slow and expensive, so owning a large, well-written body of content was a genuine advantage. That has changed. Generating plausible notes, explanations, and even passable questions is now fast and cheap. It is not always accurate, but it is no longer scarce, and anything that is no longer scarce stops being where the value sits.

So the value moves to accuracy, and to how you verify it. A great deal of what a static question bank charges for is the reassurance of an impressive set of credentials standing behind the answers. That was a reasonable model when producing content was slow and expensive, and when keeping it accurate across different health systems was hard in its own right. However, in the hierarchy of evidence that underpins modern medicine (Burns et al., 2011Plastic and Reconstructive Surgery), individual expert opinion sits at the very bottom, with systematic reviews and meta-analyses of trials at the top. Resting the case for correctness mainly on who signed it off sits a little oddly with how medicine weighs evidence everywhere else. The alternative is to ground content in structured, validated knowledge, so that whatever is generated can be checked against it. That distinction, between a confident answer and a verifiable one, is the whole game once content itself is cheap. It is the same reason a general-purpose chatbot is the wrong tool for this, and why how a system reasons matters more than how much it can recite.

This is where it matters for your money as a student. The companies that built their business selling question banks have every reason to defend that model. Their price tag depends on it: the brand, the expert review and the premium those credentials command. That is not a criticism of those companies, or of the experts they use. It is simply the reality of their business model, and what their shareholders and investors expect. For you, though, still working out which study tool is worth paying for alongside everything else, the question is clearer. Does it make sense to spend on a purely recall tool, if recall is not the most valuable part of medicine? How much are you really willing to pay for the credentials of whoever checked the question?

So what are question banks actually for?

They are for building and maintaining a floor of recall, which you genuinely need, and they are not the thing that makes you a doctor. Used well, a question bank does three useful jobs:

  1. It builds a floor of reliable recall, the facts you should not have to stop and think about, which is best laid down through steady repetition across a semester rather than a last-minute sprint.
  2. It calibrates your self-assessment, showing you where you are genuinely weaker than you felt and where your time is best concentrated.
  3. It surfaces gaps early, while there is still time to close them.

Those are real and worth doing. None of them is the same as framing the question from an ambiguous encounter, supervising an AI's output, or sitting with a patient and working out what matters. The mistake isn’t using a question bank. The mistake is using it as the whole of your preparation because it is the tool you happen to have.

It helps to see where each kind of tool actually fits:

Tool categoryWhat it builds wellWhat it leaves outWhere it fits in your stack
Question-bank recall platformsA basic level of factual recall and exam techniqueFraming and reasoning through an unstructured caseLaying down and testing the recall floor
Video and notes content platformsA first pass at understanding new contentApplying that content under real uncertaintyAcquiring and reviewing material
AI OSCE and consultation-practice platformsConfidence talking a case through out loudAny real check that your reasoning is soundGetting comfortable putting it into words
Clinical-reasoning platforms with recall built in (where Gestalt sits)The full stack, from the recall floor up to reasoning across encounters, because recall is built inThe real-patient, bedside experience your clinical placements exist to provideBuilding the whole structure, base to ceiling

For more on what to look for in the reasoning-practice category, see our piece on AI OSCE platforms.

So what should you actually study with?

The more useful question is not which question bank is best. It is what your training is actually asking you to become, and whether your study stack matches it. Assessment and practice are both moving the same way, toward clinical reasoning, communication and judgement under uncertainty. A question bank earns its place by building the recall floor those skills stand on. Beyond that, the tools worth your time are the ones that make you practise the thing you are increasingly being measured on, and the thing that will define you as a doctor: reasoning through a real, unframed problem.

So use a question bank for what it is good at, and do not mistake it for the whole job. Build the floor, then spend your scarcer hours on everything that has to stand on top of it: the reasoning, the communication, the judgement. That is the difference between preparing for the exam you used to sit and preparing for the patient in bed four on Monday morning, who will not arrive as a tidy stem with five options attached.