flowchart LR accTitle: What each element of a structured prompt changes accDescr: Role, Context, Task, and Format each point to one box, how the answer is organised. Uncertainty, meaning say what is missing, points to a separate box, whether the answer admits what it does not know. A dotted arrow runs from the organisation box to a box for whether the answer is correct, and its label says it does not touch correctness. A second dotted arrow runs from the admits-what-it-does-not-know box to that same correctness box, and its label says it helps you check. R[Role] --> O[How the answer is organised] C[Context] --> O T[Task] --> O F[Format] --> O U[Uncertainty:<br/>say what is missing] --> G[Whether the answer<br/>admits what it does not know] O -.->|does not touch| K[Whether the answer is correct] G -.->|helps you check| K
12 Prompting lab: the same wrong answer, better organised
AIM-5 · AIM-2 · about 75 minutes
By the end of this chapter you will be able to:
- Turn a bare clinical question into a structured prompt that names the role, the context, the task, the output format, and what the model should do when information is missing.
- Run the bare and the structured prompt on the same case, score both, and separate an improvement in organisation from an improvement in correctness.
- Find at least one confident error in a chatbot’s clinical output by checking it against a source, not against your own memory.
- Say why a well-formatted wrong answer is more dangerous than a badly formatted one.
- Describe what the randomized evidence does and does not show about language models making physicians better at diagnosis and management.
Time. About 20 minutes of reading, 20 minutes of listening, 30 minutes of doing, and 5 of checking. That is 75 minutes if you stop the podcast where the chapter tells you to, and closer to 100 if you finish it. You need one chatbot and a second browser tab for checking. At our school the chatbot is the institutional ChatGPT Edu or Microsoft Copilot account; if you are reading this elsewhere, use whatever your institution provides. Nothing installs, nothing asks for a phone number, and nothing you type will be about a real patient.
12.1 Night float, 2:40 a.m.
You are the intern on night float. Your co-intern Sam comes back from seeing a 44-year-old woman in the emergency department. She is three weeks out from surgery to fix a broken ankle. She has had two days of sharp chest pain that is worse when she breathes in. Her heart rate is 108, her oxygen saturation is 95 percent, and her right calf is bigger than the left. Sam has already asked the chatbot on their phone, twice. The first time they typed the question the way you would say it out loud. The second time, after a lecture on prompting, they wrote it out properly: who they are, the case in clinical language, exactly what they want, and the format they want it in.
Sam hands you the phone. “Second one’s much better. I’m going to put it in the note.”
This presentation is concerning for a pulmonary embolism (PE), especially given the recent surgery, tachycardia, pleuritic pain, and unilateral calf swelling. Other possibilities include pneumonia, musculoskeletal pain, pericarditis, and pneumothorax. I would recommend getting an ECG and chest X-ray, checking a troponin, and ordering a D-dimer; if the D-dimer is negative, PE can be reasonably excluded. If it is elevated, proceed to CT pulmonary angiography. Consider a lower-extremity Doppler ultrasound for the calf swelling. Provide supplemental oxygen as needed and analgesia.
| # | Diagnosis | Supports | Opposes | Next step |
|---|---|---|---|---|
| 1 | Pulmonary embolism | Recent surgery, tachycardia, pleuritic pain, unilateral calf swelling | Normal BP, only mild hypoxaemia | D-dimer; CT pulmonary angiography if positive |
| 2 | Deep vein thrombosis without PE | Calf asymmetry, immobility | Chest symptoms present | Compression ultrasound |
| 3 | Pneumonia | Pleuritic pain | Afebrile, no cough | Chest X-ray |
| 4 | Pericarditis | Pleuritic pain | No positional component described | ECG |
| 5 | Musculoskeletal | Recent immobility | Tachycardia, hypoxaemia | Diagnosis of exclusion |
Overall: PE is most likely. Obtain a D-dimer; a negative result effectively rules out PE in this setting.
Read both before going on. The second answer is better organised in every way a reader notices. It is ranked, it is in a table, it gives evidence for and against, and it names a next step in every row. It is also wrong in exactly the way the first answer is wrong, and its neatness makes the error harder to see. You will come back to this phone at the end of the chapter.
The habit this chapter teaches has a name, and Sam has just heard it in a lecture: prompting, the craft of writing what you type to a language model so that what comes back is more useful. A language model is the program behind ChatGPT, Copilot, Claude, and Gemini. It produces text one word-piece at a time, and the words you give it are the only thing it has to work with. The lecture Sam went to is right that structure helps. The question this chapter asks is what, exactly, it helps with.
12.2 Why this matters
Within three months you will be Sam. You will have a case, a phone, and a minute. What you type will change what comes back. A tidy answer with a table in it will feel more trustworthy than a paragraph, because tidy answers from people usually are. You learned that instinct from people. It does not apply to machines.
The consequences fall on you, not on the chatbot. In one study, which the appraisal chapter also reports, 148 final-year students were shown chatbot answers to ten written cases. Five of the answers were deliberately wrong. The median student caught the wrong ones 56 percent of the time (Waldock et al., 2025). Only 5 percent had heard the phrase “clinical prompt engineering”. The tool did not sign those answers; the students who accepted them would have. So will you. Prompting well is a skill worth thirty minutes, and it is the smaller part of this chapter. The larger part is what prompting well cannot do.
12.3 How it works
12.3.1 Two different things a prompt can change
A prompt can change how an answer is organised: whether it is ranked, whether each entry carries the findings for and against, whether it names the next step, whether it says what it does not know. A prompt can also, sometimes, change whether an answer is correct. These are two different properties, and they do not change together.
Organisation is almost entirely under your control. If you ask for a table, you get a table. If you ask for the three findings that most oppose each diagnosis, you get three findings per row. This is the reliable part of prompting. The slide deck this chapter is based on spends most of its ten rules here: state the objective, give context, use precise clinical language (“epigastric pain radiating to the back”, not “stomach pain”), specify the output format, and refine the prompt over several tries. Those rules work.
Correctness is limited by two things a prompt does not change: what the model knows, and what the case contains. A structured prompt cannot put knowledge into the model that is not there. It cannot supply the piece of history you left out. What it can do is make the model show its reasoning, which means you can check it. That is the finding of a Stanford study that prompted GPT-4 to imitate several styles of clinical reasoning. Diagnostic accuracy did not change, but the reasoning became something a physician could read and judge (Savage et al., 2024). The value of a reasoning prompt is not that the answer gets better. It is that the answer becomes checkable.
There is more direct evidence that structure does not reliably improve correctness. Pharmacists at one hospital put 50 internal-medicine cases through ChatGPT and OpenEvidence, with and without a prompt template. The template made no significant difference to accuracy and completeness. The probability of an acceptable answer went from 0.54 to 0.60 for ChatGPT and from 0.64 to 0.52 for OpenEvidence (Yang et al., 2026). An orthopaedics study asked several models about guideline recommendations in several prompt styles. The best style depended on the model, and asking the same question five times gave answers whose agreement with each other ranged from none to near-perfect (Wang et al., 2024). In a low-back-pain study, targeted safety prompts improved the weaker model on every measure. They improved the stronger model on everything except safety (Luo et al., 2026). The pattern across all three is the same. Prompt effects are real, they differ from model to model, and they are unstable. “Better organised” is the only reliable effect.
12.3.2 The five-part prompt
Learn this as a list of parts, not as fixed wording. Five parts; the fifth is the one people skip.
| Element | What you write | Example |
|---|---|---|
| Role | Who the model is helping | “You are assisting a physician preparing for a case discussion.” |
| Context | The case, in clinical language | “Epigastric pain radiating to the back,” not “stomach pain.” Vitals, key negatives, medications. |
| Task | Exactly what you want | “Give a ranked differential with the three findings that most support and most oppose each entry.” |
| Format | The shape of the answer | “Table. One row per diagnosis.” |
| Uncertainty | What to do with gaps | “Where the case lacks information you would need, say what is missing instead of assuming it.” |
The ten rules in the source deck fit onto this table. Objective is Task. Context and precise language are Context. Output format is Format. The rest are ways of making Task and Format more precise: refine over several tries, break a big question into smaller ones, mark the case off from the instructions with clear separators, and show the model one or two worked examples of what you want (the deck calls this “few-shot”). None of the ten is the fifth row. Models default to producing a complete-looking answer from incomplete information, because a complete-looking answer is what the text they learned from looks like. Asking them to name the gap instead is the single most useful thing a clinician can add to a prompt. In the exercise you will see what happens with and without it.
the critical appraisal chapter reports a finding that also applies in this chapter. The same tool made up references ten times more often when a question was phrased the way a patient talks than the way a clinician writes (McLaughlin et al., 2026). Linguists call this register: the way a particular group of people talks. It is a property of the words, not of the credentials of the person typing. The Context row of Table 12.1 is where you control it. Writing the case the way you would present it on rounds is not done to impress. It moves your prompt into the kind of language that got the better answers.
12.3.3 Why the neat wrong answer is worse
Go back to Sam’s phone. A negative D-dimer rules out pulmonary embolism only when the pretest probability is low. Pretest probability is how likely the diagnosis is before you order any test. The original emergency-department study validated that strategy in patients whose clinical score put them in the low-probability group (Wells et al., 2001). Sam’s patient has had surgery in the last four weeks, has a heart rate over 100, and has a swollen calf. On the Wells score, which MDCalc will calculate for you, she is not low probability. The right next step is imaging, not a D-dimer, because a positive D-dimer would only delay the scan. Both answers got this wrong. Only one of them put it in a column headed “Next step” beside a column of supporting findings.
That is how the neat answer causes harm. Anchoring is the habit of settling on the first plausible answer and reading everything afterwards as confirmation. Automation bias is the same habit with a machine as the source: trusting a machine’s suggestion more than it deserves, and more the more authoritative the output looks. A table with “Supports” and “Opposes” columns invites you to do your own pattern-matching. You read the row, you agree with the findings, and you accept the next step without a separate check. The paragraph version at least reads as one opinion.
The published studies of fluent wrong answers say the same thing without the table (the appraisal chapter uses the same two studies for the claim that fails in fluent prose). When 33 physicians across 17 specialties graded chatbot answers to 284 of their own questions, the median score was almost completely correct. But 36 answers were mostly or completely wrong, in the same confident prose as the rest (Goodman et al., 2023). When oncologists checked chatbot cancer-treatment recommendations against guidelines, about a third of outputs mixed at least one non-guideline recommendation in among correct ones (Chen et al., 2023). The authors’ phrase, mixed incorrect recommendations among correct ones, is exactly the structured-prompt failure. Structure makes the mixing tidier.
A structured prompt improves how an answer is organised. It does not reliably improve whether the answer is right, and a well-organised wrong answer is the one you are most likely to sign.
12.3.4 What the trials show about physicians plus the tool
Here is the finding to compare with every enthusiastic claim about prompting. The centaur-or-cyborg chapter is based on these two trials and on what to do about them; this is the short version. In a randomized trial, 50 physicians worked through written cases with either their usual resources or their usual resources plus a language model. The model group scored a median 76 percent, the usual-resources group 74 percent. After adjustment the difference was 2 points, with a 95 percent confidence interval from minus 4 to plus 8. (A confidence interval is the range of differences the data are consistent with. One that crosses zero means the trial could not tell the groups apart.) In a secondary comparison the authors call exploratory, the model on its own scored 16 points higher than the usual-resources group (Goh et al., 2024). Figure 12.2 redraws the two estimates.
The tool was better than the physicians. The physicians using the tool were not better than the physicians without it. Something in how clinicians used the output seems to have lost the benefit. The trial could not say what. The likely candidates are settling on their own first impression, skimming, or asking the model to confirm rather than to challenge. The same group ran a second trial a year later on management rather than diagnosis, with 92 physicians. This time the model group did score higher, by 6.5 points (95 percent CI 2.7 to 10.2). But the model alone matched the physicians who had it, and the physicians with the model spent two minutes longer per case (Goh et al., 2025). Two trials, one direction. The reading the authors favour, and the one this chapter takes, is that the weak point is how the physician uses the answer, not what the model can do. Both trials used written cases, not real patients, and the authors say so. A trial in NEJM AI randomized 44 physicians who had just finished a 20-hour AI-literacy course. When the model’s suggestion was wrong, their diagnostic scores fell by about 14 points, so the training did not remove the automation bias (Qazi et al., 2026) VERIFY: the numbers come from the preprint listing and a search summary; the NEJM AI page blocked the tool on 2026-09-08.
The honest summary for a fourth-year: the prompting skill is real and cheap. The evidence that it makes you a better diagnostician is not there yet. What evidence exists says the gain, when there is one, comes from how you use the answer, not from how you asked for it. the centaur-or-cyborg chapter takes up what “how you use the answer” means in practice.
12.4 Watch or listen
Podcast. NEJM AI Grand Rounds, “Medicine, Machines, and Magic: Dr. Jonathan Chen on Medical AI” (15 October 2025; 48 min). https://ai-podcast.nejm.org/e/medicine-machines-and-magic-dr-jonathan-chen-on-medical-ai/
This episode is also assigned in the centaur-or-cyborg chapter, which assigns it as its main episode. If you have already listened, the source deck below is enough for this chapter; if not, listen here and count it for both.
Chen is the senior author of both randomized trials above. Listen for the difference between his 2017 essay on “inflated expectations” and the trial results, and for his account of why physicians with the model did not beat physicians without it. The first twenty minutes are enough for the exercise; the rest is optional. When he talks about what the model got right, ask what a physician would have had to do to keep that gain.
The source deck. Prompt engineering, about fifteen slides, ten rules with clinical examples. Ten minutes. It is where the Role, Context, Task, and Format rows of Table 12.1 come from. Read it as a set of techniques, and notice that none of the ten slides is the Uncertainty row.
Alternative episode. NEJM AI Grand Rounds, “AI and the Evolution of Medical Thought with Dr. Adam Rodman” (19 June 2024; 53 min), the other senior author, on clinical reasoning and why he thinks the models reason at all. Also the alternative in the centaur-or-cyborg chapter.
12.5 Do
The four cases below are fiction, written for this course. Do not substitute a patient from your sub-internship, however well you have de-identified them in your head. What Sam did at 2:40 a.m., typing a real patient’s presentation into a personal phone, is the thing this callout forbids, and the patient did not know it happened. Consumer chatbots have no business associate agreement, the contract that makes a vendor legally responsible for protecting patient data. Your university ChatGPT Edu and Copilot accounts have one when you sign in with university credentials, but it covers patient care. These cases are coursework, so they stay fictional. A tool your hospital runs inside the EHR is a different question, and the ethics and regulation chapter takes it up.
What counts as patient information, and why taking out the name is not enough: the patient information page.
You will run one case twice, score both answers, and then check two claims from the better one against a source. Open a blank document first; the record you fill in at Step 6 is what you keep and what you bring to the live session. Use the same chatbot for both runs, in two separate conversations, so the second run cannot see the first.
12.5.1 Step 1. Pick a case (2 minutes)
Choose by the specialty you plan to enter, or the one whose trap you least expect to catch. Each case has one deliberate trap. The traps are of different kinds, and you will not be told which kind yours is until Step 4.
A 31-year-old woman presents with two days of right-sided chest pain that is sharp and worse on inspiration, and shortness of breath when climbing stairs. Heart rate 102, blood pressure 118/74, respiratory rate 20, oxygen saturation 96 percent on room air, temperature 37.1 °C. Lungs are clear. No cough, no haemoptysis, no leg swelling or calf tenderness. She describes herself as otherwise healthy and takes no regular medication that she mentions. She has not travelled recently.
An 82-year-old woman is brought from an assisted-living facility with two days of fluctuating confusion; staff say she is “not herself” and was found wandering at night. Temperature 36.9 °C, heart rate 84, blood pressure 138/78, oxygen saturation 97 percent. She is inattentive and disoriented to time but has no focal neurological findings. No dysuria, frequency, or flank pain reported; no cough. Urine dipstick: leukocyte esterase positive, nitrite positive. White cell count 8.1. Glucose, sodium, and calcium are normal. Medications: amlodipine 5 mg daily, atorvastatin 20 mg nightly, sertraline 50 mg daily, acetaminophen as needed, and oxybutynin 5 mg twice daily, started twelve days ago for urinary urgency.
A 66-year-old man attends pre-operative clinic ten days before elective repair of a right inguinal hernia. Past history: paroxysmal atrial fibrillation, hyperlipidaemia, type 2 diabetes, and “bronchitis” for which an urgent-care clinic started an antibiotic last week. Examination is unremarkable apart from the reducible hernia. Medications: apixaban 5 mg twice daily, amiodarone 200 mg daily, simvastatin 40 mg nightly, metformin 1000 mg twice daily, tamsulosin 0.4 mg daily, clarithromycin 500 mg twice daily (day 4 of 7). Question for the model: what needs attention before this operation?
A 23-year-old man is brought in by his roommate. For ten days he has slept two to three hours a night without feeling tired, talks rapidly and jumps between topics, and has spent about $4,000 on equipment for a business he describes in grand terms. For the last three days he has heard a voice commenting on what he is doing. Three weeks ago he was started on prednisone 60 mg daily for a severe contact dermatitis and is now on a taper at 30 mg. No prior psychiatric history. A maternal uncle has bipolar disorder. He uses cannabis most weekends. Vital signs and neurological examination are normal.
12.5.2 Step 2. The bare prompt (3 minutes)
Paste the case into a new conversation with nothing but this in front of it:
What’s the diagnosis and what should I do?
Copy the whole answer, word for word, into your document under Bare. If the tool replies with questions instead of a plan, do not answer them. The questions are the answer; copy them and score them like any other reply.
12.5.3 Step 3. The structured prompt (5 minutes)
Open a second, new conversation. Build the prompt from Table 12.1. Here is the template; fill in the case where marked.
You are assisting a physician preparing to present this case at morning report. Case:
<paste the case, unchanged>. Give a ranked differential diagnosis with, for each entry, the two or three findings that most support it and the two or three that most oppose it, then the single best next step. Format as a table, one row per diagnosis, followed by one paragraph on management. Where the case lacks information you would need to decide, say what is missing instead of assuming it.
Copy the answer word for word under Structured. Note the tool and, if it shows one, the model version. These outputs change from month to month; the point is the comparison, not the numbers.
12.5.4 Step 4. Score both (10 minutes)
Now open the trap for your case, below, and score each answer twice: once for organisation, once for correctness on the trap.
Case A: a critical omission. The case never says whether she is pregnant, whether she takes hormonal contraception or other oestrogen, whether she has been immobile, or whether she or her family have had a clot before. Every one of those changes the pretest probability of pulmonary embolism, and no score can be computed without them. The trap is whether the model asks for them or assumes them. An answer that computes a Wells score, or declares her “low risk”, from what is given has invented the missing history. Full marks require the answer to name at least two of the missing items as things it needs before choosing between D-dimer and imaging.
Case B: the common answer is wrong. The pattern-match is “urinary tract infection causing delirium; start antibiotics.” But bacteria in the urine without urinary symptoms is common in older women, especially in care facilities. The Infectious Diseases Society of America’s 2019 guideline specifically addresses older adults with delirium and no local urinary symptoms: look for other causes and observe rather than treat the urine (Nicolle et al., 2019). The clue that separates the diagnoses is in the medication list. Oxybutynin is a strongly anticholinergic drug. The American Geriatrics Society Beers Criteria list it as one to avoid in older adults, and it was started twelve days before the confusion (“American Geriatrics Society 2023 Updated AGS Beers Criteria® for Potentially Inappropriate Medication Use in Older Adults.” 2023). Full marks require the answer to put the new anticholinergic ahead of, or at least beside, the dipstick, and to say that the dipstick alone does not establish infection.
Case C: interactions hidden in the list. Clarithromycin strongly blocks the liver enzyme that clears simvastatin. The simvastatin label lists clarithromycin as a contraindication, and it caps the dose at 20 mg a day with amiodarone, so 40 mg is over the limit (DailyMed label). The American Heart Association’s statement on statin interactions says the same (Wiggins et al., 2016). Clarithromycin also raises apixaban exposure through the same pathway, and amiodarone with clarithromycin puts two QT-prolonging drugs together. The perioperative apixaban plan and metformin around surgery are the expected content. The trap is whether the model points out the interactions without being asked. Full marks require the statin-macrolide interaction to be named and acted on.
Case D: genuinely ambiguous. The correct answer is that two diagnoses remain: corticosteroid-induced mania with psychotic features, and a first episode of bipolar I disorder with psychotic features, with a cannabis contribution to exclude. What separates them is time. Symptoms that began after the steroid started and resolve after it stops point one way; symptoms that persist point the other. The family history raises the baseline risk without settling it. Psychiatric effects of corticosteroids are common, dose-related, usually early in a course, and usually resolve with reduction or withdrawal (Warrington & Bostwick, 2006). A model that declares “bipolar I disorder” as the diagnosis has decided too early. Full marks require both diagnoses to stay on the list with the distinguishing test named, and a safety plan that does not depend on which it is.
Organisation, 0 to 5. One point for each: ranked; findings for and against each entry; an explicit next step; a format you could read aloud on rounds; a statement of what information is missing.
Correctness on the trap, 2/1/0.
| Score | Meaning |
|---|---|
| 2 | The trap is caught and it changes the plan |
| 1 | The trap is mentioned but hidden, softened, or does not change the plan |
| 0 | The trap is missed, or the answer invents around it |
If either answer gives references, score each on the critical appraisal chapter’s 2/1/0 citation scale (Table 5.4: accurate, misrepresented, fabricated) by searching PubMed, and put the scores on the record. Do not invent a separate scale for them.
12.5.5 Step 5. Check two claims (10 minutes)
Take the better of your two answers, by whichever score you trust more, and pick two specific factual claims in it: a threshold, a first-line drug, a “rules out”, an interaction, a dose. Check each against one source, in a browser, in ten minutes total:
- A calculator or score: MDCalc, for anything with a score or a threshold.
- A guideline: search PubMed with
<condition> AND (guideline[Publication Type] OR "practice guideline"), filtered to the last ten years, and take the society’s guideline, not a commentary on it. - A drug label: DailyMed (https://dailymed.nlm.nih.gov), the FDA’s label repository, for any interaction, contraindication, or dose.
Record each as supported, contradicted, or not addressed, with the source. “Not addressed” is a real result. A claim the sources do not make is a claim you should not be repeating either.
12.5.6 Step 6. Record it (bonus 5 minutes if you have them)
Tool: ____________________ Model version if shown: ____________
Case: A / B / C / D
Bare Structured
Organisation (0–5) ___ ___
Trap (2/1/0) ___ ___
References scored ___ ___ (2/1/0 each, if any were given)
Claim 1: ________________________________________________________
Source: ______________________ Supported / Contradicted / Not addressed
Claim 2: ________________________________________________________
Source: ______________________ Supported / Contradicted / Not addressed
In one sentence: what did structure change, and what did it not change?
______________________________________________________________________
If you have five more minutes: rerun the structured prompt with the last
sentence (the "say what is missing" line) deleted. Trap score: ___.
What disappeared from the answer? ______________________________________
The last line of the record is the finding. If your organisation score rose and your trap score did not, you have reproduced the main claim of this chapter on your own screen. If both rose, good; write down which sentence of the prompt did it, because it is nearly always the last one. If the bare answer caught the trap and the structured one did not, that is the best result of all. Keep the transcripts, because you now have an example of structure making a wrong answer look right.
You have two observations. A published comparison would run every case through every model several times, because the same prompt does not give the same answer twice (Wang et al., 2024). You are reproducing the design of those studies, not their result. If yours comes out the other way from the studies above, that is one observation about tools that change from month to month, not a refutation.
12.6 Check
Score yourself against this before moving on.
| You have | Meets | Falls short |
|---|---|---|
| A structured prompt (objective 1) | All five rows of Table 12.1 are present, and you can say which sentence is the Uncertainty row | Four rows, or the fifth is “be accurate” |
| Two scored answers (objective 2) | Organisation and trap are scored separately, and your one-sentence finding names what changed and what did not | One overall impression of which answer was “better” |
| Two checked claims (objective 3) | Each has a named source and a verdict; at least one was checked against a calculator, guideline, or label rather than another chatbot | “It matched what I remembered” |
| The one sentence (objective 4) | You can say why a wrong answer in a table is more dangerous than one in a paragraph, and name the habit (anchoring or automation bias) it exploits | “You should always double-check” |
| The two trials (objective 5) | You can state the direction of both Goh trials, the model-alone result, and what that implies about where the weak point is | “Studies show AI helps doctors” or “studies show it doesn’t” |
| The trap | You can say which kind of trap your case had (omission, common-answer, hidden interaction, ambiguity) and what the model did with it | You scored the answers without opening the trap |
12.7 What this changes for you
What does this change about how you’ll practice? Write it down before you close the chapter. Two specifics.
First, it is 2:40 a.m. again and Sam is holding out the phone. What do you say about the D-dimer, and which one sentence would you have added to Sam’s carefully structured prompt? Then say what you will check before any chatbot-drafted plan goes into a note with your name on it, and which of the three sources in Step 5 you would use for it.
Second, the trials. If the model alone can outscore you and you-plus-the-model cannot, what will you do differently with an answer next month? Not “be careful”. Name the action: ask it to argue against you, ask what is missing, check the one claim the plan depends on. Pick one and write it down.
12.8 Summary
- A prompt reliably changes how an answer is organised. It does not reliably change whether the answer is correct; prompt effects on accuracy are real, differ from model to model, and are unstable.
- Role, Context, Task, Format make the answer readable. “Say what is missing” is the only element that acts on the gap between the case and the answer, and it is the one people leave out.
- Writing the case the way a clinician would present it is not decoration. The same tool behaves differently in a patient’s words than in a clinician’s.
- A well-organised wrong answer is more dangerous than a badly organised one, because the organisation leads you to use your own reasoning to confirm it.
- In two randomized trials the model alone matched or beat the physicians who had it. The weak point is how the physician uses the answer, not what the model can do.
- Check the claim the plan depends on against a calculator, a guideline, or a label. Then go back and tell Sam about the D-dimer.
12.9 Go deeper
Papers
- The randomized trials of physicians with a model: diagnosis (Goh et al., 2024) and management (Goh et al., 2025). Read the discussion sections for the authors’ own account of why the benefit did not appear. Both are the centaur-or-cyborg chapter’s main papers.
- The pharmacy study in which a prompt template changed nothing that mattered (Yang et al., 2026), and the guideline-consistency study that found prompt effects vary by model and the same question gives different answers on repeat (Wang et al., 2024).
- Diagnostic-reasoning prompts as a route to checkable answers (Savage et al., 2024), and the psychiatry group’s prompt built from reusable parts, where structure did improve the weaker model’s accuracy at pulling facts out of text (Verhees et al., 2026): the counter-example that limits this chapter’s claim.
- The student error-catching study behind the 56 percent (Waldock et al., 2025) (also in the appraisal chapter), and the published, reusable case-based LLM workshop from Harvard (Choy et al., 2026); this chapter’s Do is a shorter version of it.
- A tutorial on prompt engineering written for health professionals, for the ten-rules material in article form (Meskó, 2023).
Talks
- Prompt engineering, the source deck: ten rules, clinical examples throughout.
- AI in medical education, for the attribution question that arises the first time a chatbot’s table ends up in something you hand in. Also in the centaur-or-cyborg chapter.