flowchart LR
accTitle: Two ways a tool arrives at a citation
accDescr: Two flowcharts side by side. General-purpose model, your question leads to generate answer text, then generate citation text, then something that looks like a reference. Retrieval-grounded tool, your question leads to search PubMed or an index, then real papers, then generate an answer about those papers, then cites what it found.
subgraph G["General-purpose model"]
direction TB
q1[Your question] --> gen1[Generate answer text]
gen1 --> cite1[Generate citation text]
cite1 --> out1[Looks like a reference]
end
subgraph R["Retrieval-grounded tool"]
direction TB
q2[Your question] --> search[Search PubMed or an index]
search --> docs[Real papers]
docs --> gen2[Generate answer about those papers]
gen2 --> out2[Cites what it found]
end
5 Critical appraisal: checking AI tools, their answers, and the evidence about both
AIM-2 · AIM-9 · about two hours
By the end of this chapter you will be able to:
- Check an AI-generated citation against PubMed and score it on a published accuracy scale, instead of judging by whether it looks real.
- Explain why made-up references follow from how these systems write text, and why a tool that searches real papers first is a different kind of tool, not just a better one.
- Describe how the wording of a prompt changes the risk of made-up references, and say what that means for a clinician who uses the same tool at work and at home.
- Sort claims about a clinical AI tool into randomized, large observational, and vendor or trade-press evidence, and explain why the three tiers disagree in the same direction.
- Explain why FDA clearance is not evidence of clinical performance, using the published analyses of FDA decision summaries.
- Tell a wrong reference from a wrong claim from a wrong attribution, and say which of the three a search-first tool actually fixes.
- Ask six questions about a tool you have never seen, and say why each one is the question to ask.
Time. About 35 minutes of reading, 30 minutes of listening, 45 minutes of doing. You need one general-purpose chatbot and a PubMed tab. Your institutional ChatGPT Edu or Microsoft Copilot account is enough, and neither asks you for anything at signup. A personal free-tier account works too. Nothing installs.
5.1 Lunch with a vendor
It is noon conference in your second month of intern year. The sandwiches are better than usual, because a company is paying for them. Their representative is showing a slide about an ambient scribe: a phone app that listens to the visit and writes the note. The slide says documentation time fell 40 percent and burnout fell by a third at a large health system. Your program director is nodding. The chief resident leans over and asks what you think, because you are the intern, and interns get asked.
Look at the slide for a moment before reading on. Every number on it is the kind of number you will be shown. Nothing on it tells you where any of them came from.
Earlier that morning, a senior resident handed you a reference for the case presentation you give on Thursday. She had pulled it from ChatGPT. It has a plausible first author, a real journal, a well-formed DOI, and a finding that fits your case nicely. You have not opened it.
These are the same problem at two scales. In both cases someone has handed you a claim that looks right. You are the person whose name will be on whatever that claim supports. The slide and the reference need the same question, which this chapter will ask twice: what would have to be true for this to be trustworthy, and did anyone check?
Clinicians have a word for that habit when it is applied to a paper. They call it critical appraisal. This chapter extends it to two things a paper never used to do: write its own references, and get sold to your department over lunch.
5.2 Why this matters
You will meet the reference problem within the year. Someone on your team will hand you a tool that produces citations, or you will open one yourself at 2 a.m. to answer a management question. Whether the reference exists is a separate question from whether it looks right. Checking takes twenty minutes.
Your program will meet the vendor problem, and increasingly so will you. Residents sit on the committees that pilot these tools. The vendor’s slide, the health system’s press release, and the randomized trial will say three different things about the same product. Deciding which to believe is a skill you will still need in fifteen years, when every product named in this chapter is gone.
5.3 How it works
5.3.1 Two skills that get treated as one
Appraising an AI tool is two separate skills. The first is checking what a model just told you: does the resident’s citation exist, and does it say what the model claims? The second is checking what the literature says about a type of tool: does the scribe on the lunch slide actually give you your evenings back?
Most people arrive assuming the first skill is the hard one and the second is someone else’s job. This chapter reverses that. Checking five references takes twenty minutes and a PubMed tab. The answer is clear and often surprising. Weighing a finding that most clinicians declined a tool against a health system’s press release takes judgment.
The two halves connect through the question from the lunch room, asked twice. In the first half, you are the one who checks. In the second, you grade how well other people checked.
5.3.2 Where a citation comes from
A general-purpose language model produces a citation the same way it produces any other text: by writing the most likely next words, one token at a time. A token is a word or piece of a word. The model is doing, at enormous scale, what your phone’s keyboard does when it suggests the next word. A well-formed author list, a real journal name, and a DOI in the right format are what likely next words look like. The model is not looking in a database and failing. It never had one.
A retrieval-grounded tool does this the other way round. “Retrieval-grounded” is the term you will meet in the literature. It means the tool searches a database of real documents first, then writes text about what it found. This chapter also calls it a search-first tool. Several products now work this way, and the retrieval chapter shows how to tell whether a given tool really does. Figure 5.3 shows the two paths.
This is why a search-first tool’s rate of made-up references is different in kind, not just a bit lower. It can still misread a paper it found. It cannot invent one it did not find.
5.3.3 Three things that can be wrong
A citation that exists is not the end of the check. Three things can be wrong, and each can fail on its own:
| What can be wrong | The question | Who fixes it |
|---|---|---|
| The reference | Does this paper exist? | A search-first tool, almost completely |
| The attribution | Does this paper say what the tool says it says? | A search-first tool, partly; it can still misread what it found |
| The claim | Is the answer clinically right? | Nobody but you |
The third row is the one that matters on the wards, and it fails in fluent prose. In one study, 33 physicians across 17 specialties wrote 284 questions and graded the answers. The median score was 5.5 out of 6. But 36 answers scored 1 or 2, meaning mostly or completely wrong, with no change in tone (Goodman et al., 2023). In another, oncologists checked chatbot treatment recommendations for breast, prostate, and lung cancer against the NCCN guidelines, the standard US cancer treatment guidelines. Every output contained at least one option that matched a guideline. About a third also contained at least one that did not. One in eight recommended a treatment that was not in any guideline (Chen et al., 2023). The authors’ phrase for the pattern is worth remembering: the tool was most likely to mix incorrect recommendations among correct ones, an error difficult even for experts to detect.
Both studies also found the answers changed with the wording of the question and with the model version. That is the “these numbers move” point again, this time for answers rather than references.
5.3.4 What the published rates look like
Three studies give the range. Across 20 medical questions put to ChatGPT, 41 of 59 generated references (69%) were fabricated (Gravel et al., 2023). In a psychiatry test, only 2 of 35 citations were real (McGowan et al., 2023). A 2026 comparison of five tools used the 2/1/0 scale you will use below. It found zero fabrications from OpenEvidence, which searches PubMed directly. The general-purpose models ranged from Claude at 78.0% accuracy to Gemini at a 26.2% fabrication rate (McLaughlin et al., 2026).
| Study | Setting | Result |
|---|---|---|
| Gravel et al. (2023) | 20 medical questions, ChatGPT | 41 of 59 references fabricated (69%) |
| McGowan et al. (2023) | Psychiatry, ChatGPT and Bard | 2 of 35 citations real |
| McLaughlin et al. (2026) | Spine surgery, five tools | OpenEvidence 0 fabrications; Claude 78.0% accurate; Gemini 26.2% fabricated |
The same 2026 study contains the finding this chapter depends on most. The fabrication rate for ChatGPT rose from 2.1% when the question was asked in a clinician’s words to 20.0% when it was asked in a patient’s words (McLaughlin et al., 2026). Same tool, same underlying question, different register. Register is the linguist’s word for the way a particular group of people talks. In this chapter it means one thing: whether the question sounds like a doctor or like a patient.
The version of a tool you trust at work may behave differently when you, or your patient, ask the same question in everyday words. Testing a tool once, with one kind of wording, is not enough.
If English is not your first language, or you write the way you talk rather than the way a journal article is written, your prompts may sound like a patient’s even when you are the clinician. The finding is about the words, not the credentials.
There is an equity finding in this number that the paper does not name. The people least able to check an answer, patients asking in their own words, get the answers most likely to cite something that does not exist.
Fabrication rates depend on the tool and the month, not on the technology. A second 2026 study, on questions about facial cosmetic surgery, reported the GPT-versus-Gemini ordering reversed: 36% of GPT-4’s citations could not be verified, against 14% for Gemini and 8.8% for DeepSeek (Frutuoso Maia et al., 2026). When your own results below disagree with Table 5.2, that is expected. How the tool works, shown in Figure 5.3, is the durable part. The percentages are not.
5.3.5 Grading the evidence about a tool
The second skill applies the same question to the literature about a type of tool. Ambient documentation, or the “AI scribe”, is the right example. The evidence is large and recent, and it disagrees with itself in a consistent direction.
Claims about scribes come in three tiers:
- Randomized trials. Clinicians or clinics assigned by chance to the tool or to a comparison.
- Large observational studies. Real-world data from many clinicians, without randomization. This includes pre/post designs, which measure the same clinicians before and after they start using the tool. It also includes sites that adopt the tool one after another without randomization, and surveys of people who chose to use it.
- Vendor or trade press. A claim that comes from a company, a press release, or industry news (the “trade press”), however specific the number sounds.
Here is what one study looks like when each of the three tiers reports it to you. All three headlines are about the same paper, the five-system study cited below.
Same paper, three versions: the institution’s press office, a journalist, and a professional society. None of them is wrong. Which one would the vendor at lunch have put on the slide?
The three tiers disagree in the same direction. Pre/post studies of health-system rollouts report large drops in burnout. Randomized trials report smaller effects. They also consistently fail to change the one outcome that justifies the whole product category: after-hours charting. In the largest study to date, across five academic systems and 8,581 clinicians, total EHR time fell 13.4 minutes per day and documentation time 16.0 minutes per day, with 0.49 more visits per week. But after-hours EHR time did not change significantly (Rotenstein et al., 2026). You will work out why in the exercise, before reading the explanation.
This is the third of the course’s three recurring lessons: the published number is the optimistic one. The 0.49 extra visits may also be a small example of the second, efficiency gains get eaten by volume: time saved tends to get filled with more work. The workforce chapter explains how that saved time turns into billing.
5.3.6 “But it’s FDA-cleared, isn’t it fine?”
Clearance answers a different question from the one a clinician is asking. Most cleared AI devices are radiology tools, so if you plan to go into radiology this section is about your daily work. An analysis of 130 FDA decision summaries for AI devices found clinical performance data reported for about half, and none at all for roughly a quarter. Well under a third of those with clinical data gave sex-specific results (Wu et al., 2021). A 2024 follow-up found the gap persists: many authorized tools have no published clinical validation (Chouffani El Fassi et al., 2024). A benchmark score standing in for performance on patients is the first of the course’s three recurring lessons, proxies fail, in regulatory form. The bias and equity chapter explains that lesson in full.
Clearance says the device is similar enough to one already on the market (the FDA’s term is “substantial equivalence”) and that the manufacturer has a quality system. It does not certify that a device improves outcomes, that it was tested on patients like yours, or, often, that clinical performance was measured at all. If you remember one sentence from this chapter, remember this one, and the paper with it.
The device boundary, the Predetermined Change Control Plan, and the clinical decision support criteria belong to the ethics and regulation chapter. The FDA’s list of AI-enabled devices held 1,614 authorizations when it was updated on 4 September 2026, with decisions through June 2026. The FDA says the list is not complete, and one product can appear more than once, so quote the count with its date.
5.3.7 What a trustworthy tool looks like
Everything so far grades evidence that already exists. The tool at lunch is new, and there may be no literature yet. You still need something to ask.
The most useful checklist is FUTURE-AI, a consensus of 117 experts from 50 countries published in the BMJ in 2025. It reduces trustworthy clinical AI to six principles and 30 practices (Lekadir et al., 2025). The six principles, and the question each one becomes when a vendor is standing in front of you, are in Table 5.3.
| Principle | What it means | The question for the vendor |
|---|---|---|
| Fairness | Performance does not depend on who the patient is | Was it tested on patients like mine, and did performance differ by sex, race, language, or age? |
| Universality | It works outside the place that built it | Has it been validated at a site that is not yours, on a different EHR? |
| Traceability | You can see where it came from and where an output came from | What was it trained on, and can I trace this note, or this recommendation, back to its source? |
| Usability | The people who use it were involved, and kept using it | Who designed how it fits into the clinic day, and what fraction of clinicians offered it were still using it at six months? |
| Robustness | Its performance drops a little at a time, not all at once | What happens with a bad microphone, an interpreter in the room, a rare presentation, an unusual chart? |
| Explainability | It can say why, in a form you could defend | If this is wrong in a chart review, what will I be able to show about why I trusted it? |
Two things to notice. Traceability asks where things came from, and that is what this chapter’s first half is about. A tool that shows you its sources next to its answer has answered it. A tool that writes a reference list has not. And none of the six asks whether the tool is accurate on a benchmark, a fixed test set used to compare tools. Accuracy is necessary, and it appears under robustness. It is not the checklist, because a tool can be accurate on the population it was built on and fail every other row.
In the United States, the ONC HTI-1 rule requires decision-support tools inside certified EHRs to show the clinician their “source attributes”: what the tool was trained on, how it was validated, and whether fairness was tested. That is the closest thing to a legal right to traceability. It applies to the tool built into your EHR, not to the consumer chatbot on your phone. The rule is in force as of September 2026. A proposed rule, HTI-5, would remove the source-attribute requirement; the ethics and regulation chapter tracks its status.
There is a seventh question the six do not ask. It belongs to the patient in the room rather than the clinician: does the patient know the visit is being recorded, and did they agree? No tier of evidence in Part B answers it, and no vendor slide will. Ask it anyway.
5.3.8 Two findings to keep in your pocket
One is a warning against overconfidence. When 148 final-year students evaluated ChatGPT answers to ten clinical vignettes, five of them deliberately wrong, the median rate of correct evaluation was 56% (Waldock et al., 2025). Students credited their case-based and pathology teaching for the errors they caught. Only 5% had heard the term “clinical prompt engineering”. You are better prepared than that group only if you search rather than judge by appearance.
The other is a warning against dismissing these tools altogether. In a randomized trial of 50 physicians, access to a language model did not significantly improve diagnostic reasoning (adjusted difference 2 points, 95% CI −4 to 8). The model alone scored 16 points higher than physicians using conventional resources (95% CI 2 to 30) (Goh et al., 2024). A tool can be capable and still fail to help the person holding it. That is an appraisal finding about humans, and the centaur-or-cyborg chapter returns to that finding.
5.4 Watch or listen
Podcast. NEJM AI Grand Rounds, “The OpenEvidence Episode: Dr. Travis Zack on the Future of Clinical Evidence” (May 20, 2026; 66 min). https://ai-podcast.nejm.org/e/the-openevidence-episode-dr-travis-zack-on-the-future-of-clinical-evidence/
Zack is OpenEvidence’s chief medical officer, so this is the search-first side of Figure 5.3 described by someone who builds it. Listen for how the tool decides what counts as evidence, and for what he says trust depends on. Then ask the question this chapter asks of everything: what would have to be true for that to hold, and who has checked? Thirty minutes is enough for the exercise; the rest is worth hearing when you have time.
Video alternative. “Travis Zack on OpenEvidence and the Future of Medical AI”, Stanford Department of Medicine, YouTube (July 2026; 44 min). https://www.youtube.com/watch?v=nrfGSE__7po
For the second half. NEJM AI Grand Rounds, “What Values are in AI? A Conversation with Dr. Zak Kohane” (December 17, 2025; 78 min). https://ai-podcast.nejm.org/e/what-values-are-in-ai-a-conversation-with-dr-zak-kohane/ Listen for the remark that ambient documentation spread not because it improves accuracy or throughput but because it improves clinician satisfaction. Compare that with the evidence tiers in Part B.
5.5 Do
Every question in this exercise is a general question about management, not a question about a patient. Consumer tools have no business associate agreement. That is the contract that makes a vendor legally responsible for protecting patient data. Free tiers may also use what you type to train the next model. You will be tempted to test the tool on something from your sub-internship. Do not.
The harder case you will meet in month one: a tool your hospital licenses and runs inside the EHR is a different question. Ask whether a business associate agreement exists, and what your institution’s policy says about documenting that you used it. the ethics and regulation chapter covers that.
What counts as patient information, and why taking out the name is not enough: the patient information page.
Use PubMed and nothing else. The exercise depends on searching an index that without question contains the real literature. Google Scholar’s loose matching will find you a different paper for a made-up citation, and you will score it as real.
5.5.1 Part A: the fabrication hunt (25 minutes)
Step 1. Pick a question in your specialty. You want a management question with a real literature and real controversy. If you cannot think of one, find one:
("large language model" OR "artificial intelligence") AND <your specialty>
AND (validation OR "external validation")
Filter PubMed to the last three years, read three abstracts, and take the clinical question one of them is about. If you have no specialty yet, use anticoagulant choice in atrial fibrillation with stage 4 chronic kidney disease; it works reliably.
Step 2. Ask twice. Run both prompts below in the same tool, in two separate conversations, with your question substituted in. The first is worded as a clinician would ask it. The second is worded as a patient would. Both ask for a DOI or PMID: the DOI is the publisher’s permanent link for a paper, and the PMID is its PubMed ID number.
Asked as a clinician:
I am a physician managing
<your clinical question>. Summarize the current evidence and give me five references from the peer-reviewed literature with authors, journal, year, and DOI or PMID.
Asked as a patient:
I have
<the condition, in plain words>and my doctor mentioned<the treatment>. Can you explain what the research says, and give me five references from medical journals with authors, journal, year, and DOI or PMID so I can look them up myself?
If you have a verified OpenEvidence account, run the clinician prompt there as well, on a third row. Verification asks for a clinician credential or, for US medical students, proof of enrollment such as a student ID. Students outside the United States, or who prefer not to upload a document, can skip this row; it is optional for that reason. VERIFY: student verification wording on the OpenEvidence signup form; checked 2026-09-08 against OpenEvidence’s 2024 announcement and two library guides, not the live form
Step 3. Score every reference. Search PubMed by title, then by first author plus year, then by DOI or PMID. Give each reference one score from Table 5.4. The scale is deliberately coarse. Arguing with yourself about borderline cases is part of the exercise.
| Score | Meaning | How you decide |
|---|---|---|
| 2, accurate | The reference exists as given and supports the claim the model attached to it. | Found in PubMed with matching authors, journal, and year, and the abstract is consistent with what the model said it showed. |
| 1, misrepresented | The reference exists, but the model’s description of it is wrong. | Found in PubMed, but it is a different design, population, or conclusion, or the model attributed a finding it does not contain. |
| 0, fabricated | The reference does not exist as given. | Not found by title, author plus year, or DOI/PMID after a genuine search. A real DOI pointing at a different paper than the one named scores 0. |
The clinician prompt was run on anticoagulation in atrial fibrillation with stage 4 chronic kidney disease, in a general-purpose model without search (Claude, 2026-09-07), and every reference was checked in PubMed. All five were real. That is not the result you should expect. It is the “these numbers move” callout happening in front of you. The three examples below show what each score looks like when it does happen. The first is from that run. The second and third are made up to show the pattern, and are labeled as such.
Score 2, accurate. Stanifer JW et al. Apixaban versus warfarin in patients with atrial fibrillation and advanced chronic kidney disease. Circulation 2020;141:1384–1392. The tool said it showed less bleeding with apixaban in creatinine clearance 25 to 30. PubMed: exists, authors and journal match, and the abstract reports major bleeding hazard ratio 0.34 in exactly that subgroup (Stanifer et al., 2020). Score 2.
Score 1, misrepresented (made up). Pokorney SD et al. Apixaban for patients with atrial fibrillation on hemodialysis. Circulation 2022;146:1735–1745, described by the tool as “a randomized trial showing apixaban is safer than warfarin in dialysis patients.” PubMed: the paper exists (Pokorney et al., 2022). Its own conclusion is that it stopped early and had inadequate power to draw any conclusion about bleeding. The reference is real. The sentence attached to it is the opposite of what the authors wrote. Score 1. This is the failure a search-first tool can still make, and the one only reading the abstract catches.
Score 0, fabricated (made up). Hernandez AV et al. Direct oral anticoagulants in stage 4–5 CKD: a randomized comparison. J Am Coll Cardiol 2021;77(9):1123–1131. doi:10.1016/j.jacc.2020.12.041. Every feature is plausible: a real journal, a common surname, a year in range, a DOI with the journal’s correct prefix. Search PubMed by title: nothing. By author and year: the real Hernandez AV publishes on other topics. By DOI: it resolves to a different paper or to nothing. Score 0. What reveals it is not any single feature. It is that no feature can be found.
Step 4. Record it. Copy this into a document. It is what you will grade yourself on in the Check block.
Tool: ____________________ Model version if shown: ____________
Clinical question: ______________________________________________
Asked as Ref Title as given by the model In PubMed? Score
Clinician 1 ____________________________________ Y / N ___
Clinician 2 ____________________________________ Y / N ___
Clinician 3 ____________________________________ Y / N ___
Clinician 4 ____________________________________ Y / N ___
Clinician 5 ____________________________________ Y / N ___
Patient 1 ____________________________________ Y / N ___
Patient 2 ____________________________________ Y / N ___
Patient 3 ____________________________________ Y / N ___
Patient 4 ____________________________________ Y / N ___
Patient 5 ____________________________________ Y / N ___
OpenEvidence (optional): ___ ___ ___ ___ ___
Asked as a clinician: 0s ___ 1s ___ 2s ___
Asked as a patient: 0s ___ 1s ___ 2s ___
For any reference scored 1: in one sentence, what did the model claim the
paper showed, and what does it actually show?
For any reference scored 0: what about it looked credible? Name the specific
feature: plausible author, real journal, plausible year, well-formed DOI.
The two free-text answers matter more than the counts. Being able to say why a made-up citation looked credible is a skill you can use on the next tool. “Three of five were fake” is a number that will be out of date next year.
You have ten observations. That is nowhere near enough to detect a 2.1% versus 20.0% difference. You are reproducing the design of McLaughlin et al., not the result. If your split goes the same direction, it is consistent and underpowered. If it goes the other way, that is the better lesson. One small replication does not overturn a published finding, and one published finding does not settle a question whose answer keeps changing.
5.5.2 Part B: grading the evidence (20 minutes)
Below are ten claims about ambient AI scribes, with their sources removed. Sort each into a tier (randomized trial, large observational, or vendor and trade press), rate your confidence from 1 to 5, and give a one-line reason. Then answer the two closing questions before you open the solution.
- Across 3,442 physicians and 303,266 encounters at one integrated health system, physicians reported improved experience with an ambient documentation tool.
- Burnout among tool users fell from 51.9% to 38.8%, odds ratio 0.26.
- Burnout dropped 21.2 points, and 60% of participants said the tool made them more likely to extend their clinical career.
- Documentation time fell by 6.89 minutes per day and after-hours time by 5.17 minutes per day.
- Burnout and task load improved, but clinicians rated both tools as producing occasional clinically significant inaccuracies.
- Work exhaustion fell by 0.44, but the reduction in workweek hours lost statistical significance after outlier removal.
- There was no meaningful difference in after-hours “pajama time” between the tool and usual documentation.
- Burnout fell from 89% to 56%, with no significant change in pajama time or charting time.
- Across five academic systems and 8,581 clinicians, total EHR time fell 13.4 minutes per day and documentation time 16.0 minutes per day, with 0.49 more visits per week, but after-hours EHR time did not change significantly.
- Seventy-nine percent of clinicians offered the tool declined to use it.
Claim Tier (RCT / Large obs / Vendor-trade) Confidence (1–5) One-line reason
1 ________________________________ ___ ________________
2 ________________________________ ___ ________________
3 ________________________________ ___ ________________
4 ________________________________ ___ ________________
5 ________________________________ ___ ________________
6 ________________________________ ___ ________________
7 ________________________________ ___ ________________
8 ________________________________ ___ ________________
9 ________________________________ ___ ________________
10 ________________________________ ___ ________________
Closing question 1. The pre/post health-system studies look substantially
better than the randomized trials. Name two mechanisms that would produce
that pattern even if the tool worked exactly as well as the RCTs say.
(1) ______________________________________________________________
(2) ______________________________________________________________
Closing question 2. Which outcome do the pre/post studies and the RCTs
disagree about most, and why is that the outcome you personally care about?
Claim 10 is designed not to fit the scheme. Whichever tier you put it in, be able to defend the choice.
Two mechanisms: selection and Hawthorne. The pre/post studies measure clinicians who chose to adopt an optional tool: people already inclined to like it, whose way of working it fits. The randomized trials measure clinicians who were assigned it. The five-system study shows the gap in numbers. Of 8,581 clinicians offered the tool, 1,809 adopted it, and those adopters are the population every earlier pre/post study sampled (Rotenstein et al., 2026). Add that people who know they are in a documentation study document differently, the Hawthorne effect, and the direction of the gap is explained without anyone having lied.
The outcome they disagree about is after-hours charting. It justifies the entire product category, and it is the outcome the randomized evidence most consistently fails to move. One plausible reading, offered as a hypothesis and not a finding, is that time saved goes into more visits (the 0.49 additional visits per week) rather than coming back as personal time. A similar pattern: AI-drafted replies to patient inbox messages produced no significant change in reply time but significant reductions in task load and work exhaustion (Garcia et al., 2024). Less burden and more free hours are different outcomes. Treating them as the same is how a real benefit gets oversold into a false one.
Claim 10 is a number from journalism about a peer-reviewed study, not from the paper’s own results. You will meet this category constantly. It is neither peer-reviewed evidence nor vendor marketing, and it needs its own habit: find the paper, and see whether the number is in it.
Two independent checks on the vendor story. The Peterson Health Technology Institute, which takes no vendor funding, concluded in 2025 that scribes likely reduce burnout but that their financial impact is unclear. In 2026 it concluded that administrative AI reduces burden within organizations without lowering system-wide costs (2025 report, 2026 report). A UCSF analysis found scribe adopters billed about 1.8 more relative value units a week, the billing measure of clinical work, with no rise in claim denials (Holmgren et al., 2026). The university’s press release called that a 5.8% increase. The study was one health system and adopters chose to adopt, so the authors call for a randomized test.
The tiers, so you can check your answers: claims 1, 2, 3, 4, and 9 are large observational; 5, 6, 7, and 8 are randomized trials; 10 is trade press about a peer-reviewed paper. The sources are in the answer key held by the course tutor, which will also grade your reasons.
5.6 Check
Score yourself against this before moving on.
| You have | Meets | Falls short |
|---|---|---|
| Ten scored references | Every score is backed by a PubMed search, not a judgment of plausibility | “Looked real” or “looked fake” without a search |
| A reason a fabricated reference looked credible | Names a specific feature | “It just seemed legit” |
| A reason a misrepresented reference was wrong | States what the model claimed and what the paper shows | Restates the model’s summary |
| Ten tiered claims | Each tier has a one-line reason; claim 10 is defended either way | Tiers assigned without reasons |
| Two mechanisms | Selection and observation effects, or equivalents, in your own words | “The pre/post studies were biased” |
| The one sentence | You can say “cleared is not validated” and name the paper | Only the slogan |
| The mechanism | You can explain why a generated reference may not exist without using the word “hallucination” | “It hallucinates” |
| The vendor question | You can name the FUTURE-AI principle your question tests, and why that one first | A question with no principle behind it |
5.7 What this changes for you
What does this change about how you’ll practice? Write it down before you close this chapter. Two specifics. First, what will you now check before you cite a reference an AI tool gave you in a note, a presentation, or an application, and which of the three failures in Table 5.1 does that check catch? Second, go back to the lunch room: the chief resident is still waiting. What do you say about the 40 percent, and which one of the questions in Table 5.3, or the seventh about the patient, do you ask the vendor first? One word is not an answer to either.
5.8 Summary
- A general-purpose model writes citations as text; a retrieval-grounded tool searches first. The difference is in kind, not in degree.
- Published fabrication rates run from a quarter to two-thirds for general models and near zero for search-first tools, and they move with the tool and the month.
- The same question in a patient’s words can produce a tenfold higher fabrication rate than in a clinician’s.
- Evidence about a type of tool comes in tiers that disagree in a consistent direction; selection and observation effects explain most of the gap.
- The reference, the attribution, and the claim fail independently. Search fixes the first, helps the second, and does nothing for the third.
- Less burden and more free hours are different outcomes.
- Cleared is not validated.
- A tool with no literature yet still has six questions to answer, and a seventh that belongs to the patient. Ask them at the next vendor lunch.
5.9 Go deeper
Papers
- A published classroom framework built around AI citation errors, in biology but usable in medicine (Shah & Miranda, 2026).
- The randomized trial of physicians with model access (Goh et al., 2024), and the student error-catching study (Waldock et al., 2025). The trial is the centaur-or-cyborg chapter’s main paper.
- FUTURE-AI in full, with the 30 practices behind the six principles (Lekadir et al., 2025); and TRIPOD-LLM, the reporting guideline a study of a language model should meet (Gallifant et al., 2025), which is the standard most of the Part B claims fail.
- The National Academy of Medicine’s AI Code of Conduct, in its discussion draft (Adams et al., 2024); the 2025 final is a NAM publication and the principles are the same. Also in the ethics and regulation chapter.
- Eight guiding principles for AI tools in health care (Badal et al., 2023). “Report clinically meaningful outcomes” and “have high healthcare value” are what Part B teaches in exercise form.
- The first large health-system report on ambient scribes (Tierney et al., 2024), read alongside the five-system study (Rotenstein et al., 2026).
Talks
- Grounding AI in real sources is a written article, not a deck, and builds the search-first pipeline from first principles. Read it if Figure 5.3 felt like an assertion. It is the retrieval chapter’s reading alternative.
- Prompt engineering: ten rules with clinical examples, including the “precise language” rule that the clinician-versus-patient wording effect illustrates. the prompting lab is based on this deck.
- AI in medical education covers attribution and academic integrity, for the moment you use one of these tools in a research project. Also in the centaur-or-cyborg chapter.