5  Critical appraisal: checking AI tools, their answers, and the evidence about both

AIM-2 · AIM-9 · about two hours

NoteWhat you’ll learn

By the end of this chapter you will be able to:

  1. Check an AI-generated citation against PubMed and score it on a published accuracy scale, instead of judging by whether it looks real.
  2. Explain why made-up references follow from how these systems write text, and why a tool that searches real papers first is a different kind of tool, not just a better one.
  3. Describe how the wording of a prompt changes the risk of made-up references, and say what that means for a clinician who uses the same tool at work and at home.
  4. Sort claims about a clinical AI tool into randomized, large observational, and vendor or trade-press evidence, and explain why the three tiers disagree in the same direction.
  5. Explain why FDA clearance is not evidence of clinical performance, using the published analyses of FDA decision summaries.
  6. Tell a wrong reference from a wrong claim from a wrong attribution, and say which of the three a search-first tool actually fixes.
  7. Ask six questions about a tool you have never seen, and say why each one is the question to ask.

Time. About 35 minutes of reading, 30 minutes of listening, 45 minutes of doing. You need one general-purpose chatbot and a PubMed tab. Your institutional ChatGPT Edu or Microsoft Copilot account is enough, and neither asks you for anything at signup. A personal free-tier account works too. Nothing installs.

5.1 Lunch with a vendor

It is noon conference in your second month of intern year. The sandwiches are better than usual, because a company is paying for them. Their representative is showing a slide about an ambient scribe: a phone app that listens to the visit and writes the note. The slide says documentation time fell 40 percent and burnout fell by a third at a large health system. Your program director is nodding. The chief resident leans over and asks what you think, because you are the intern, and interns get asked.

Ambient documentation · outcomes at scale

Give your clinicians their evenings back

−40%documentation time
−33%burnout
96%would recommend
Results from a large integrated health system, 2024–2025
Illustrative slide for this chapter; not a real vendor’s material.

Look at the slide for a moment before reading on. Every number on it is the kind of number you will be shown. Nothing on it tells you where any of them came from.

Earlier that morning, a senior resident handed you a reference for the case presentation you give on Thursday. She had pulled it from ChatGPT. It has a plausible first author, a real journal, a well-formed DOI, and a finding that fits your case nicely. You have not opened it.

These are the same problem at two scales. In both cases someone has handed you a claim that looks right. You are the person whose name will be on whatever that claim supports. The slide and the reference need the same question, which this chapter will ask twice: what would have to be true for this to be trustworthy, and did anyone check?

Clinicians have a word for that habit when it is applied to a paper. They call it critical appraisal. This chapter extends it to two things a paper never used to do: write its own references, and get sold to your department over lunch.

5.2 Why this matters

A physician in a white coat seen from behind, facing four monitors showing an ECG, an echocardiogram, and a video call with two people.
Figure 5.1: Every tool in this chapter arrives through a screen like these, between a patient and a decision. Cardiologist Juan Manuel Romero in a telemedicine consultation, Ciudad Obregón, Mexico. Photo: Intel Free Press, CC BY-SA 2.0, via Wikimedia Commons.

You will meet the reference problem within the year. Someone on your team will hand you a tool that produces citations, or you will open one yourself at 2 a.m. to answer a management question. Whether the reference exists is a separate question from whether it looks right. Checking takes twenty minutes.

Your program will meet the vendor problem, and increasingly so will you. Residents sit on the committees that pilot these tools. The vendor’s slide, the health system’s press release, and the randomized trial will say three different things about the same product. Deciding which to believe is a skill you will still need in fifteen years, when every product named in this chapter is gone.

5.3 How it works

5.3.1 Two skills that get treated as one

Appraising an AI tool is two separate skills. The first is checking what a model just told you: does the resident’s citation exist, and does it say what the model claims? The second is checking what the literature says about a type of tool: does the scribe on the lunch slide actually give you your evenings back?

Most people arrive assuming the first skill is the hard one and the second is someone else’s job. This chapter reverses that. Checking five references takes twenty minutes and a PubMed tab. The answer is clear and often surprising. Weighing a finding that most clinicians declined a tool against a health system’s press release takes judgment.

The two halves connect through the question from the lunch room, asked twice. In the first half, you are the one who checks. In the second, you grade how well other people checked.

5.3.2 Where a citation comes from

A general-purpose language model produces a citation the same way it produces any other text: by writing the most likely next words, one token at a time. A token is a word or piece of a word. The model is doing, at enormous scale, what your phone’s keyboard does when it suggests the next word. A well-formed author list, a real journal name, and a DOI in the right format are what likely next words look like. The model is not looking in a database and failing. It never had one.

xkcd comic. One figure asks how a machine learning system works; the other explains you pour data into a big pile of linear algebra and collect the answers on the other side, and if they are wrong you stir the pile until they look right.
Figure 5.2: What is inside. The mechanism really is closer to this than to a librarian. xkcd 1838, Machine Learning, by Randall Munroe, CC BY-NC 2.5.

A retrieval-grounded tool does this the other way round. “Retrieval-grounded” is the term you will meet in the literature. It means the tool searches a database of real documents first, then writes text about what it found. This chapter also calls it a search-first tool. Several products now work this way, and the retrieval chapter shows how to tell whether a given tool really does. Figure 5.3 shows the two paths.

flowchart LR
  accTitle: Two ways a tool arrives at a citation
  accDescr: Two flowcharts side by side. General-purpose model, your question leads to generate answer text, then generate citation text, then something that looks like a reference. Retrieval-grounded tool, your question leads to search PubMed or an index, then real papers, then generate an answer about those papers, then cites what it found.
  subgraph G["General-purpose model"]
    direction TB
    q1[Your question] --> gen1[Generate answer text]
    gen1 --> cite1[Generate citation text]
    cite1 --> out1[Looks like a reference]
  end
  subgraph R["Retrieval-grounded tool"]
    direction TB
    q2[Your question] --> search[Search PubMed or an index]
    search --> docs[Real papers]
    docs --> gen2[Generate answer about those papers]
    gen2 --> out2[Cites what it found]
  end
Figure 5.3: Two ways a tool arrives at a citation. On the left, the citation is generated as text; whether it exists is a matter of chance. On the right, the citation is retrieved from an index before any text is written.

This is why a search-first tool’s rate of made-up references is different in kind, not just a bit lower. It can still misread a paper it found. It cannot invent one it did not find.

5.3.3 Three things that can be wrong

A citation that exists is not the end of the check. Three things can be wrong, and each can fail on its own:

Table 5.1: Three independent failures. Fixing the first does nothing for the third.
What can be wrong The question Who fixes it
The reference Does this paper exist? A search-first tool, almost completely
The attribution Does this paper say what the tool says it says? A search-first tool, partly; it can still misread what it found
The claim Is the answer clinically right? Nobody but you
xkcd comic in four panels titled Citogenesis: a writer invents a fact and adds it to Wikipedia, a journalist cites Wikipedia, Wikipedia then cites the journalist, and the loop closes with the fact now sourced.
Figure 5.4: The attribution failure, before language models existed. A claim gets a citation, the citation gets cited, and the loop closes without anyone checking the source. Retrieval-grounded tools trained on the web inherit every loop like this. xkcd 978, Citogenesis, by Randall Munroe, CC BY-NC 2.5.

The third row is the one that matters on the wards, and it fails in fluent prose. In one study, 33 physicians across 17 specialties wrote 284 questions and graded the answers. The median score was 5.5 out of 6. But 36 answers scored 1 or 2, meaning mostly or completely wrong, with no change in tone (Goodman et al., 2023). In another, oncologists checked chatbot treatment recommendations for breast, prostate, and lung cancer against the NCCN guidelines, the standard US cancer treatment guidelines. Every output contained at least one option that matched a guideline. About a third also contained at least one that did not. One in eight recommended a treatment that was not in any guideline (Chen et al., 2023). The authors’ phrase for the pattern is worth remembering: the tool was most likely to mix incorrect recommendations among correct ones, an error difficult even for experts to detect.

Both studies also found the answers changed with the wording of the question and with the model version. That is the “these numbers move” point again, this time for answers rather than references.

5.3.4 What the published rates look like

Three studies give the range. Across 20 medical questions put to ChatGPT, 41 of 59 generated references (69%) were fabricated (Gravel et al., 2023). In a psychiatry test, only 2 of 35 citations were real (McGowan et al., 2023). A 2026 comparison of five tools used the 2/1/0 scale you will use below. It found zero fabrications from OpenEvidence, which searches PubMed directly. The general-purpose models ranged from Claude at 78.0% accuracy to Gemini at a 26.2% fabrication rate (McLaughlin et al., 2026).

Table 5.2: Published fabrication rates. Each is one snapshot of one tool in one domain.
Study Setting Result
Gravel et al. (2023) 20 medical questions, ChatGPT 41 of 59 references fabricated (69%)
McGowan et al. (2023) Psychiatry, ChatGPT and Bard 2 of 35 citations real
McLaughlin et al. (2026) Spine surgery, five tools OpenEvidence 0 fabrications; Claude 78.0% accurate; Gemini 26.2% fabricated

The same 2026 study contains the finding this chapter depends on most. The fabrication rate for ChatGPT rose from 2.1% when the question was asked in a clinician’s words to 20.0% when it was asked in a patient’s words (McLaughlin et al., 2026). Same tool, same underlying question, different register. Register is the linguist’s word for the way a particular group of people talks. In this chapter it means one thing: whether the question sounds like a doctor or like a patient.

ImportantThe practical consequence

The version of a tool you trust at work may behave differently when you, or your patient, ask the same question in everyday words. Testing a tool once, with one kind of wording, is not enough.

If English is not your first language, or you write the way you talk rather than the way a journal article is written, your prompts may sound like a patient’s even when you are the clinician. The finding is about the words, not the credentials.

There is an equity finding in this number that the paper does not name. The people least able to check an answer, patients asking in their own words, get the answers most likely to cite something that does not exist.

NoteThese numbers move

Fabrication rates depend on the tool and the month, not on the technology. A second 2026 study, on questions about facial cosmetic surgery, reported the GPT-versus-Gemini ordering reversed: 36% of GPT-4’s citations could not be verified, against 14% for Gemini and 8.8% for DeepSeek (Frutuoso Maia et al., 2026). When your own results below disagree with Table 5.2, that is expected. How the tool works, shown in Figure 5.3, is the durable part. The percentages are not.

5.3.5 Grading the evidence about a tool

The second skill applies the same question to the literature about a type of tool. Ambient documentation, or the “AI scribe”, is the right example. The evidence is large and recent, and it disagrees with itself in a consistent direction.

Claims about scribes come in three tiers:

  • Randomized trials. Clinicians or clinics assigned by chance to the tool or to a comparison.
  • Large observational studies. Real-world data from many clinicians, without randomization. This includes pre/post designs, which measure the same clinicians before and after they start using the tool. It also includes sites that adopt the tool one after another without randomization, and surveys of people who chose to use it.
  • Vendor or trade press. A claim that comes from a company, a press release, or industry news (the “trade press”), however specific the number sounds.

Here is what one study looks like when each of the three tiers reports it to you. All three headlines are about the same paper, the five-system study cited below.

Same paper, three versions: the institution’s press office, a journalist, and a professional society. None of them is wrong. Which one would the vendor at lunch have put on the slide?

The three tiers disagree in the same direction. Pre/post studies of health-system rollouts report large drops in burnout. Randomized trials report smaller effects. They also consistently fail to change the one outcome that justifies the whole product category: after-hours charting. In the largest study to date, across five academic systems and 8,581 clinicians, total EHR time fell 13.4 minutes per day and documentation time 16.0 minutes per day, with 0.49 more visits per week. But after-hours EHR time did not change significantly (Rotenstein et al., 2026). You will work out why in the exercise, before reading the explanation.

This is the third of the course’s three recurring lessons: the published number is the optimistic one. The 0.49 extra visits may also be a small example of the second, efficiency gains get eaten by volume: time saved tends to get filled with more work. The workforce chapter explains how that saved time turns into billing.

5.3.6 “But it’s FDA-cleared, isn’t it fine?”

Clearance answers a different question from the one a clinician is asking. Most cleared AI devices are radiology tools, so if you plan to go into radiology this section is about your daily work. An analysis of 130 FDA decision summaries for AI devices found clinical performance data reported for about half, and none at all for roughly a quarter. Well under a third of those with clinical data gave sex-specific results (Wu et al., 2021). A 2024 follow-up found the gap persists: many authorized tools have no published clinical validation (Chouffani El Fassi et al., 2024). A benchmark score standing in for performance on patients is the first of the course’s three recurring lessons, proxies fail, in regulatory form. The bias and equity chapter explains that lesson in full.

ImportantCleared is not validated

Clearance says the device is similar enough to one already on the market (the FDA’s term is “substantial equivalence”) and that the manufacturer has a quality system. It does not certify that a device improves outcomes, that it was tested on patients like yours, or, often, that clinical performance was measured at all. If you remember one sentence from this chapter, remember this one, and the paper with it.

The device boundary, the Predetermined Change Control Plan, and the clinical decision support criteria belong to the ethics and regulation chapter. The FDA’s list of AI-enabled devices held 1,614 authorizations when it was updated on 4 September 2026, with decisions through June 2026. The FDA says the list is not complete, and one product can appear more than once, so quote the count with its date.

5.3.7 What a trustworthy tool looks like

Everything so far grades evidence that already exists. The tool at lunch is new, and there may be no literature yet. You still need something to ask.

The most useful checklist is FUTURE-AI, a consensus of 117 experts from 50 countries published in the BMJ in 2025. It reduces trustworthy clinical AI to six principles and 30 practices (Lekadir et al., 2025). The six principles, and the question each one becomes when a vendor is standing in front of you, are in Table 5.3.

Table 5.3: FUTURE-AI’s six principles as six questions.
Principle What it means The question for the vendor
Fairness Performance does not depend on who the patient is Was it tested on patients like mine, and did performance differ by sex, race, language, or age?
Universality It works outside the place that built it Has it been validated at a site that is not yours, on a different EHR?
Traceability You can see where it came from and where an output came from What was it trained on, and can I trace this note, or this recommendation, back to its source?
Usability The people who use it were involved, and kept using it Who designed how it fits into the clinic day, and what fraction of clinicians offered it were still using it at six months?
Robustness Its performance drops a little at a time, not all at once What happens with a bad microphone, an interpreter in the room, a rare presentation, an unusual chart?
Explainability It can say why, in a form you could defend If this is wrong in a chart review, what will I be able to show about why I trusted it?

Two things to notice. Traceability asks where things came from, and that is what this chapter’s first half is about. A tool that shows you its sources next to its answer has answered it. A tool that writes a reference list has not. And none of the six asks whether the tool is accurate on a benchmark, a fixed test set used to compare tools. Accuracy is necessary, and it appears under robustness. It is not the checklist, because a tool can be accurate on the population it was built on and fail every other row.

NoteYou have a right to some of these answers

In the United States, the ONC HTI-1 rule requires decision-support tools inside certified EHRs to show the clinician their “source attributes”: what the tool was trained on, how it was validated, and whether fairness was tested. That is the closest thing to a legal right to traceability. It applies to the tool built into your EHR, not to the consumer chatbot on your phone. The rule is in force as of September 2026. A proposed rule, HTI-5, would remove the source-attribute requirement; the ethics and regulation chapter tracks its status.

There is a seventh question the six do not ask. It belongs to the patient in the room rather than the clinician: does the patient know the visit is being recorded, and did they agree? No tier of evidence in Part B answers it, and no vendor slide will. Ask it anyway.

5.3.8 Two findings to keep in your pocket

One is a warning against overconfidence. When 148 final-year students evaluated ChatGPT answers to ten clinical vignettes, five of them deliberately wrong, the median rate of correct evaluation was 56% (Waldock et al., 2025). Students credited their case-based and pathology teaching for the errors they caught. Only 5% had heard the term “clinical prompt engineering”. You are better prepared than that group only if you search rather than judge by appearance.

The other is a warning against dismissing these tools altogether. In a randomized trial of 50 physicians, access to a language model did not significantly improve diagnostic reasoning (adjusted difference 2 points, 95% CI −4 to 8). The model alone scored 16 points higher than physicians using conventional resources (95% CI 2 to 30) (Goh et al., 2024). A tool can be capable and still fail to help the person holding it. That is an appraisal finding about humans, and the centaur-or-cyborg chapter returns to that finding.

5.4 Watch or listen

Podcast. NEJM AI Grand Rounds, “The OpenEvidence Episode: Dr. Travis Zack on the Future of Clinical Evidence” (May 20, 2026; 66 min). https://ai-podcast.nejm.org/e/the-openevidence-episode-dr-travis-zack-on-the-future-of-clinical-evidence/

Zack is OpenEvidence’s chief medical officer, so this is the search-first side of Figure 5.3 described by someone who builds it. Listen for how the tool decides what counts as evidence, and for what he says trust depends on. Then ask the question this chapter asks of everything: what would have to be true for that to hold, and who has checked? Thirty minutes is enough for the exercise; the rest is worth hearing when you have time.

Video alternative. “Travis Zack on OpenEvidence and the Future of Medical AI”, Stanford Department of Medicine, YouTube (July 2026; 44 min). https://www.youtube.com/watch?v=nrfGSE__7po

For the second half. NEJM AI Grand Rounds, “What Values are in AI? A Conversation with Dr. Zak Kohane” (December 17, 2025; 78 min). https://ai-podcast.nejm.org/e/what-values-are-in-ai-a-conversation-with-dr-zak-kohane/ Listen for the remark that ambient documentation spread not because it improves accuracy or throughput but because it improves clinician satisfaction. Compare that with the evidence tiers in Part B.

5.5 Do

WarningNo patient information, ever

Every question in this exercise is a general question about management, not a question about a patient. Consumer tools have no business associate agreement. That is the contract that makes a vendor legally responsible for protecting patient data. Free tiers may also use what you type to train the next model. You will be tempted to test the tool on something from your sub-internship. Do not.

The harder case you will meet in month one: a tool your hospital licenses and runs inside the EHR is a different question. Ask whether a business associate agreement exists, and what your institution’s policy says about documenting that you used it. the ethics and regulation chapter covers that.

What counts as patient information, and why taking out the name is not enough: the patient information page.

TipPubMed is the verification tool

Use PubMed and nothing else. The exercise depends on searching an index that without question contains the real literature. Google Scholar’s loose matching will find you a different paper for a made-up citation, and you will score it as real.

5.5.1 Part A: the fabrication hunt (25 minutes)

Step 1. Pick a question in your specialty. You want a management question with a real literature and real controversy. If you cannot think of one, find one:

("large language model" OR "artificial intelligence") AND <your specialty>
AND (validation OR "external validation")

Filter PubMed to the last three years, read three abstracts, and take the clinical question one of them is about. If you have no specialty yet, use anticoagulant choice in atrial fibrillation with stage 4 chronic kidney disease; it works reliably.

Step 2. Ask twice. Run both prompts below in the same tool, in two separate conversations, with your question substituted in. The first is worded as a clinician would ask it. The second is worded as a patient would. Both ask for a DOI or PMID: the DOI is the publisher’s permanent link for a paper, and the PMID is its PubMed ID number.

Asked as a clinician:

I am a physician managing <your clinical question>. Summarize the current evidence and give me five references from the peer-reviewed literature with authors, journal, year, and DOI or PMID.

Asked as a patient:

I have <the condition, in plain words> and my doctor mentioned <the treatment>. Can you explain what the research says, and give me five references from medical journals with authors, journal, year, and DOI or PMID so I can look them up myself?

If you have a verified OpenEvidence account, run the clinician prompt there as well, on a third row. Verification asks for a clinician credential or, for US medical students, proof of enrollment such as a student ID. Students outside the United States, or who prefer not to upload a document, can skip this row; it is optional for that reason. VERIFY: student verification wording on the OpenEvidence signup form; checked 2026-09-08 against OpenEvidence’s 2024 announcement and two library guides, not the live form

Step 3. Score every reference. Search PubMed by title, then by first author plus year, then by DOI or PMID. Give each reference one score from Table 5.4. The scale is deliberately coarse. Arguing with yourself about borderline cases is part of the exercise.

Table 5.4: The 2/1/0 citation scoring scale of McLaughlin et al. (2026).
Score Meaning How you decide
2, accurate The reference exists as given and supports the claim the model attached to it. Found in PubMed with matching authors, journal, and year, and the abstract is consistent with what the model said it showed.
1, misrepresented The reference exists, but the model’s description of it is wrong. Found in PubMed, but it is a different design, population, or conclusion, or the model attributed a finding it does not contain.
0, fabricated The reference does not exist as given. Not found by title, author plus year, or DOI/PMID after a genuine search. A real DOI pointing at a different paper than the one named scores 0.

The clinician prompt was run on anticoagulation in atrial fibrillation with stage 4 chronic kidney disease, in a general-purpose model without search (Claude, 2026-09-07), and every reference was checked in PubMed. All five were real. That is not the result you should expect. It is the “these numbers move” callout happening in front of you. The three examples below show what each score looks like when it does happen. The first is from that run. The second and third are made up to show the pattern, and are labeled as such.

Score 2, accurate. Stanifer JW et al. Apixaban versus warfarin in patients with atrial fibrillation and advanced chronic kidney disease. Circulation 2020;141:1384–1392. The tool said it showed less bleeding with apixaban in creatinine clearance 25 to 30. PubMed: exists, authors and journal match, and the abstract reports major bleeding hazard ratio 0.34 in exactly that subgroup (Stanifer et al., 2020). Score 2.

Score 1, misrepresented (made up). Pokorney SD et al. Apixaban for patients with atrial fibrillation on hemodialysis. Circulation 2022;146:1735–1745, described by the tool as “a randomized trial showing apixaban is safer than warfarin in dialysis patients.” PubMed: the paper exists (Pokorney et al., 2022). Its own conclusion is that it stopped early and had inadequate power to draw any conclusion about bleeding. The reference is real. The sentence attached to it is the opposite of what the authors wrote. Score 1. This is the failure a search-first tool can still make, and the one only reading the abstract catches.

Score 0, fabricated (made up). Hernandez AV et al. Direct oral anticoagulants in stage 4–5 CKD: a randomized comparison. J Am Coll Cardiol 2021;77(9):1123–1131. doi:10.1016/j.jacc.2020.12.041. Every feature is plausible: a real journal, a common surname, a year in range, a DOI with the journal’s correct prefix. Search PubMed by title: nothing. By author and year: the real Hernandez AV publishes on other topics. By DOI: it resolves to a different paper or to nothing. Score 0. What reveals it is not any single feature. It is that no feature can be found.

Step 4. Record it. Copy this into a document. It is what you will grade yourself on in the Check block.

Tool: ____________________   Model version if shown: ____________
Clinical question: ______________________________________________

Asked as    Ref  Title as given by the model           In PubMed?  Score
Clinician    1   ____________________________________   Y / N      ___
Clinician    2   ____________________________________   Y / N      ___
Clinician    3   ____________________________________   Y / N      ___
Clinician    4   ____________________________________   Y / N      ___
Clinician    5   ____________________________________   Y / N      ___
Patient      1   ____________________________________   Y / N      ___
Patient      2   ____________________________________   Y / N      ___
Patient      3   ____________________________________   Y / N      ___
Patient      4   ____________________________________   Y / N      ___
Patient      5   ____________________________________   Y / N      ___
OpenEvidence (optional):  ___ ___ ___ ___ ___

Asked as a clinician: 0s ___  1s ___  2s ___
Asked as a patient:   0s ___  1s ___  2s ___

For any reference scored 1: in one sentence, what did the model claim the
paper showed, and what does it actually show?

For any reference scored 0: what about it looked credible? Name the specific
feature: plausible author, real journal, plausible year, well-formed DOI.

The two free-text answers matter more than the counts. Being able to say why a made-up citation looked credible is a skill you can use on the next tool. “Three of five were fake” is a number that will be out of date next year.

NoteTen references is not a study

You have ten observations. That is nowhere near enough to detect a 2.1% versus 20.0% difference. You are reproducing the design of McLaughlin et al., not the result. If your split goes the same direction, it is consistent and underpowered. If it goes the other way, that is the better lesson. One small replication does not overturn a published finding, and one published finding does not settle a question whose answer keeps changing.

5.5.2 Part B: grading the evidence (20 minutes)

Below are ten claims about ambient AI scribes, with their sources removed. Sort each into a tier (randomized trial, large observational, or vendor and trade press), rate your confidence from 1 to 5, and give a one-line reason. Then answer the two closing questions before you open the solution.

  1. Across 3,442 physicians and 303,266 encounters at one integrated health system, physicians reported improved experience with an ambient documentation tool.
  2. Burnout among tool users fell from 51.9% to 38.8%, odds ratio 0.26.
  3. Burnout dropped 21.2 points, and 60% of participants said the tool made them more likely to extend their clinical career.
  4. Documentation time fell by 6.89 minutes per day and after-hours time by 5.17 minutes per day.
  5. Burnout and task load improved, but clinicians rated both tools as producing occasional clinically significant inaccuracies.
  6. Work exhaustion fell by 0.44, but the reduction in workweek hours lost statistical significance after outlier removal.
  7. There was no meaningful difference in after-hours “pajama time” between the tool and usual documentation.
  8. Burnout fell from 89% to 56%, with no significant change in pajama time or charting time.
  9. Across five academic systems and 8,581 clinicians, total EHR time fell 13.4 minutes per day and documentation time 16.0 minutes per day, with 0.49 more visits per week, but after-hours EHR time did not change significantly.
  10. Seventy-nine percent of clinicians offered the tool declined to use it.
Claim   Tier (RCT / Large obs / Vendor-trade)   Confidence (1–5)   One-line reason
  1     ________________________________        ___                ________________
  2     ________________________________        ___                ________________
  3     ________________________________        ___                ________________
  4     ________________________________        ___                ________________
  5     ________________________________        ___                ________________
  6     ________________________________        ___                ________________
  7     ________________________________        ___                ________________
  8     ________________________________        ___                ________________
  9     ________________________________        ___                ________________
 10     ________________________________        ___                ________________

Closing question 1. The pre/post health-system studies look substantially
better than the randomized trials. Name two mechanisms that would produce
that pattern even if the tool worked exactly as well as the RCTs say.

  (1) ______________________________________________________________
  (2) ______________________________________________________________

Closing question 2. Which outcome do the pre/post studies and the RCTs
disagree about most, and why is that the outcome you personally care about?

Claim 10 is designed not to fit the scheme. Whichever tier you put it in, be able to defend the choice.

Two mechanisms: selection and Hawthorne. The pre/post studies measure clinicians who chose to adopt an optional tool: people already inclined to like it, whose way of working it fits. The randomized trials measure clinicians who were assigned it. The five-system study shows the gap in numbers. Of 8,581 clinicians offered the tool, 1,809 adopted it, and those adopters are the population every earlier pre/post study sampled (Rotenstein et al., 2026). Add that people who know they are in a documentation study document differently, the Hawthorne effect, and the direction of the gap is explained without anyone having lied.

The outcome they disagree about is after-hours charting. It justifies the entire product category, and it is the outcome the randomized evidence most consistently fails to move. One plausible reading, offered as a hypothesis and not a finding, is that time saved goes into more visits (the 0.49 additional visits per week) rather than coming back as personal time. A similar pattern: AI-drafted replies to patient inbox messages produced no significant change in reply time but significant reductions in task load and work exhaustion (Garcia et al., 2024). Less burden and more free hours are different outcomes. Treating them as the same is how a real benefit gets oversold into a false one.

Claim 10 is a number from journalism about a peer-reviewed study, not from the paper’s own results. You will meet this category constantly. It is neither peer-reviewed evidence nor vendor marketing, and it needs its own habit: find the paper, and see whether the number is in it.

Two independent checks on the vendor story. The Peterson Health Technology Institute, which takes no vendor funding, concluded in 2025 that scribes likely reduce burnout but that their financial impact is unclear. In 2026 it concluded that administrative AI reduces burden within organizations without lowering system-wide costs (2025 report, 2026 report). A UCSF analysis found scribe adopters billed about 1.8 more relative value units a week, the billing measure of clinical work, with no rise in claim denials (Holmgren et al., 2026). The university’s press release called that a 5.8% increase. The study was one health system and adopters chose to adopt, so the authors call for a randomized test.

The tiers, so you can check your answers: claims 1, 2, 3, 4, and 9 are large observational; 5, 6, 7, and 8 are randomized trials; 10 is trade press about a peer-reviewed paper. The sources are in the answer key held by the course tutor, which will also grade your reasons.

5.6 Check

Score yourself against this before moving on.

You have Meets Falls short
Ten scored references Every score is backed by a PubMed search, not a judgment of plausibility “Looked real” or “looked fake” without a search
A reason a fabricated reference looked credible Names a specific feature “It just seemed legit”
A reason a misrepresented reference was wrong States what the model claimed and what the paper shows Restates the model’s summary
Ten tiered claims Each tier has a one-line reason; claim 10 is defended either way Tiers assigned without reasons
Two mechanisms Selection and observation effects, or equivalents, in your own words “The pre/post studies were biased”
The one sentence You can say “cleared is not validated” and name the paper Only the slogan
The mechanism You can explain why a generated reference may not exist without using the word “hallucination” “It hallucinates”
The vendor question You can name the FUTURE-AI principle your question tests, and why that one first A question with no principle behind it

5.7 What this changes for you

What does this change about how you’ll practice? Write it down before you close this chapter. Two specifics. First, what will you now check before you cite a reference an AI tool gave you in a note, a presentation, or an application, and which of the three failures in Table 5.1 does that check catch? Second, go back to the lunch room: the chief resident is still waiting. What do you say about the 40 percent, and which one of the questions in Table 5.3, or the seventh about the patient, do you ask the vendor first? One word is not an answer to either.

5.8 Summary

  • A general-purpose model writes citations as text; a retrieval-grounded tool searches first. The difference is in kind, not in degree.
  • Published fabrication rates run from a quarter to two-thirds for general models and near zero for search-first tools, and they move with the tool and the month.
  • The same question in a patient’s words can produce a tenfold higher fabrication rate than in a clinician’s.
  • Evidence about a type of tool comes in tiers that disagree in a consistent direction; selection and observation effects explain most of the gap.
  • The reference, the attribution, and the claim fail independently. Search fixes the first, helps the second, and does nothing for the third.
  • Less burden and more free hours are different outcomes.
  • Cleared is not validated.
  • A tool with no literature yet still has six questions to answer, and a seventh that belongs to the patient. Ask them at the next vendor lunch.

5.9 Go deeper

Papers

Talks

Adams, L., Fontaine, E., Lin, S., Crowell, T., Chung, V. C. H., & Gonzalez, A. A. (2024). Artificial Intelligence in Health, Health Care, and Biomedical Science: An AI Code of Conduct Principles and Commitments Discussion Draft. NAM Perspectives, 2024. https://doi.org/10.31478/202403a
Badal, K., Lee, C. M., & Esserman, L. J. (2023). Guiding principles for the responsible development of artificial intelligence tools for healthcare. Communications Medicine, 3(1). https://doi.org/10.1038/s43856-023-00279-9
Chen, S., Kann, B. H., Foote, M. B., Aerts, H. J. W. L., Savova, G. K., Mak, R. H., & Bitterman, D. S. (2023). Use of Artificial Intelligence Chatbots for Cancer Treatment Information. JAMA Oncology, 9(10), 1459–1462. https://doi.org/10.1001/jamaoncol.2023.2954
Chouffani El Fassi, S., Abdullah, A., Fang, Y., Natarajan, S., Masroor, A. B., Kayali, N., Prakash, S., & Henderson, G. E. (2024). Not all AI health tools with regulatory authorization are clinically validated. Nature Medicine, 30(10), 2718–2720. https://doi.org/10.1038/s41591-024-03203-3
Frutuoso Maia, B. K., Morais, E. F. de, Santana Santos, T. de, & Charles Pagotto, L. E. (2026). Performance of large language models in preoperative and postoperative counselling for aesthetic facial procedures. The British Journal of Oral & Maxillofacial Surgery, 64(3), 216–222. https://doi.org/10.1016/j.bjoms.2026.01.002
Gallifant, J., Afshar, M., Ameen, S., Aphinyanaphongs, Y., Chen, S., Cacciamani, G., Demner-Fushman, D., Dligach, D., Daneshjou, R., Fernandes, C., Hansen, L. H., Landman, A., Lehmann, L., McCoy, L. G., Miller, T., Moreno, A., Munch, N., Restrepo, D., Savova, G., … Bitterman, D. S. (2025). The TRIPOD-LLM reporting guideline for studies using large language models. Nature Medicine, 31(1), 60–69. https://doi.org/10.1038/s41591-024-03425-5
Garcia, P., Ma, S. P., Shah, S., Smith, M., Jeong, Y., Devon-Sand, A., Tai-Seale, M., Takazawa, K., Clutter, D., Vogt, K., Lugtu, C., Rojo, M., Lin, S., Shanafelt, T., Pfeffer, M. A., & Sharp, C. (2024). Artificial Intelligence-Generated Draft Replies to Patient Inbox Messages. JAMA Network Open, 7(3), e243201. https://doi.org/10.1001/jamanetworkopen.2024.3201
Goh, E., Gallo, R., Hom, J., Strong, E., Weng, Y., Kerman, H., Cool, J. A., Kanjee, Z., Parsons, A. S., Ahuja, N., Horvitz, E., Yang, D., Milstein, A., Olson, A. P. J., Rodman, A., & Chen, J. H. (2024). Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Network Open, 7(10), e2440969. https://doi.org/10.1001/jamanetworkopen.2024.40969
Goodman, R. S., Patrinely, J. R., Stone, C. A., Zimmerman, E., Donald, R. R., Chang, S. S., Berkowitz, S. T., Finn, A. P., Jahangir, E., Scoville, E. A., Reese, T. S., Friedman, D. L., Bastarache, J. A., Heijden, Y. F. van der, Wright, J. J., Ye, F., Carter, N., Alexander, M. R., Choe, J. H., … Johnson, D. B. (2023). Accuracy and Reliability of Chatbot Responses to Physician Questions. JAMA Network Open, 6(10), e2336483. https://doi.org/10.1001/jamanetworkopen.2023.36483
Gravel, J., D’Amours-Gravel, M., & Osmanlliu, E. (2023). Learning to Fake It: Limited Responses and Fabricated References Provided by ChatGPT for Medical Questions. Mayo Clinic Proceedings. Digital Health, 1(3), 226–234. https://doi.org/10.1016/j.mcpdig.2023.05.004
Holmgren, A. J., Fenton, C. L., Thombley, R., Soleimani, H., Croci, R., DeMasi, O., Byron, M. E., Murray, S. G., Adler-Milstein, J. R., & Yazdany, J. (2026). Ambient Artificial Intelligence Scribes and Physician Financial Productivity. JAMA Network Open, 9(1), e2553233. https://doi.org/10.1001/jamanetworkopen.2025.53233
Lekadir, K., Frangi, A. F., Porras, A. R., Glocker, B., Cintas, C., Langlotz, C. P., Weicken, E., Asselbergs, F. W., Prior, F., Collins, G. S., Kaissis, G., Tsakou, G., Buvat, I., Kalpathy-Cramer, J., Mongan, J., Schnabel, J. A., Kushibar, K., Riklund, K., Marias, K., … Starmans, M. P. A. and. (2025). FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ (Clinical Research Ed.), 388, e081554. https://doi.org/10.1136/bmj-2024-081554
McGowan, A., Gui, Y., Dobbs, M., Shuster, S., Cotter, M., Selloni, A., Goodman, M., Srivastava, A., Cecchi, G. A., & Corcoran, C. M. (2023). ChatGPT and Bard exhibit spontaneous citation fabrication during psychiatry literature search. Psychiatry Research, 326, 115334. https://doi.org/10.1016/j.psychres.2023.115334
McLaughlin, N. D., Srinivas, A. N., Lowe, Z. F., Botterbush, K. S., Patel, M. S., & Avila, M. J. (2026). Large Language Model Hallucinations in Spine Surgery: A Comparative Analysis of Clinician vs Patient-Level Prompts. Neurosurgery Practice, 7(3), e000244. https://doi.org/10.1227/neuprac.0000000000000244
Pokorney, S. D., Chertow, G. M., Al-Khalidi, H. R., Gallup, D., Dignacco, P., Mussina, K., Bansal, N., Gadegbeku, C. A., Garcia, D. A., Garonzik, S., Lopes, R. D., Mahaffey, K. W., Matsuda, K., Middleton, J. P., Rymer, J. A., Sands, G. H., Thadhani, R., Thomas, K. L., Washam, J. B., … Granger, C. B. and. (2022). Apixaban for Patients With Atrial Fibrillation on Hemodialysis: A Multicenter Randomized Controlled Trial. Circulation, 146(23), 1735–1745. https://doi.org/10.1161/circulationaha.121.054990
Rotenstein, L. S., Holmgren, A. J., Thombley, R., Sriram, A., Dbouk, R. H., Jost, M., Aizenberg, D., MacDonald, S., Kanaparthy, N., Williams, B., Hsiao, A., Schwamm, L., Murray, S., Byron, M., You, J. G., Centi, A. J., Iannaccone, C., Frits, M., Landman, A. B., … Mishuris, R. G. (2026). Changes in Clinician Time Expenditure and Visit Quantity With Adoption of Artificial Intelligence-Powered Scribes: A Multisite Study. JAMA, 335(16), 1408–1417. https://doi.org/10.1001/jama.2026.2253
Shah, R., & Miranda, J. L. (2026). Leveraging generative artificial intelligence errors to teach appropriate citation usage. Journal of Microbiology & Biology Education, 27(2), e0019525. https://doi.org/10.1128/jmbe.00195-25
Stanifer, J. W., Pokorney, S. D., Chertow, G. M., Hohnloser, S. H., Wojdyla, D. M., Garonzik, S., Byon, W., Hijazi, Z., Lopes, R. D., Alexander, J. H., Wallentin, L., & Granger, C. B. (2020). Apixaban Versus Warfarin in Patients With Atrial Fibrillation and Advanced Chronic Kidney Disease. Circulation, 141(17), 1384–1392. https://doi.org/10.1161/circulationaha.119.044059
Tierney, A. A., Gayre, G., Hoberman, B., Mattern, B., Ballesca, M., Kipnis, P., Liu, V., & Lee, K. (2024). Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst, 5(3). https://doi.org/10.1056/cat.23.0404
Waldock, W. J., Lam, G., Baptista, A., Walls, R., & Sam, A. H. (2025). Which curriculum components do medical students find most helpful for evaluating AI outputs? BMC Medical Education, 25(1), 195. https://doi.org/10.1186/s12909-025-06735-5
Wu, E., Wu, K., Daneshjou, R., Ouyang, D., Ho, D. E., & Zou, J. (2021). How medical AI devices are evaluated: Limitations and recommendations from an analysis of FDA approvals. Nature Medicine, 27(4), 582–584. https://doi.org/10.1038/s41591-021-01312-x