11  Ethics, liability, and the device line: who answers for the note

AIM-4 · AIM-7 · AIM-X · about 75 minutes

NoteWhat you’ll learn

By the end of this chapter you will be able to:

  1. State that no US malpractice claim involving a clinical AI tool has reached a verdict, and describe what the person suing would have to prove anyway.
  2. Place a clinical scenario in the Price, Gerke, and Cohen 2×2 (a four-box grid) and say which box carries the most legal risk and why.
  3. Describe the Jabbour finding, including the part where explanations did not help, and say what it means for “explainable AI” as a safety measure.
  4. State FDA’s four criteria for non-device clinical decision support and explain why the fourth decides most real cases.
  5. Explain what a Predetermined Change Control Plan is and why it is what makes a model that keeps changing possible to regulate at all.
  6. Say what to document, whom to tell, and what contract to ask about when an AI tool touches a note. Find the current primary-source answer to a regulatory question that has moved since this chapter was written.
  7. Write down one part of clinical judgment you will not hand over to an AI, and what would have to be true for you to change your mind.

Time. About 25 minutes of reading, 20 minutes of listening, 25 minutes of doing, and five minutes of writing at the end. That writing is the point of the whole course. You need a browser and nothing else. No chatbot is required. Where one helps, your institutional ChatGPT Edu or Copilot account is enough. Nothing installs, and nothing in this chapter touches a real patient’s data.

11.1 The sentence you did not write

It is the third month of intern year. Your clinic uses an ambient scribe: a phone app that listens to the visit and drafts the note. Mr. Alvarez, 68, has atrial fibrillation and takes apixaban. Halfway through the visit he mentions, almost in passing, that he stopped it two weeks ago because of nosebleeds. You talk about that for a minute, then move on to his knee. The scribe’s draft is waiting when you get back to the workroom. It is 6:40 in the evening, you have fourteen notes to sign, and this one reads well.

Progress note · Internal Medicine ClinicSigned 18:47 · Intern (PGY-1)
Subjective

68M with paroxysmal atrial fibrillation (CHA2DS2-VASc 4), HTN, and right knee OA presents for follow-up. Continues apixaban 5 mg BID without bleeding complications. Reports 3 weeks of right knee pain, worse with stairs, no trauma. Denies chest pain, dyspnea, palpitations.

Assessment and plan

AF, rate controlled: continue apixaban and metoprolol.
Right knee OA: trial of acetaminophen; PT referral; RTC 8 weeks.

Draft generated by ambient documentation. Reviewed and signed by the clinician. Illustrative note for this chapter; synthetic patient; not any vendor’s or institution’s material.

Read the highlighted sentence. This is the kind of error the published studies of these tools describe. It is not a slip in the transcript. It is a fluent, believable sentence that says the opposite of what the patient said. In a pilot at one academic health system, physicians reviewed 356 scribe-drafted notes. 18% contained an omission, 11.5% contained a made-up statement, and 5.3% contained an error rated as a serious risk if not corrected. 14.9% of all notes were signed with no edits at all (Taylor et al., 2026).

Three weeks later Mr. Alvarez is admitted with a stroke. The admitting team reads your note. Then the hospital’s risk manager reads it, and asks you a question with one answer: who wrote this sentence? The note has your name on it. The waiting room has a sign about “AI-assisted documentation”, and the medical assistant mentions it at check-in. Whether that counts as consent is, as you will see, a question courts are being asked right now.

Nobody in this scenario has done anything unusual. That is the problem this chapter is about. A tool that is allowed to be in the room is wrong, and your name is on the result. What do the law, the regulator, and your own judgment have to say about that?

11.2 Why this matters

You will sign a note drafted by software within a year of starting residency, possibly within a week. You will silence a prediction alert. You will follow a recommendation you cannot fully explain, and override one you can. Each of those is a decision that a lawyer, a regulator, and a patient could later ask you to account for. The rules for doing so are partly written, partly moving, and partly not written at all.

The honest version of the law is short. No US malpractice claim involving a clinical AI tool has reached a verdict as of the date on this chapter. Students who arrive expecting a rule to memorize leave with something that lasts longer: a way to reason about legal risk, one regulatory line you can actually learn, one clear experimental finding about what a bad model does to the clinicians who trust it, and a written position of your own. That last item is the reason this chapter closes the course.

11.3 How it works

11.3.1 How a scribe turns a conversation into a note

An ambient scribe does four things in order. Errors can enter at the second and third steps, and they look different.

1. RecordA phone or room microphone records the visit. This is the step that needs the patient’s consent.
2. TranscribeSpeech recognition turns the audio into text and tries to tell the speakers apart.Errors here: a word misheard, a “no” dropped, a daughter’s words put in the patient’s mouth.
3. SummarizeA language model reads the transcript and writes a draft note, section by section.Errors here: a sentence that was never said, a detail mentioned in passing left out.
4. Review and signYou edit the draft. Your signature makes every sentence yours.
Illustrative diagram for this chapter, not any vendor’s design.

The difference matters when you read the draft. A transcription error often looks wrong: an odd word, a drug name that does not fit. A summarizing error usually looks right. The model writes the kind of sentence that usually appears in a note like this one. Mr. Alvarez’s draft is that kind. The transcript could have been perfect, and the summary could still say what is usual for a patient with atrial fibrillation: that he continues apixaban. So read the draft against what you remember the patient said, not against what sounds like a normal note.

You will likely meet one of these tools on your first day of residency. In February 2026, UCHealth announced it was rolling out Abridge, an ambient scribe, across its system, after a nine-month pilot with 250 clinicians (press release). The scribe open lab lets you watch each step make its own kind of error, on a made-up visit.

11.3.2 No court has answered yet, and what a plaintiff must prove anyway

Malpractice is a type of negligence. To win a negligence case, the person suing (the plaintiff) must prove four things, which lawyers call the four elements. First, that you owed the patient a duty. You did; he was your patient. Second, that you fell below the standard of care, meaning what a reasonably competent physician in your position would have done. Third, that this failure caused the harm. Fourth, that the harm was real. The tool is not a defendant in that list. It is evidence. Under current law, courts measure the physician against the standard of care. They do not measure the AI against anything (Mello & Guha, 2024; Price et al., 2019).

Two things follow. First, “the AI was wrong” is not, by itself, a legal argument. The question a court will ask is whether you should have known better than to rely on the output, given what you knew at the time (Mello & Guha, 2024). Second, suing the vendor instead is harder than you might expect. Courts have usually treated clinical software as a service rather than a product. And under the learned intermediary doctrine, a manufacturer’s duty to warn runs to the physician, who stands between the product and the patient. That rule places the treating clinician’s judgment between the manufacturer and the patient. Price, Gerke, and Cohen note that this may shift as tools become too opaque for a physician to check. But that is the law today (Price, Gerke, and Cohen, 2024).

So Mr. Alvarez’s note would not be argued in court as “the scribe made an error.” It would be argued as “the physician signed a note saying the patient was taking a medication he had told her he had stopped.” Is that below the standard of care at 6:40 p.m. with fourteen notes to sign? That is exactly the question no court has yet answered.

NoteYour insurer, briefly

Malpractice insurers have not written AI out of their policies. The chief operating officer of The Doctors Company told Medical Economics in 2025 that the carrier has no AI exclusion, and that it would defend a physician and pay a covered claim when AI played a role (interview). The carrier reprinted the interview in its own newsletter. Across the wider market, insurers decide case by case, and they advise physicians to tell their insurer before adopting a clinical AI tool VERIFY: quote against the Medical Economics interview; the carrier’s site links it but does not repeat the wording; checked 2026-09-08. For a resident, the coverage that matters first is your institution’s. Someone in the GME office can tell you what it says. Ask in orientation, not after.

11.3.3 Four boxes

Price, Gerke, and Cohen published the framework that everyone since has used (Price et al., 2019). It asks two questions. Did the AI’s recommendation match the standard of care? Did the physician follow it? Two yes-or-no questions give four boxes. In this chapter, “exposure” means how much legal risk the physician carries.

Table 11.1: The Price, Gerke, and Cohen 2×2. Rows are what the AI said; columns are what the physician did.
Physician followed the AI Physician overrode the AI
AI matched the standard of care A. Good outcome likely. If harm occurs, exposure is low. One open question: can following a correct recommendation you could not explain still create exposure? C. High exposure. The physician left the standard of care and disagreed with a tool that got it right.
AI departed from the standard of care B. High exposure. The physician did the nonstandard thing because the machine said so. D. The physician was right. What risk remains lives in the documentation, in hospital policy, and in the fact that the tool stayed in use.

The uncomfortable box is B. Under current law, the physician who follows the AI into nonstandard care carries the most risk. That discourages using AI in exactly the cases where it would most change what you do (Price et al., 2019). The other high-exposure box is C, and it raises a different worry. Once a tool is in use, does a clinician have to justify disagreeing with it?

Two experiments have asked potential jurors. In a nationally representative sample of 2,000 US adults, physicians who followed an AI recommendation for standard care were judged less liable than those who rejected it. Rejecting a nonstandard recommendation gave no similar protection (Tobia et al., 2020). A 2026 replication studied US adults, German adults, and German physicians. Those are the people who would sit as lay jurors in one system and as court-appointed experts in the other. It found the same pattern. Accepting standard-care AI advice was rated more reasonable than rejecting it. Accepting or rejecting nonstandard advice was judged about the same (Tacconelli et al., 2026). Applied to Table 11.1, the practical lesson is this. The tool is safest to follow when it agrees with what you would have done anyway. When it disagrees, you carry the risk yourself, whichever way you decide.

A related question comes up in every class: can you be sued for not using a tool that was available? Today, no. The standard of care is what competent physicians do now, and most do not use these tools. Price, Gerke, and Cohen expect that to change. Once using AI is routine, not using it could itself fall below the standard of care (Price et al., 2019).

Where does Mr. Alvarez’s note sit? Nowhere on the grid, it seems at first. The scribe made no recommendation. But the note is a recommendation once signed. “Continue apixaban” was the plan. It was written on a premise the patient had contradicted, and you followed it. Box B, reached indirectly.

11.3.4 Explanations did not help

The usual answer from institutions is “we will require explainable AI.” That answer has been tested once, in a randomized study, and in that test it did not work. (The bias and equity chapter uses the same study for what a clinician can catch in practice; here it is the legal reading.)

Jabbour and colleagues gave 457 hospitalists, nurse practitioners, and physician assistants written cases of acute respiratory failure. They asked them to judge the likelihood of pneumonia, heart failure, and COPD. Baseline accuracy was 73.0%. A standard AI model raised accuracy by 2.9 points. When the model also showed its reasoning as an image-based explanation, the gain was 4.4 points. A systematically biased model lowered accuracy by 11.3 points. Adding explanations to the biased model recovered 2.3 points. The confidence interval for that difference crossed zero, so the true recovery may be nothing (Jabbour et al., 2023).

Bar chart of change in clinician diagnostic accuracy in percentage points from a 73.0% baseline. Standard AI model without explanation plus 2.9; with explanation plus 4.4; biased model without explanation minus 11.3; biased model with explanation minus 9.1. Whiskers show 95% confidence intervals; the two biased-model intervals overlap.
Figure 11.1: The Jabbour finding, redrawn from the paper’s own numbers. Explanations helped a little when the model was right and did not significantly help when the model was wrong. Data from Jabbour et al., JAMA 2023;330:2275–2284; figure drawn for this chapter.

Compare that with Table 11.1. Imagine a defense that says “the tool showed its reasoning, so the physician should have caught it.” That defense must answer evidence that showing the reasoning did not help the clinicians who saw it. A plaintiff in box B has the same paper: the physician was given a biased model, and the safeguard everyone relies on was shown not to work. Automation bias is the habit of trusting a machine’s suggestion more than it deserves. It is not fixed by adding a paragraph of reasoning. It is a property of tired humans under time pressure, which is what you will be.

11.3.5 Where a tool becomes a medical device

Software that gives clinicians recommendations is either a medical device, regulated by FDA, or it is not. The law draws the line with four tests, which it calls criteria (section 520(o)(1)(E) of the Food, Drug, and Cosmetic Act). FDA’s guidance on how it reads them is the one regulatory document worth memorizing. Software counts as non-device clinical decision support (non-device CDS), and so avoids FDA review, only if it meets all four:

  1. It is not intended to acquire, process, or analyze a medical image, a signal from a lab test (an in vitro diagnostic), or a signal from a monitor.
  2. It is intended to display, analyze, or print medical information.
  3. It is intended to support or provide recommendations to a health care professional about prevention, diagnosis, or treatment.
  4. It is intended to let that professional independently review the basis for the recommendation, so that they do not rely mainly on it to make the decision for a patient.

Criterion 4 decides most real cases. A risk score with no visible reasoning fails it. The same score may pass if it shows its inputs, its logic, and the validation behind it, in plain language. In the current guidance, FDA says outright that software meant for a time-critical decision does not meet criterion 4. The clinician will not have time to review the basis. FDA names automation bias as the reason (FDA, Clinical Decision Support Software, guidance issued January 29, 2026).

flowchart TD
  accTitle: The four criteria as a path
  accDescr: Flowchart of the four criteria as yes-or-no questions. Analyzes an image or signal, yes leads to device. Displays or analyzes medical information, no leads to device. Supports or provides recommendations to a clinician, no leads to device. Clinician can independently review the basis with time to do so, no leads to device, yes leads to non-device CDS with no FDA review.
  q1{"Does it analyze an image,<br/>an IVD signal, or a signal<br/>from a monitor?"} -- yes --> dev[Device: FDA oversight]
  q1 -- no --> q2{"Does it display or analyze<br/>medical information?"}
  q2 -- no --> dev
  q2 -- yes --> q3{"Does it support or provide<br/>recommendations to a clinician?"}
  q3 -- "no (directs the decision,<br/>or is patient-facing)" --> dev
  q3 -- yes --> q4{"Can the clinician independently<br/>review the basis, with time to do so?"}
  q4 -- "no (black box,<br/>or time-critical)" --> dev
  q4 -- yes --> nd[Non-device CDS:<br/>no FDA review]
Figure 11.2: The four criteria as a path. Any ‘no’ makes the software a device. Criterion 4 is where most real tools are decided. It is a question about what the clinician can see, not about how accurate the model is.

A sepsis alert that fires from vital-sign trends and tells the nurse to page you is a device. It fails criterion 1 (it processes monitor signals) and probably criterion 4 (it is time-critical). Any tool that reads an image is a device on criterion 1 alone, which is why most of the AI devices on FDA’s list are radiology tools. A tool that matches a patient’s diagnosis to a guideline and shows you the guideline is non-device CDS. A large language model that answers a clinical question in fluent prose, with no visible source, is a harder case than either. FDA has not yet said which it is. In August 2026 the agency published a discussion paper on generative-AI-enabled devices. It names hallucination, meaning made-up content stated as fact, as the property that sets these tools apart, and asks for comment by 19 October 2026. But a discussion paper is not guidance (FDA, Considerations for the Regulation of Generative AI-Enabled Medical Devices, August 18, 2026). Your scribe, notably, is probably not a device under any of this, as FDA reads the criteria today. It makes no recommendation. It writes what it heard, or what it thinks it heard, and the regulator that reviews it is you.

ImportantThis line moved twice while the course was being written

The four criteria come from a 2016 statute. FDA’s guidance on them was issued in September 2022, replaced on January 6, 2026, and replaced again on January 29, 2026. The January 29 version is the one linked above. The January revisions were presented as a loosening of the rules. The main real change is that FDA will now accept a tool that gives a single recommendation, where only one is clinically appropriate, instead of insisting on a list. The four criteria themselves did not change. Before you rely on anything in this section, open the FDA page and check the date at the top. That check is the skill. The content is the example.

11.3.6 A model that changes

FDA clearance is permission to sell a specific version of a specific algorithm. Most modern models are updated all the time. So what, exactly, is cleared?

FDA’s answer, made final in December 2024, is the Predetermined Change Control Plan, or PCCP. In the original submission, the manufacturer describes the changes it expects to make, the steps it will follow to make and test them, and what effect it expects them to have. Updates that fall inside the plan do not trigger a new review. This is what lets a learning model release improvements without going back to FDA each time. It is the single most important regulatory idea for a physician who will use these tools (FDA, PCCP guidance, final, December 4, 2024). It also raises the question to ask any vendor: if the model on my unit today is not the model that was cleared, what did clearance tell me?

The critical appraisal chapter makes the related point in full, and the short version belongs here too. Clearance is a regulatory status, not proof of validation: of 130 FDA decision summaries for AI devices, about half reported any clinical performance data (Wu et al., 2021), and a 2024 follow-up found the gap persists (Chouffani El Fassi et al., 2024). “Cleared” tells you the manufacturer convinced FDA that the device is about as safe and effective as something already on the market. It does not tell you it was tested on patients like yours.

One small irony, if you want it. FDA’s own internal document-review model, Elsa, was released to agency staff in 2025 (FDA press release, December 2025). In July 2025, CNN reported that FDA reviewers had seen Elsa cite studies that did not exist. The Commissioner said he had not heard those specific concerns, and an HHS spokesperson said that large language models “inherently hallucinate” (CNN, 23 July 2025) VERIFY: quotes as reported by Engadget and NOTUS; the CNN page could not be opened by tool on 2026-09-08. The agency writing the rules is working through the same problems as everyone else.

11.3.7 The note, the patient, and the contract

This is the section the critical appraisal chapter promised you. Three things follow when an AI tool touches a note, and they are separate.

Documentation. You sign it, you answer for it. That is not new law. It is the same rule that applied to a human scribe or a resident’s draft. What is new is the kind of error: a fluent sentence that was never said. Two habits will protect you in a chart review, and both take little time at 6:40 p.m. First, read the plan against the conversation, not against the draft’s own subjective section. The draft agrees with itself even when it is wrong. If you find the error after you have signed, add an addendum. Do not quietly edit the signed note. Second, where your institution’s policy calls for it, record that the note was drafted with ambient documentation and reviewed by you. Many systems add that line automatically. If yours does, know what it says. It is an attestation, a statement in your name that something is true. The Sharp HealthCare complaint below alleges that the vendor’s tool inserted a statement that the patient “consented” to recording when, the plaintiff says, no one asked. An attestation you did not write, in a note you signed, is Mr. Alvarez’s problem in a different sentence.

Disclosure. Whether the patient knows an AI is in the room is an ethical question first, and more and more a legal one. One academic center studied its own ambient-documentation pilot. The most common consent process was a spoken conversation before the visit. 81.6% of patients consented when given basic information. That fell to 55.3% when they were told about data storage and the company involved. Asked who should answer for a documentation error linked to the tool, 64.1% of patients said the physician (Lawrence et al., 2025). Those patients have already decided who wrote the sentence.

These are complaints, not findings. Nothing in them has been decided by a court, and you should treat them as allegations. The claims are based on California’s all-party-consent recording law, its medical-confidentiality statute, and the federal Wiretap Act. In the Sutter case, the plaintiffs argue that following HIPAA is not, on its own, a defense to recording without informed consent. About a dozen states require every party’s consent to record a conversation. Whether a waiting-room sign meets that requirement is what the courts are being asked. the patient-communication chapter covers how to have the conversation. This chapter’s point is that the conversation is the safeguard, not the sign.

The contract. HIPAA is the federal health privacy law. It lets a covered entity, such as your hospital, share protected health information with a vendor only under a business associate agreement (BAA). That is a contract that makes the vendor legally responsible for protecting the data. An institution-licensed scribe or a model built into the EHR runs under one. The consumer chatbot on your phone does not. At CU Anschutz, ChatGPT Edu and Microsoft Copilot are covered when you sign in with university credentials; the same tools through a personal account are not. The month-one boundary case from the appraisal chapter, a tool your hospital licenses inside the EHR, depends on one question: is there a BAA for this tool, and does the policy say what I may put into it? If nobody can answer in a sentence, the answer is no.

NoteYour right to see inside the EHR’s tools: status as of September 2026

the appraisal chapter said you have something close to a legal right to see how a decision-support tool inside your EHR was built and tested, and pointed here for its status. Here is the answer. The federal HTI-1 rule (2024) requires certified EHRs to show “source attributes” for predictive decision-support tools: what they were trained on, how they were validated, whether fairness was tested. It is in force. Developers began collecting a year of real-world performance data in January 2026. But in January 2026 the same office proposed a rule to loosen this, called HTI-5. It would remove the “model card” source-attribute requirements entirely. The stated reason is that clinicians rarely opened them. The comment period closed on 27 February 2026. As of this writing, HTI-5 has not been finalized (ASTP/ONC, HTI-5 proposed rule). So: the right exists today and may not next year. The scavenger hunt below asks someone to check.

11.3.8 Rules elsewhere, and which way things are moving

The European Union wrote the most comprehensive AI statute in force. Under the AI Act, software that is already a regulated medical device automatically counts as high-risk. High-risk systems carry extra duties, for risk management, data governance, human oversight, and record-keeping, on top of device certification. Those duties were due to apply in August 2026. In July 2026, under a “Digital Omnibus” amendment, they were delayed: to December 2027 for standalone high-risk systems, and to August 2028 for AI built into regulated products, which is where medical devices sit (Regulation (EU) 2026/1744; see the Orrick summary, July 2026). Over the same period, the United States withdrew its 2023 executive order on AI safety, loosened rules at both FDA and ONC, and tried, so far without a statute, to override state AI laws. Note the contrast. The EU delayed binding oversight under pressure. The US removed federal oversight and is now arguing with its own states about whether they may add any. Both facts will be out of date within a year. That is not a reason to skip them. It is the reason the Do block exists.

11.3.9 What you keep

The last idea is not legal. Four Polish endoscopy centers adopted AI polyp detection. Experienced endoscopists doing colonoscopy without the AI then found precancerous polyps less often than before. Their adenoma detection rate fell from 28.4% before exposure to 22.4% after, a drop of six points (Budzyń et al., 2025). The tool did not malfunction. The clinicians changed. Every framework in this course says the physician must stay “in the loop.” This is the study that shows the loop can quietly change the physician. The centaur-or-cyborg chapter asked you to write a policy for which tasks you hand over. This chapter asks the harder version. Which skills will you keep practicing without the tool, so that on the day the tool is wrong, you are still the person the standard of care assumes?

11.4 Watch or listen

Podcast. Stanford Legal, “Exploring AI in Healthcare: Legal, Regulatory, and Safety Challenges,” with Michelle Mello and Neel Guha (November 21, 2024; 48 min). A full transcript is on the episode page. https://law.stanford.edu/stanford-legal/exploring-ai-in-healthcare-legal-regulatory-and-safety-challenges/

Mello and Guha wrote the NEJM liability analysis cited above (Mello & Guha, 2024). Twenty minutes is enough. Listen for three things. First, the list of who can be sued when a tool causes harm: the hospital for how it chose and monitored the tool, the physician for failing to catch the error, and the developer for the design. Second, the point that courts will ask whether the physician should have known better than to rely on this output, rather than comparing the AI with an unaided human. Third, the pediatric example of a tool used outside the population it was built for. Then ask which of the three defendants Mr. Alvarez’s lawyer would name first, and why.

If you would rather read. The NEJM piece itself, or the Price, Gerke, and Cohen chapter on liability linked above, which is free on the NCBI Bookshelf. The older NEJM AI Grand Rounds archive has no episode centered on liability as of this writing. If one appears, it belongs here.

11.5 Do

WarningNo patient information, ever

Nothing in this exercise involves a real patient, a real note, or a classmate’s data. Mr. Alvarez is synthetic. If you use a chatbot at all in Part A, use it to find documents, not to answer the question. The answer must come from the primary source, with its URL and the date you checked it.

What counts as patient information, and why taking out the name is not enough: the patient information page.

11.5.1 Part A: the regulatory scavenger hunt (20 minutes)

This chapter has just told you several things that were true on 7 September 2026. The exercise is to find out whether they still are. Pick two questions below, one from each list. For each, find the current answer from the primary source: the agency, the statute, the court docket (the court’s own record of the case), or the organization’s own page. Record the URL and the date you checked.

List 1: answered in this chapter. Has it moved?

  1. What is the issue date on FDA’s current Clinical Decision Support Software guidance, and do the four criteria still read as listed above?
  2. Has ONC finalized HTI-5? If so, does the decision-support “model card” requirement still exist?
  3. Has FDA issued generative-AI guidance (not a discussion paper) for AI-enabled devices?
  4. What is the current EU AI Act deadline for high-risk AI built into medical devices?
  5. Has any US malpractice claim involving a clinical AI tool reached a verdict? (Search the legal press. If you find one, you have found something this chapter’s author could not.)

List 2: open when this chapter was written.

  1. What do ACGME or LCME currently require of residency programs or medical schools regarding AI training? Find the standard’s own text.
  2. What is the status of the National Academy of Medicine’s AI Code of Conduct, and where is the current version published?
  3. Has a patient-facing clinical chatbot safety incident been documented since the 2023 National Eating Disorders Association “Tessa” case? Name it and its primary source.
  4. What has happened in the Sharp HealthCare or Sutter Health ambient-scribe cases since April 2026? (Court docket or a law-firm summary, not a vendor blog.)
  5. Does your own institution have a written policy on ambient documentation, on consent for it, and on which AI tools carry a business associate agreement? Find the document, or record who told you none exists.

Record it like this. This is what you hand in.

Question ___ (List 1)
Primary source URL: ______________________________________________
Date checked: ____________   Document date shown on the source: ____________
Has it moved since 2026-09-07?   Y / N / cannot tell
If yes, what changed, in one sentence: ____________________________________
Did two credible sources disagree?  Y / N   If yes, which and how: __________

Question ___ (List 2)
Primary source URL: ______________________________________________
Date checked: ____________
The answer, in one sentence: _______________________________________________
Confidence that this is the *current* answer (1-5): ___   Why: ____________

Students usually bring back the same finding. Several of these have no clear answer even at the primary source, and two credible outlets say different things. That is not a failure of the exercise. It is the honest state of the field. Being able to say “the primary source is dated X and says Y, and I could not confirm Z” is a skill you will need in five years, when everything specific in this chapter has been replaced.

11.5.2 Part B: your side of the case (10 minutes)

Go back to Mr. Alvarez. Answer in writing:

  1. Which box of Table 11.1 is the note in, and why? (One sentence. If you think it is not on the grid, say what the “recommendation” was.)
  2. What would a plaintiff have to prove, element by element, and which element is hardest for them?
  3. The single fact that, if changed, would reverse your answer. This is the real product of the exercise. Was it the time of day? The number of notes? Whether the attestation line was inserted automatically? Whether the clinic’s policy required reading the plan against the conversation? Whether Mr. Alvarez knew the scribe was running? Name one, and say which way it changes your answer.
  4. Which of the four CDS criteria would the sepsis alert on your next inpatient rotation fail, and does it matter to you that it is a device?

You are not asked what a court would decide. There is no case law to be right about. You are asked what a court should decide, and why.

11.6 Check

Score yourself against this before moving on.

You have Meets Falls short
Two scavenger-hunt answers Each has a primary-source URL, the date you checked, and the date on the source A secondary summary, or no date
The “has it moved” judgment You can say whether the source is newer than 2026-09-07 and what changed “Seems the same”
A box for the note A box of Table 11.1, with the recommendation named “It depends”
The four elements Duty, breach, causation, damages, and which is hardest here “The AI was wrong”
The deciding fact One specific fact and the direction it changes the ruling A list of everything that matters
The device question Which criterion the alert fails, in your own words “It’s FDA-cleared so it’s fine”
The Jabbour sentence You can say what explanations did and did not do, with the direction of the numbers “Explainable AI helps”
The PCCP sentence You can say what clearance covers when the model has been updated “It was cleared”
Your position Three sentences below, with a condition that could prove you wrong A slogan

11.7 What this changes for you

This is the closing question of the course, and it is the reason this chapter is last. Write three sentences. Do not write a slogan.

What part of clinical judgment will you not hand over to an AI system, and what would have to be true for you to change your mind?

The second half is the important part. Without it, you will write something that sounds like a value. With it, you have to state a condition under which you would be wrong. That turns a first reaction into a position you can defend and, more importantly, revise. Go back to the policy you wrote in the centaur-or-cyborg chapter and put this sentence at the bottom of it. Then put Mr. Alvarez at the bottom of that. It is 6:40 p.m., the draft reads well, and you now know what a sentence you did not write can cost. What, specifically, do you do differently with the fourteenth note?

Date the file. Reread it at the end of intern year. Half of your class will have changed their answer, and the useful skill is noticing why.

11.8 Summary

  • No US malpractice claim involving a clinical AI tool has reached a verdict. What exists is a framework, a line, one experiment, and your position.
  • The physician is measured against the standard of care. The AI is evidence, not a defendant. “The AI was wrong” is not a legal argument.
  • The riskiest box is following a tool into nonstandard care. The second is overriding a tool that was right. In two studies, jurors reward following the tool when it agrees with the standard.
  • A biased model cut clinician accuracy by 11 points, and explanations did not significantly repair it.
  • Four criteria draw the device line. The fourth, whether you can independently review the basis with time to do so, decides most cases. The guidance moved twice in January 2026.
  • A PCCP is how a changing model stays cleared. Cleared is not validated.
  • You sign it, you answer for it. The patient has to know. And there is a contract, or there is not.
  • The tool can change you. Decide what you keep.
  • Mr. Alvarez’s sentence was written by nobody. Your name is on it.

11.9 Go deeper

Papers

Primary sources to bookmark (check the date on each before quoting it)

Talks

Adams, L., Fontaine, E., Lin, S., Crowell, T., Chung, V. C. H., & Gonzalez, A. A. (2024). Artificial Intelligence in Health, Health Care, and Biomedical Science: An AI Code of Conduct Principles and Commitments Discussion Draft. NAM Perspectives, 2024. https://doi.org/10.31478/202403a
Al-Anezi, F. M. (2026). Generative Artificial Intelligence in Healthcare: Automation Bias, Deskilling, and Cognitive Implications - A Systematic Review. Journal of Healthcare Leadership, 18, 590498. https://doi.org/10.2147/jhl.s590498
Anderson, T. N., Mohan, V., Dorr, D. A., Ratwani, R. M., Biro, J. M., & Gold, J. A. (2025). Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters. Mayo Clinic Proceedings. Digital Health, 3(4), 100292. https://doi.org/10.1016/j.mcpdig.2025.100292
Badal, K., Lee, C. M., & Esserman, L. J. (2023). Guiding principles for the responsible development of artificial intelligence tools for healthcare. Communications Medicine, 3(1). https://doi.org/10.1038/s43856-023-00279-9
Budzyń, K., Romańczyk, M., Kitala, D., Kołodziej, P., Bugajski, M., Adami, H. O., Blom, J., Buszkiewicz, M., Halvorsen, N., Hassan, C., Romańczyk, T., Holme, Ø., Jarus, K., Fielding, S., Kunar, M., Pellise, M., Pilonis, N., Kamiński, M. F., Kalager, M., … Mori, Y. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: A multicentre, observational study. The Lancet. Gastroenterology & Hepatology, 10(10), 896–903. https://doi.org/10.1016/s2468-1253(25)00133-5
Chouffani El Fassi, S., Abdullah, A., Fang, Y., Natarajan, S., Masroor, A. B., Kayali, N., Prakash, S., & Henderson, G. E. (2024). Not all AI health tools with regulatory authorization are clinically validated. Nature Medicine, 30(10), 2718–2720. https://doi.org/10.1038/s41591-024-03203-3
Jabbour, S., Fouhey, D., Shepard, S., Valley, T. S., Kazerooni, E. A., Banovic, N., Wiens, J., & Sjoding, M. W. (2023). Measuring the Impact of AI in the Diagnosis of Hospitalized Patients: A Randomized Clinical Vignette Survey Study. JAMA, 330(23), 2275–2284. https://doi.org/10.1001/jama.2023.22295
Lawrence, K., Kuram, V. S., Levine, D. L., Sharif, S., Polet, C., Malhotra, K., & Owens, K. (2025). Informed Consent for Ambient Documentation Using Generative AI in Ambulatory Care. JAMA Network Open, 8(7), e2522400. https://doi.org/10.1001/jamanetworkopen.2025.22400
Mello, M. M., & Guha, N. (2024). Understanding Liability Risk from Using Health Care Artificial Intelligence Tools. The New England Journal of Medicine, 390(3), 271–278. https://doi.org/10.1056/nejmhle2308901
Payne, P. R. O., Johnson, K. B., Maddox, T. M., Embi, P. J., Mandl, K. D., McGraw, D., Saria, S., & Adams, L. (2025). Toward an artificial intelligence code of conduct for health and healthcare: Implications for the biomedical informatics community. Journal of the American Medical Informatics Association : JAMIA, 32(2), 408–412. https://doi.org/10.1093/jamia/ocae306
Price, W. N., Gerke, S., & Cohen, I. G. (2019). Potential Liability for Physicians Using Artificial Intelligence. JAMA, 322(18), 1765–1766. https://doi.org/10.1001/jama.2019.15064
Tacconelli, A., Merane, J., Nielsen, A., Tobia, K., Hackanson, B., & Stremitzer, A. (2026). How Following Medical Artificial Intelligence Advice Can Mitigate Malpractice Liability: Cross-National Insights from a Randomized Trial. Journal of Nuclear Medicine : Official Publication, Society of Nuclear Medicine, 67(9), 1474–1480. https://doi.org/10.2967/jnumed.126.272292
Taylor, S. L., Jost, M., MacDonald, S., Ren, Y., Hilton, S., Davenport, S., Aizenberg, D., Hall, B., Lyles, C. R., & Adams, J. Y. (2026). Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study. JMIR Medical Informatics, 14, e86474. https://doi.org/10.2196/86474
Tobia, K., Nielsen, A., & Stremitzer, A. (2020). When Does Physician Use of AI Increase Liability? Journal of Nuclear Medicine : Official Publication, Society of Nuclear Medicine, 62(1), 17–21. https://doi.org/10.2967/jnumed.120.256032
Wolfe, J., & Welsh, S. M. (2026). Ambient Documentation Systems in Emergency Medicine: A Scoping Review of Clinical Precision, Patient Experience, Throughput, and Quality. Cureus, 18(4), e106643. https://doi.org/10.7759/cureus.106643
Wu, E., Wu, K., Daneshjou, R., Ouyang, D., Ho, D. E., & Zou, J. (2021). How medical AI devices are evaluated: Limitations and recommendations from an analysis of FDA approvals. Nature Medicine, 27(4), 582–584. https://doi.org/10.1038/s41591-021-01312-x