flowchart LR
accTitle: Two ways to use the same model
accDescr: Flowchart with two boxes. Left, Chatbot, your prompt leads to the model writing text, then to you copying what you want. Right, Coding agent, your goal leads to plan a step, then act by writing a file or running code, then observe the result, which loops back to plan a step or ends by showing you what it built.
subgraph C["Chatbot"]
direction TB
p1[Your prompt] --> t1[Model writes text]
t1 --> u1[You copy what you want]
end
subgraph A["Coding agent"]
direction TB
p2[Your goal] --> plan[Plan a step]
plan --> act[Act: write a file, run code]
act --> obs[Observe the result]
obs --> plan
obs --> done[Show you what it built]
end
7 Agentic coding: build a clinical tool in an hour, then find what is wrong with it
AIM-5 · AIM-2 · AIM-4 · about 75 minutes
By the end of this chapter you will be able to:
- Say what a coding agent does that a chatbot does not, and why that changes who is responsible for checking the result.
- Build a working single-page clinical calculator from a plain-language prompt, in a browser, with no install.
- Verify the generated clinical logic against the original publication or MDCalc, and find at least one discrepancy or edge-case failure in the tool you just built.
- Explain why that tool is not a medical device, and name what would have to change for it to become one.
- State the rule about patient information in consumer and institutional AI tools, and say why it is not a formality.
Time. About 20 minutes of reading, 15 minutes of watching, 35 minutes of building and checking, and 5 minutes for the Check. You need your institutional ChatGPT Edu account and one tab open on MDCalc, the free calculator website most residents use. (ChatGPT’s canvas is a side panel that shows the page you build, running.) Nothing installs. Other tools, and what each one asks of you at signup, are listed in the Do block.
7.2 Why this matters
You will see a tool like the resident’s within the year. They are cheap to make, and every team has someone who makes them. Some are spreadsheets. Some are web pages. More and more, they are built by describing what you want to a model and accepting what comes back. The person who built it is usually proud of it. The attending is usually pleased. Nobody has been asked to check it, because checking a colleague’s tool is not a job anyone is assigned.
Unchecked tools that a whole team relies on have a history older than language models. In October 2020, Public Health England’s official figures missed nearly 16,000 positive COVID-19 results over a week. Days later, only half of those patients had been contacted for tracing (Mahase, 2020). The reported cause, in the clippings below, was that results were being collected in an Excel file format with a row limit. Rows past the limit were silently dropped. The spreadsheet did not crash. It did what it was told.
A quieter version appears in the research literature. Roughly one in five genomics papers with gene lists in Excel contained gene names that the spreadsheet had silently turned into dates (Ziemann et al., 2016). The tool was fine. The tool was not checked against what it was being used for.
In three months you will be the intern with the bookmark, or the intern who built it. Either way the skill you need is the same. Take a tool that seems to work, decide what it would take to trust it, and go and look.
7.3 How it works
7.3.1 A chatbot answers; an agent acts
Everything you have used so far in this course is a chatbot. You type, it writes, you copy what you want out of the reply. A coding agent is the same kind of model wrapped in a loop. It plans a step. It takes an action in the world: writes a file, runs a program, reads the output. It looks at what happened, and plans the next step. The actions are real. Files change, code runs. Because the actions are real, agents run with permissions, a list of what they are allowed to touch. Figure 7.1 shows the two shapes.
The tools you will use today are somewhere between the two. Your chatbot opens a side panel, writes a web page into it, and shows you the running page. That is the smallest possible version of the loop: the model wrote a file and the browser ran it. You are the one watching the result. The full agents work inside a folder of files on your own computer. They are mentioned at the end of the chapter and are not used here on purpose. They need an install and a terminal, and neither is the point.
The point is what changes when the output is something that runs. A wrong sentence in a chatbot reply sits there until you read it. A wrong rule in a calculator does its work silently, on every patient, until somebody compares its answer to another source. The error comes from the same place as in the critical appraisal chapter: the model produces a plausible next piece of text, not a looked-up fact. But now the error has a place to hide.
7.3.2 Two names for the same thing
In February 2025 Andrej Karpathy, a founding member of OpenAI, invented the phrase “vibe coding”. It means describing what you want, accepting what the model produces, and never reading the code. Collins Dictionary named it word of the year for 2025. A year later Karpathy wrote that the phrase was out of date and, as widely reported, proposed “agentic engineering” instead: directing and reviewing an agent’s work rather than accepting it VERIFY: wording of Karpathy’s 4 Feb 2026 post; the post exists but its full text could not be read on 2026-09-08, and the “agentic engineering” phrasing is from trade coverage.
Both terms will be gone before you finish residency. The shift between them is the part that lasts. The first name treats your judgment as something you set aside. The second treats it as the scarce input. You are about to do the first for fifteen minutes so that you can practise the second for twenty.
The most useful way to think of the model, borrowed from the author’s talks on the subject, is as a capable but junior colleague: fast, fluent, willing, and in need of clear instructions and review. You review its work the way a senior reviews a junior’s. You do not trust the confident tone. You check the parts that would hurt someone if wrong.
7.3.3 Three ways the tool you just built is wrong
A generated calculator can fail in three ways, and each can happen on its own. Each has a matching check in the Do block.
| Failure | What it looks like | How you catch it |
|---|---|---|
| The rule | A point value, cutoff, or coefficient is wrong, or a rule from the original paper is missing (the dialysis rule; the sodium limits; whether a score is only applied above a threshold) | Compare every rule to the original publication, not to your memory of it |
| The edge | The tool does something the builder never imagined, for inputs the builder never tried: nothing checked, two boxes checked that cannot both be true, a value outside the range the score was built on | Try the inputs the builder did not: empty, contradictory, extreme |
| The version | The score has more than one published form, and the tool does not say which one it uses (Wells’ criteria exist in a three-tier and a two-tier form; the transplant system replaced MELD-Na with MELD 3.0 in 2023) | Ask which version, and check that against what your team, and MDCalc, actually use |
The first row is where models are weakest and where the evidence is clearest. When ChatGPT was asked to perform 48 kinds of medical calculation directly, it got one trial in three wrong. Giving the model a purpose-built calculation tool to call, instead of doing the arithmetic itself, cut the error rate thirteenfold, from 64% to 4.8% (Goodell et al., 2025). Read that finding carefully, because it says two things. A model doing clinical arithmetic in its reply is unreliable. A model writing code that does the arithmetic is much better. That is exactly why the tool you build today will mostly work. But the code only contains the rules the model put in it, and the rules it leaves out leave no trace.
The second row is where non-programmers are weakest. Medical students spent three hours building health apps by prompting. Every group produced a working prototype, and no group added any protection for the data the app collected. The authors’ summary was that making it easier to build is not the same as making it easier to build safely (Quispe et al., 2026). An editorial in the laboratory medicine literature gives the pattern a name you will recognise from the wards, the illusion of competence. Code that is correctly written and looks professional is most convincing to the people least able to find its flaws. There is a real difference between code that runs and code that is safe to run on patients (De Bruyne et al., 2026).
The third row is the one people notice least. A nursing research group used a coding agent to run a statistical analysis (a logistic regression) and checked every output against a published reference. The coefficients and p values matched. An error in one confidence interval was missed on first review and caught only on a second review (Tolentino et al., 2026). The tool was mostly right, which is the condition under which people stop checking.
7.3.4 What “checked” looks like
The counterexample is worth knowing because it shows the standard can be met. A cardiology group built a web calculator for three acute coronary syndrome risk scores, with ChatGPT writing the code. They then compared its output against 226 registry patients whose scores had been computed independently. Agreement was exact (r = 1.000 for both scores that had a registry comparison), and the paper says so in the methods, with the registry named (Egaimi et al., 2025). That is not a large study. It is not clinical validation, which would mean showing the tool helps patients, and the authors are careful to say so. It is what a technical check looks like: an outside source of truth, a fixed set of cases, and a written result.
The resident’s tool in the workroom had none of those. Yours, in thirty minutes, will have a small version of all three.
7.3.5 The named hazard
A 2026 framework for governing clinician-built software, VIBE-HI, names one hazard as the most important and gives it a name: comprehension abdication. It means handing your understanding over to the system that generated the code. If you cannot say in plain language what your tool does for every input, you cannot predict how it fails, and you cannot safely change it (Alqheedan & Alzughaibi, 2026). The framework sorts tools into four risk levels, or tiers, using four questions. Table 7.2 redraws them from the paper’s description.
| The question | Pushes the tier up when |
|---|---|
| Does it touch identifiable patient data? | Real names, medical record numbers, or anything that could be traced back to a person |
| How close is it to a clinical decision? | Its output changes what is done to a patient, rather than what a student studies |
| Who uses it, and how much do they understand it? | People other than the builder, who did not see it made |
| How many patients does it reach? | A team, a service, a hospital |
Run the resident’s tool through the four questions. It handles pasted labs, so the first answer depends on what people paste, which nobody controls. Its output is a listing decision. It is used by a whole team, most of whom never saw it built. It reaches every patient on the service. A study aid for one person answers all four questions the other way. The tool did not change between those two descriptions. Its use did.
7.3.6 Where a toy becomes a device
In the United States, the FDA decides whether software like this counts as a medical device. For tools like this one, the decision depends on the FDA’s criteria for non-device clinical decision support, the category of software that advises clinicians without being regulated as a device. The criterion that matters for a calculator is whether the clinician can independently review the basis for the recommendation. A tool that shows the score, the inputs, the rule it applied, and where the rule came from is one a clinician can check. A bare number is not. A personal study aid that a student uses to practise scores is not a device. The same page, bookmarked on the workroom computers and used to decide who gets listed, is a different question. The answer depends on what it shows and who relies on it. The FDA reissued its decision-support guidance on 29 January 2026, replacing a version dated 6 January. What the revision changed, and the rest of the regulation, belong to the ethics and regulation chapter.
For this chapter, the working rule is simpler than the regulation. Toy, not device. Your tool carries a banner saying so, in the prompt, before it exists. That is not done for show. It is the professional habit of labelling AI-assisted work, rehearsed on something harmless.
The course’s standing rule, you sign the note, covers a number you copied from a calculator. When the team’s number is wrong, the person accountable is the clinician who acted on it, not the resident who built the page. Regulation may one day hold the builder responsible; the chart holds you responsible now. The practical consequence for a note: record the score and the source you took it from. Make the source one that has been checked by someone other than a colleague. Today that means the original publication or an established calculator such as MDCalc. If a team tool has been validated locally, your institution’s policy will say so, and the ethics and regulation chapter covers what to document when it has.
7.3.7 Patient information, once more
At CU Anschutz, ChatGPT Edu is covered by a business associate agreement when you sign in with university credentials. A business associate agreement is the contract that makes a vendor legally responsible for protecting patient data. It covers patient care, not a tool you built this afternoon. Without one, as with a personal account, pasting a real patient’s information into the tool can be an unauthorized disclosure, whatever you intended. Lab values alone may not identify anyone, but you are not the person who decides that; see what counts as patient information. Consumer and free tiers never have one, and free tiers may use what you type to train the next model.
The rule for today, and for every tool you build until someone with authority tells you otherwise: every number you type is one you invented. Not a de-identified real patient, not a classmate, not yourself. Invented. The habit you are rehearsing, a chat window with a patient’s labs in it, is exactly how patient data leaks when the same window is reused for real work.
Around 40% of programs generated by an early coding assistant in security-relevant scenarios contained a vulnerability (Pearce et al., 2022). In a controlled study, people with an AI assistant wrote less secure code and were more confident that it was secure (Perry et al., 2023). Both studies predate current models and the numbers will have changed. The direction is what matters. You are not going to security-review a page you built in fifteen minutes. So the rule is not “make it secure”; it is put nothing in it worth stealing. No passwords, no identifiers, no real data.
7.4 Watch or listen
Video. Andrej Karpathy, Software Is Changing (Again), keynote at Y Combinator’s AI Startup School, San Francisco, 17 June 2025 (published 19 June 2025; 39 min). https://www.youtube.com/watch?v=LCEmiRjPEtQ. Also on the Y Combinator podcast feed if you prefer audio. YouTube’s captions are automatic and adequate.
Fifteen minutes is enough. Skip the opening on “Software 3.0” if you are short of time and go to the middle third. There he describes partial autonomy apps and the autonomy slider. The idea is that the useful question about an AI tool is not whether it is autonomous, but how much freedom you have given it for this task, and how fast you can check what it did. Listen for the phrase about keeping the AI “on a leash.” Then ask, for the tool you are about to build, where the slider is set and who is holding the leash. The person who invented “vibe coding” is also the person telling you to verify.
Instead, or as well. The author’s own slides: From Chatbots to Coding Agents (June 2026) goes through the plan-act-observe loop in Figure 7.1 with live examples. The recording of the October 2025 live vibe-coding session from this course’s predecessor shows a calculator being built in front of a class, with the mistakes left in.
7.5 Do
Every value you type into the tool today is one you made up. No patient from a rotation, however de-identified you think it is; no classmate; no family member. Consumer tools carry no business associate agreement, and your university account’s agreement covers patient care, not testing a tool you built in an hour. If a test case needs a sodium of 118, invent a patient with a sodium of 118.
ChatGPT Edu (your institutional account). Ask for the page and it should open in the canvas side panel, with a preview button that runs it. This is the expected route. If canvas does not appear, your workspace administrator may have turned off canvas code execution; use the fallback below. VERIFY: log in to the institutional ChatGPT Edu workspace and run the starter prompt the week before the cohort; OpenAI’s admin documentation confirms canvas code execution is a workspace setting, checked 2026-09-08
Fallback that works with any chatbot, including Microsoft Copilot. Ask for “a single HTML file”. Copy the whole code block into a plain text editor (Notepad, or TextEdit in plain-text mode). Save it as calc.html and double-click it. Your browser opens it. Nothing is installed. You have made a file, which is what the agent would have done for you.
Claude.ai free tier builds and hosts the page as what it calls an artifact, a page with a shareable link, which is convenient for the live session. Signup requires a phone number. If you would rather not give one, you lose nothing by using the options above. Google AI Studio needs a Google account and no phone; it is the second alternative.
Free tiers may use your prompts for training. That is fine for invented numbers and a prompt about a public scoring system. It would not be fine for anything else.
7.5.1 Part A: build it (12 minutes)
Step 1. Pick a score. Choose the one closest to your specialty from Table 7.3. Each has a single original publication you will check against, and each has at least one edge the model tends to miss. If none fits, use CHA₂DS₂-VASc; it is the one the worked example below is written for.
Step 2. Paste the starter prompt, with your score substituted, into a fresh conversation. Use it word for word the first time. The safety wording is inside it on purpose. What comes back depends on how the request is phrased, and the same prompt produces different code on different days. Using the fixed text means your result reflects the model, not how fluently you write prompts. If English is not your first language, that is the point.
Build a single-page web app, as one self-contained HTML file, that calculates the
<score>. Show each input as a checkbox or numeric field with its point value or coefficient visible. Display the running total and the interpretation the original publication attaches to it. Add a prominent red banner: “Educational prototype only. Not validated for clinical decision-making.” In a footnote, cite the original publication for the score and say exactly which version of the score you implemented, so I can independently verify every rule against MDCalc and the paper before I trust the output.
Step 3. Ask for two changes. Whatever comes back, ask for two changes, one at a time. Watch what the model does to the rest of the page each time:
- “Colour the interpretation by risk tier.”
- “Add a ‘clear all’ button, and tell me what the score shows when nothing is entered.”
The second request is really a check, phrased as a feature. Read the answer.
You now have a working tool. It took about ten minutes, and it probably looks better than the resident’s. Do not share it yet.
7.5.2 Part B: break it (15 minutes)
Open MDCalc and the original paper from Table 7.3 side by side with your tool. Run three invented cases through all three sources and record the results in the template below.
Case 1, ordinary. A patient clearly inside the population the score was made for, with nothing unusual. All three should agree. If they do not, you have found a rule error and can stop here, because you have already met the objective. Note it and continue anyway.
Case 2, the edge. Take the edges column for your score and build a patient whose values fall on one of them: a dialysis patient for MELD-Na, a sodium of 118, a Glasgow score of 14 for qSOFA, a 6 mm subsolid nodule in a patient under 35 for Fleischner, a 74-year-old and a 75-year-old for CHA₂DS₂-VASc. Compare the three sources.
Case 3, the nonsense. Give the tool something no patient has. Both age bands ticked. A negative creatinine. Every box checked and then every box unchecked. A blank form. The question is not whether the tool refuses (it usually will not). The question is whether you can say, before pressing anything, what it will do, and whether what it does is safe or silent.
Expected behaviour, not a recorded run. The two examples below are built from the failure patterns in Table 7.3 and the published evidence above. A live run of the starter prompt was attempted for this chapter and did not complete in time to be recorded. So nothing here is presented as what a particular model produced on a particular day. Run the prompt yourself and compare.
CHA₂DS₂-VASc, expected. The page will almost certainly get the eight point values right and add them up correctly. That is the easy part. The one-in-three arithmetic failure above is about the model doing sums in prose, not code. Look instead at three places. First, the two age boxes. If both “65–74” and “≥75” can be ticked at once, the page can produce a score of 3 for age, which no patient has. Second, the stroke-risk percentages beside each score. The page will print a column of numbers with a footnote naming Lip et al. 2010. But the numbers most tools print come from a different group of patients (Friberg et al., 2012), and the two tables disagree. If the footnote does not say which, that is the version failure. Third, the “female sex” box. It scored 1 in the 2010 paper and is treated differently by newer guidelines. A page that scores it without saying which guideline it follows has chosen for you.
MELD-Na, expected. The formula will be right. What tends to be missing is exactly what the opening scenario depended on. No dialysis field, so creatinine is never set to 4.0. Values below 1.0 are not raised to it. No limits on sodium, so a sodium of 118 enters the equation as typed. And no check that the sodium term applies only above the MELD threshold. Ask the page for a patient on dialysis and one with a sodium of 118, then compare to MDCalc.
Your run will differ from these expectations in both directions. Models change monthly, and the same prompt produces different code on different days. The pattern is what to look for. The arithmetic is usually right; the rules at the edges are where to read.
Record it. Copy this into a document. It is what you bring to the live session.
Tool used: ___________________ Model version if shown: ____________
Score built: __________________ Version the tool says it implements: ________
Case Inputs (invented) My tool MDCalc Paper Agree?
1 ordinary _______________________ _____ _____ _____ Y / N
2 edge _______________________ _____ _____ _____ Y / N
3 nonsense _______________________ _____ _____ _____ Y / N
Discrepancy or edge failure found (at least one; "none" needs a reason):
________________________________________________________________________
Which of the three failures in the table is it: rule / edge / version?
________________________________________________________________________
The one change I would make before letting anyone else use it:
________________________________________________________________________
“None found” is an acceptable answer only with a sentence saying which edges you tried. A tool that survived three cases is not validated. It is three cases old.
7.5.3 Part C: explain it, then place it (8 minutes)
The comprehension test. In four sentences, without opening the code, write what your tool does for: an ordinary patient, the edge case, the nonsense case, and a blank form. If you cannot write the fourth sentence, ask the model “what does the page show when nothing is entered, and why?”. Read the answer against the page, and then write it. This is the test the resident never took.
The device paragraph. One paragraph answering three questions. Is your tool, as it is, a medical device? (No; say why, using the phrase “study aid” and the word “banner”.) What would have to change for that answer to change? (Name at least two of the four questions in Table 7.2.) And if your tool showed only a bare score, with no inputs, rule, or citation visible, which non-device criterion would it fail?
The one sentence. Write the sentence you would say to a co-intern who shows you a calculator they built last night. It should contain a verb that means compare against something outside the tool.
7.6 Check
Score yourself against this before moving on.
| You have | Meets | Falls short |
|---|---|---|
| The chatbot-versus-agent distinction (objective 1) | You can say what the loop adds (acts on real files, observes, repeats, inside permissions) and why a running tool hides errors a reply does not | “An agent is a smarter chatbot” |
| A working tool (objective 2) | Runs in a browser, shows inputs and total, carries the banner, cites the paper and names the version | Runs, but the banner or citation is missing, or the version is unstated |
| Three recorded cases (objective 3) | Each compared against MDCalc and the original paper; at least one discrepancy or edge failure named, and sorted as rule, edge, or version | “It matched MDCalc” with no edge case tried |
| The comprehension test (objective 3) | Four sentences, including the blank form, written before opening the code | The fourth sentence is missing or copied from the model |
| The device paragraph (objective 4) | Says why it is a study aid now, names two of the four questions that would change that, and names the criterion a bare score fails | “It’s not a device because it’s just a prototype” |
| The PHI rule (objective 5) | You can say what a business associate agreement is, that your university-login ChatGPT Edu account has one and a personal account does not, and why invented numbers are the rule | “Don’t put patient data in ChatGPT” without the reason |
| The one sentence | Contains a verb meaning compare against an outside source | “Be careful with it” |
7.7 What this changes for you
What does this change about how you’ll practice? Write it down before you close this chapter. Two specifics. First, go back to the workroom. The coordinator is on the phone, the team’s number is 24, hers is 32, and the bookmark is on every computer. What do you say to her? What do you say to the resident who built it? What happens to the bookmark? Second, think of the next tool you build yourself, or are handed by a colleague. Name the three cases you will run through it before it touches a decision, and the outside source each one is compared against. One word is not an answer to either.
7.8 Summary
- A chatbot writes; an agent acts, observes, and repeats inside permissions you set. The tools you used today are the smallest version of that loop. The full ones need an install this chapter avoids on purpose.
- A running tool hides errors a reply does not. The arithmetic is usually right. The rules at the edges, and the version of the score, are where a generated calculator fails.
- A model doing clinical arithmetic in its reply is wrong about a third of the time. A model writing code that does it is far better, and still only contains the rules it was told about.
- Comprehension abdication is the named hazard: if you cannot say what your tool does for every input, you cannot say how it fails.
- Four questions move a tool from study aid toward device: patient data, closeness to a decision, who uses it, how many patients. The tool need not change for the answers to.
- Every number you type is one you invented. Nothing worth stealing goes into a page you will not security-review. The banner goes in the prompt.
- The coordinator was right, and the team’s spreadsheet was three cases old.
7.9 Go deeper
Papers
- VIBE-HI, the governance framework behind Table 7.2, with its tiers mapped to regulation (Alqheedan & Alzughaibi, 2026); and the laboratory-medicine editorial on the illusion of competence (De Bruyne et al., 2026).
- The clinical-calculation study behind the one-in-three figure (Goodell et al., 2025), and the cardiology group’s technical validation of a ChatGPT-built calculator against a registry (Egaimi et al., 2025), which shows what your Part B could become.
- Two views of the same practice from inside medicine: an editorial arguing it opens biomedical software to everyone (Moore & Tatonetti, 2025), and a medical-education piece that built two teaching simulations this way (Chow & Ng, 2025).
- The one evaluated medical-education app built by prompting, an ECG tutor compared against control sites, with mixed and inconsistent effects on exam performance (Al Janabi & Bland, 2026). Building fast is not the same as building something that works.
- The security studies quoted above (Pearce et al., 2022; Perry et al., 2023), for the numbers and their dates.
Talks
- From Chatbots to Coding Agents and AI Agents for Biomedical Knowledge Work, the author’s slides this chapter’s approach comes from; the second has the privacy-and-trust section in full.
If you want to go further
- The full agents, Claude Code, Codex, Gemini CLI, and GitHub Copilot in an editor, need a local install, a terminal, and in most cases a paid plan; new signups for the student Copilot plan have been paused since April 2026 with no announced end date, though Copilot Free still accepts signups VERIFY: pause still in force? checked 2026-09-08 on the GitHub changelog and community FAQ. Try one after the course, on a project with nothing in it worth stealing.
- Google Colab runs Python notebooks in a browser with a Google account and no install. Its built-in agent can run a whole data analysis, start to finish, from a prompt. This course’s predecessor used it for a public insurance-cost dataset. The completed notebook is a good thing to appraise: ask what the agent decided without being asked, and what you would want to know before believing its R².