flowchart LR
accTitle: The retrieval pipeline
accDescr: Flowchart in two columns. Left column, when you add a source, your PDF is chunked into passages, each passage is embedded as a list of numbers, and the numbers are indexed. Right column, when you ask a question, the question is embedded and the nearest passages are retrieved from the index, those passages are pasted into the context window, and the model generates an answer pointing at them.
subgraph I["When you add a source"]
direction TB
doc[Your PDF] --> chunk[1. Chunk: cut into passages of a few hundred words]
chunk --> embed[2. Embed: turn each passage into a list of numbers]
embed --> index[3. Index: store the numbers so similar ones are findable]
end
subgraph Q["When you ask a question"]
direction TB
q[Your question] --> retrieve[4. Retrieve: embed the question, find the nearest passages]
retrieve --> augment[5. Augment: paste those passages into the context window]
augment --> generate[6. Generate: write an answer, pointing at the passages]
end
index -.-> retrieve
13 Grounding AI in real sources: what a retrieval tool fixes, and what it cannot
AIM-1 · AIM-9 · about one hour
By the end of this chapter you will be able to:
- Explain what a context window is, and why a chatbot only seems to remember your conversation: the whole conversation is sent again every turn.
- State plainly that a language model does not learn from you, and say what does change between one turn and the next.
- Describe the six retrieval steps in order: chunk, embed, index, retrieve, augment, generate. Say what chunking does to a results table.
- Explain why a tool that answers from your sources invents citations far less often than a general chatbot, and why that comes from how it is built rather than from care.
- Catch the error that grounding does not fix: a real, correctly linked source that the summary has misread or stretched too far.
- Build a notebook from sources you chose, and say when that is and is not a literature search.
Time. About 20 minutes of reading, 7 minutes of video, 30 minutes of doing. You need a Google account for NotebookLM. A personal account is fine; see the access note in the Do block. You also need one general chatbot in a second tab, such as your institutional ChatGPT Edu or Microsoft Copilot. Nothing installs. NotebookLM’s free tier allows 50 sources per notebook and 50 chat questions a day, according to Google’s own help page (checked September 2026). That is more than this chapter needs, but do not spend the day’s questions in the first ten minutes.
13.1 Journal club is Thursday
You are three weeks into a sub-internship. The chief resident hands you a paper to present at Thursday’s journal club, and four more “for background.” The topic is whether AI tools that answer from a set of documents are ready for the clinic, and she knows you have opinions. It is Monday. You have five PDFs, two call shifts, and a personal statement to finish.
A co-intern says: put them all into NotebookLM and ask it your questions. You do. Ten seconds later you have this.
Every number in that answer is real. Every little blue “1” links to a passage that exists. The first sentence is going on your slide. And the first sentence is not what the paper says.
This chapter is about why the numbers are real, why the sentence is not, and how to tell in under a minute. the appraisal chapter showed you that a general chatbot invents references and a retrieval tool does not. This chapter looks more closely at how the tool works. How it works is what tells you which mistakes are still possible after the references are fixed.
Two words first, because vendors use both. A grounded tool is one that answers from documents it was given and points back to them. Retrieval is how it finds the right passages in those documents before it writes. The rest of the chapter explains both in detail.
13.2 Why this matters
In residency you will read less than you do now and rely on summaries more. Some of those summaries will come from tools like the one above. A grounded tool inside the EHR that answers from your hospital’s protocols. A literature tool that searches PubMed. A notebook you built from a rotation’s reading list. The vendors will tell you, correctly, that these tools “cite their sources.” What they will not say is what “cite” means, step by step. So they will not say which errors citing protects you from, and which it does not.
When a notebook misreads a paper, the result is a wrong slide. When the grounded tool inside the EHR misreads a protocol, the result can harm a patient. The clinician who repeated the sentence is responsible for it.
The second reason is smaller and closer. The same mechanism that makes a grounded tool trustworthy about where a sentence came from also explains three other things. Why a chatbot has no memory. Why nothing you type teaches it anything. And why pasting a patient’s note into a consumer tool is a question about data storage, not about the model learning. Those are the AIM-1 basics most clinicians get wrong, all in the same direction, and one hour fixes them.
13.3 How it works
13.3.1 The model has no memory
Start with something you can see. Open a chatbot and ask it who won the 2020 World Series, the baseball championship. It answers: the Dodgers. Ask “where was it played?” It answers correctly: Arlington, Texas, because of the pandemic. It seems to remember the first question.
It does not. Here is what the model actually received on the second turn:
[user] Who won the 2020 World Series?
[assistant] The Los Angeles Dodgers won the 2020 World Series ...
[user] Where was it played?
The entire conversation is sent again, every turn. The model then writes the next piece of text given all of it. That block of text is the context window: everything the model can see at the moment it answers. It is measured in tokens, which are words or pieces of words, and it has a fixed size. (The foundations chapter introduced the context window as a shared budget; this is the same object, now explained in full.) Today that size ranges from one long book to a small shelf of them, depending on the model. When you close the tab, the context is gone. The model is exactly what it was before you opened it.
Two things follow, and the rest of the chapter depends on both.
First, a language model does not learn from its users. Nothing you type changes the model’s weights, the billions of numbers set during training that make it what it is. What changes between turns is the context, and only the context. “It got better as we talked” means “the context got longer.” A new conversation starts from zero. Privacy is a separate question. The company may store what you type, and some free tiers use it to train the next model, months later. That is a question about how your data is handled, not the model remembering you. It is the reason for the no-patient-information rule in every chapter of this course.
Second, the context window is a budget, and the model does not use all of it equally. In the study the foundations chapter calls “lost in the middle”, researchers placed the one relevant document at different positions among many others. Accuracy was highest when it sat at the beginning or end of the context. It fell when the document was in the middle, in some settings below the model’s accuracy with no documents at all (N. F. Liu et al., 2024). A long context window does not mean the model reads all of it well. That matters for what comes next, because a retrieval tool’s whole job is to decide which few passages go into the window.
13.3.2 What “grounded” means, step by step
The appraisal chapter’s figure showed two ways a citation can be produced. A general model writes a reference as text. A retrieval tool searches first and writes second (see Figure 5.3 there). The search-first way has six steps. Together they are the retrieval pipeline, and each step is a place where something can go wrong. Google does not publish exactly how NotebookLM does each step. The six steps are the general design that every tool of this kind follows.
Chunking cuts the document into passages. A whole paper will not fit in the context window next to your question and forty other papers. A typical passage is a few hundred to a thousand tokens. One published clinical pipeline used 1,024-token chunks and reported 97.5% citation accuracy on landmark spine-surgery papers (Kurland et al., 2025). A dental-trauma tool built from five guideline sources was a knowledge base of exactly 250 chunks (Espona et al., 2026). Chunking is simple by design. It cuts on length, not meaning.
Embedding turns each passage into a list of numbers, called an embedding. Passages about similar things get similar numbers. “Apixaban reduced major bleeding” and “less hemorrhage with the factor Xa inhibitor” get numbers that are close together. “The patient’s dog is named Apixaban” does not. This is the method from the word-embeddings talk in Go deeper. You do not need to know more than “similar meaning, nearby numbers.”
Indexing stores those numbers so that, given a new list of numbers, the nearest ones can be found in milliseconds. An index is a phonebook sorted by meaning instead of by surname.
Retrieval turns your question into numbers the same way and returns the nearest few passages, usually between three and twenty. Augmentation pastes them into the context window, in front of your question, as plain text. Then generation is the ordinary thing a language model does: write the most likely next words. The difference is that now the most likely continuation of “according to the passages above” is something the passages actually say. A citation is a pointer back to a passage that is already in the window.
That is the entire method. The model did not become honest. It was given the answer and asked to paraphrase it.
13.3.3 What chunking does to a table
Here is the step most explanations skip, and the one that produces the strangest errors. Take a results table in a paper you uploaded. It is a few lines of text. If the chunk boundary falls inside it, the tool stores something like this. The block below is an illustration built from the differences the guideline study reports; it is not a screenshot of any tool.
chunk 7 of 23... rated by three blinded specialists.
Table 6. Judge accuracy by model, with and without retrieval (mean, 95% CI).
Model Without With Difference
Mistral Small 3.40 4.34 +0.94 (0.58–1.30)
Qwen3-Next 3.52 4.38 +0.86 (0.50–1.22)
chunk 8 of 23GPT-OSS-120B 3.62 4.34 +0.72 (0.36–1.08)
DeepSeek-V3.2 4.08 4.44 +0.36 (−0.01–0.72)
GPT-5-chat 4.18 4.48 +0.30 (−0.07–0.66)
The gain was large for the weaker models and small and not statistically significant for the already-strong bases ...
Chunk 8 has the rows that matter for the question in the opening scenario. Chunk 8 has no column headings. When it is retrieved alone, the model sees five numbers per line and has to guess what they mean. When only chunk 7 is retrieved, the model sees the headings and the two rows where retrieval helped most. It never sees the sentence that qualifies them. Both chunks are real. Both citations will resolve. Neither is the table.
Nothing in the basic pipeline knows that a table is a table. Nothing knows that a figure legend belongs to a figure, or that “not statistically significant” three lines later applies to the number above it. Better tools add a page-layout step that tries to keep tables whole, but from the outside you cannot tell whether yours does, or how well. Rows of a table cut apart. A footnote separated from the value it qualifies. A “however” at the top of the next page. These are the retrieval tool’s typical errors. They are errors of reading, made on a source that is real.
13.3.4 Why grounded tools fabricate less: the numbers
the appraisal chapter gave you the main result. On the 2/1/0 citation scale in Table 5.4 (2, accurate: the source supports the sentence; 1, misrepresented: the source is real but misdescribed; 0, fabricated: the source does not exist as given), a tool that indexes PubMed produced no fabricated citations, while general chatbots fabricated between 2% and 26% (McLaughlin et al., 2026). The steps above are the reason, and the same pattern repeats wherever it has been measured.
In one study, six models answered 50 questions from a head-and-neck oncology guideline, with and without retrieval of that guideline. Hallucination, the field’s word for made-up content stated as fact (here, claims that contradicted the guideline or invented specifics), fell from 42% to 4%. The share of citations that pointed at a real retrieved passage rose from 0% to between 51% and 89%, depending on the model (Vollmer et al., 2026). In implant dentistry, two retrieval systems answered with 0% fabricated references. A general model fabricated 82% of its references and 86% of its statistics (Hooshiar, 2025). A 2026 systematic review of 44 studies found retrieval was the most often tested way of reducing hallucination. It reported 30% to 50% reductions from retrieval alone (Basu & Huynh, 2026).
The drop comes from how the tool is built, not from the tool being careful. The same paper found the accuracy gain from retrieval was large for weak models. For strong models it was small and not statistically significant. The drop in hallucination, and the gain in how checkable the answers were, held across all of them (Vollmer et al., 2026). The authors add that the strong models may simply have been near the top of the scale already, so read the accuracy result with care. The safer summary is this: grounding does not reliably make a model smarter. It makes it checkable.
There is one more thing a grounded tool does that a general one cannot, and you will produce it yourself in the Do block. Asked a question its sources do not cover, a well-built grounded tool declines. A guideline chatbot for lung cancer was asked about cancers the guideline did not cover. It “consistently declined to answer and indicated that the available information was insufficient” (Nishisako et al., 2026). A general chatbot, asked the same question, writes a paragraph. “I cannot find that in your sources” is the sentence that tells you the tool is doing what it claims.
Every rate above is one tool, one field, one month. The study that used the 2/1/0 scale found the ranking of the general chatbots changed between question types (McLaughlin et al., 2026). The overgeneralization study below found newer models did worse than older ones (Peters & Chin-Yee, 2025). When your own run in the Do block disagrees with a number here, that is expected. The pipeline in Figure 13.1 is the part that lasts.
13.3.5 The failure grounding does not fix
Go back to the opening answer. Its sources are real. Its numbers are in Table 6. Its first sentence says retrieval “improved accuracy across all models.” The paper’s own words are that the gain was “small and not statistically significant for the already-strong bases.” The answer’s second sentence half-corrects this. But the second sentence is not the one that goes on the slide.
On the 2/1/0 scale this is a score 1: a real reference, but a wrong description of what it says. It is also the most common kind of error a grounded tool makes. And it is the kind that grounding, because of how it is built, cannot catch. The pipeline guarantees that the text came from the passages. It does not guarantee that the model read the passages the way the authors meant them. The dentistry study’s phrasing is exact: the retrieval systems “avoided fabricated references but still generated incorrect interpretations” (Hooshiar, 2025).
The misreading usually goes one way. When 10 widely used models were asked to summarize scientific texts, most produced conclusions broader than the original, even when told to be accurate. Some did this in 26% to 73% of cases. Model-written summaries were nearly five times more likely than human-written ones to contain a broad generalization (odds ratio 4.85, 95% CI 3.06 to 7.70) (Peters & Chin-Yee, 2025). “In this subgroup” becomes “in patients.” “Was associated with” becomes “improves.” “Not significant for the strong models” becomes “across all models.” The summary reads better than the paper. That is exactly the problem.
The run. On 2026-09-07 the full text of Vollmer et al., Diagnostics 2026 (Vollmer et al., 2026) was given to a summarizing tool that answers only from the page it is given (the model behind the WebFetch tool in Claude Code, which the vendor does not name). The question was the journal-club question from the opening scenario: did adding retrieval improve accuracy, and did it work for all the models? The three-sentence answer is reproduced word for word at the top of this chapter. A second run on a different paper in the same session (Duran et al., 2026) produced a score-2 answer. That one correctly carried the authors’ own caution that the tool had not been tested at other sites. The misreading is not inevitable. It is common.
The check. Click the citation on the first sentence. It takes you to Table 6. Two of the five confidence intervals cross zero, which means the gain for those two models may be nothing. That is the whole check, and it took less time than reading this paragraph. The general rule: for any sentence you are going to repeat, click its citation and read one line above and one line below the highlighted passage. The line below is where the “however” usually is.
Before the cohort. Re-run the same question in NotebookLM with the PMC page as the source. Record the model and date. Replace this specimen if the live run gives a cleaner score 1. The point is not this tool. The point is that you can produce one yourself.
13.3.6 Three more things grounding cannot see
The pipeline answers from what you gave it. That sets three limits, and no citation link will show you any of them.
The source set. A confident summary of five weak papers is a confident summary of five weak papers, with perfect citations. Grounding moves the appraisal problem from the answer to the sources. If you want to trust the notebook, you have to trust what you put in it. That is the appraisal chapter’s job.
What is missing. A notebook cannot cite a paper it does not contain. It will not tell you that the trial that would change your mind was published last month and is not in your five PDFs. That is the difference between a grounded search and a systematic one, a search that tries to find every relevant paper. The tool is excellent for reading a defined set of documents. It is the wrong tool for what does the literature say, because there the papers you failed to include are the ones that matter.
Disagreement. When two sources disagree, retrieval pulls passages from both, and generation writes one paragraph. A frequent result is one blended answer that no source holds, with the disagreement smoothed away. Nobody has measured how often; you will test whether your tool does it in the Do block, with a pair of papers chosen because they disagree about the same tool.
A grounded tool fixes the reference. It helps you see where each sentence came from. It does nothing for the claim itself. So: for any sentence you will repeat, click its citation, read a line above and below, and ask whether the sentence is narrower or broader than the passage. Broader is the usual direction. The tool is a way of finding your way through sources. It is not a source. The citation on your slide is the paper, and you have now read that part of it.
13.4 Watch or listen
Video. IBM Technology, “What is Retrieval-Augmented Generation (RAG)?”, presented by Marina Danilevsky (23 August 2023; 6 min 36 s). https://www.youtube.com/watch?v=T-D1OfcDW1M
Six and a half minutes, nothing you need to know first, and the clearest whiteboard version of Figure 13.1 that exists. Listen for the story about which planet has the most moons. The speaker’s confident wrong answer is what happens when the model relies on what it learned in training, which was out of date. Her fix, going to look it up, is retrieval. Then notice what the video does not say. It presents retrieval as the answer to being wrong and out of date. It never mentions that the model can retrieve the right passage and misread it. That omission is what this chapter covers. Every vendor’s sales presentation leaves it out too.
Reading alternative. The written article this chapter is built on, Grounding AI in real sources, builds up the context window from the World Series example and shows the conversation as raw text. About fifteen minutes.
13.5 Do
Sources for this exercise are published papers, guidelines, or things you wrote yourself. Not a de-identified case you typed up. Not a screenshot of a note. Not your classmate’s data. A consumer Google account carries no business associate agreement, the contract that makes a vendor legally responsible for protecting patient data. The course rule holds in every account, institutional or personal. The rule is strict on purpose. The unusual cases are where people get it wrong.
What counts as patient information, and why taking out the name is not enough: the patient information page.
NotebookLM is free with any Google account and runs in the browser. A personal Gmail account works and asks for nothing beyond what you already gave Google. Making a new Google account may ask for a phone number; if you will not give one, use the no-Google path below. If you sign in with your school Google account and NotebookLM is missing or refuses to load, your school’s administrator has turned it off for your part of the organization. They can do that. Do not spend time on it; switch to a personal account. Google’s own help page confirms that most work and school accounts have access and that administrators control it VERIFY: sign in with a school Google account the week before the cohort and confirm NotebookLM loads; this cannot be checked remotely.
If you have no Google account and will not make one: your institutional ChatGPT Edu or Copilot lets you attach files and say “answer only from the attached documents.” That gives you the same exercise with a weaker guarantee. A general chatbot told to stay inside its sources will still use its training data at times. Note in your write-up which tool you used.
13.5.1 Step 1: choose your sources (10 minutes)
Pick three to five papers in your own specialty that you were already going to read. If you do not have them, find them:
("large language model" OR "retrieval-augmented generation")
AND <your specialty> AND free full text[filter]
Filter PubMed to the last three years, read three abstracts, and take the three to five papers that share a question. “Free full text” matters because NotebookLM needs the paper, not the abstract. For each one, open the PMC version and paste its URL as a website source. That is faster than downloading and uploading the PDF, and it works the same way.
If you have no specialty yet, or want a set built to fail in interesting ways, use this ready-made set. Every paper is open access on PMC. Every one was checked in PubMed on 2026-09-07. And they are, on purpose, about the tool you are using:
| Paper | What it is | PMC |
|---|---|---|
| Vollmer et al., Diagnostics 2026 (Vollmer et al., 2026) | Six models, 50 guideline questions, with and without retrieval | PMC13465324 |
| Duran et al., Digital Health 2026 (Duran et al., 2026) | NotebookLM for orthopedic disability rating: 91% vs 68% accuracy | PMC13420063 |
| Seki et al., Clinics and Practice 2026 (Seki et al., 2026) | NotebookLM classifying radiographs: agreement with dentists near zero | PMC13511330 |
| Basu and Huynh, BMC Health Serv Res 2026 (Basu & Huynh, 2026) | Systematic review of 44 hallucination-mitigation studies | PMC13470931 |
Make a new notebook, add the sources, and wait for each to finish processing. That wait is steps 1 to 3 of Figure 13.1 happening.
13.5.2 Step 2: ask three questions and click every citation (10 minutes)
Ask three questions you would want answered for journal club. One of them should be a question you already know the answer to, so you can judge the answer rather than accept it. For the ready-made set, use these:
- What did each paper measure, and what was its main result?
- Did retrieval work for all the models in the guideline study?
- Is NotebookLM accurate?
For every citation in every answer, not a sample, click it. Read the highlighted passage plus one line above and one line below. Score each sentence-and-citation pair on the appraisal chapter’s scale (Table 5.4). 2 if the passage supports the sentence as written. 1 if the passage is real but the sentence is broader, narrower, or different. 0 if the cited passage does not exist, or the citation points at something other than what the sentence names. You are looking for a 1. There is usually at least one.
13.5.3 Step 3: ask what the sources disagree about (5 minutes)
Ask a question your sources answer differently. For the ready-made set, question 3 above is it. One paper found NotebookLM raised accuracy from 68% to 91% on a rating task done from a single written regulation. Another found its agreement with dentists on radiographs was slight or worse. Watch what the tool does with that. Does it present both, with the task each one measured? Or does it produce a sentence like “NotebookLM shows promise but has limitations,” which is true of everything and cites both?
13.5.4 Step 4: try to make it fabricate, then ask the chatbot (5 minutes)
Ask the notebook something its sources cannot support: a statistic that is not in them, a named trial that is not there, a population none of the papers studied. For the ready-made set: What is the fabrication rate of NotebookLM on citations in cardiology? No source contains that. A grounded tool should say so, give a cautious partial answer, or give a short, hedged answer.
Now ask the identical question in your general chatbot tab. Do not prompt it carefully. Ask it the way you asked the notebook. Compare on one thing only: can you check where each sentence came from? The chatbot’s answer may be better written and more useful. That is not the thing you are comparing.
If your chatbot searched the web and cited pages, notice that this is a different question about where the answer came from. The page is real. Whether it is evidence is the appraisal chapter’s job.
13.5.5 Record it
Tool: NotebookLM / other: ____________ Date: __________
Sources (3–5, with PMC IDs or titles):
1 ______________________________________
2 ______________________________________
3 ______________________________________
4 ______________________________________
5 ______________________________________
Q1 ____________________________________ citations clicked ___ scores: 2s __ 1s __ 0s __
Q2 ____________________________________ citations clicked ___ scores: 2s __ 1s __ 0s __
Q3 ____________________________________ citations clicked ___ scores: 2s __ 1s __ 0s __
For the score-1 you found (or the closest to one): in one sentence, what did
the tool say, and what does the passage say?
________________________________________________________________________
Disagreement question: did the tool present both positions, or smooth them?
________________________________________________________________________
Fabrication attempt. What you asked: __________________________________
Notebook said: ________________________________________________________
Chatbot said: ________________________________________________________
Which one could you trace, sentence by sentence? _____________________
The score-1 sentence is the thing to keep from this exercise. “Three of twelve citations were stretched” is a number that will change next month. Being able to say how one was stretched is the skill.
13.6 Check
Score yourself against this before moving on.
| You have | Meets | Falls short |
|---|---|---|
| The no-memory explanation | You can say what is re-sent each turn and why a new chat starts from zero, without the word “memory” | “It remembers the conversation” |
| The no-learning statement | You can say what changes between turns (the context) and what does not (the model), and separate that from where your text is stored | “It learns from what you type” |
| The pipeline | Six steps in order, and one sentence on what chunking does to a table | Fewer than six, or “it searches the documents” |
| Why the citations are real | You can say why a citation from a grounded tool points at a passage, and why that is not the same as the model being careful | “It’s more accurate” |
| A scored set of citations | Every citation clicked; one score-1 sentence quoted beside its passage | Scores assigned by reading the answer alone |
| The disagreement result | You can say whether the tool presented both sources or smoothed them, and quote the smoothing sentence if it did | “It handled it fine” |
| The fabrication contrast | Both answers recorded; you can say which one you could trace and why the other reads better anyway | Only the notebook’s answer |
| The search limit | You can say in one sentence when a grounded notebook is not a literature search | “It’s good for research” |
13.7 What this changes for you
What does this change about how you’ll practice? Answer in writing before you close this chapter. Two specifics. First: the next time a summary with citations is put in front of you, from a notebook, an EHR tool, or a co-intern, what is the one-minute check you will do before repeating a sentence from it? Which of the failures in this chapter does that check catch? Second: go back to Monday. The five PDFs are in the notebook, the first sentence of its answer is on your draft slide, and journal club is Thursday. What do you change on the slide? What do you say when the chief asks how you prepared? What do you do with the two papers that disagree? One word is not an answer to either.
13.8 Summary
- A chatbot has no memory. The conversation is re-sent every turn inside a fixed context window, and a new chat starts from zero.
- A model does not learn from you. The context changes; the model does not. Where your text is stored is a separate question, and the reason for the no-patient-information rule.
- Grounding is six steps: chunk, embed, index, retrieve, augment, generate. The model is given passages and asked to paraphrase them.
- Chunking cuts on length, not meaning. A table split across two chunks is two real citations and no table.
- Grounded tools fabricate far less because a citation is a pointer to a passage in the window. That comes from how the tool is built, not from care, which is why it holds for tools that do not exist yet.
- Grounding does not fix reading. A real, linked source can be misread, and the usual direction is broader than the paper. Click, read a line above and below, ask narrower or broader.
- A notebook answers from what you gave it. It is a way to read a defined set of documents, not a way to find out what the literature says.
- Thursday’s slide: the number stays, the sentence changes, and the two papers that disagree get their own line each.
13.9 Go deeper
Papers
- The overgeneralization study, 10 models and 4,900 summaries (Peters & Chin-Yee, 2025). Its finding that newer models did worse is the “these numbers move” point applied to reading rather than references.
- The guideline-retrieval benchmark behind Figure 13.2, whose pipeline was released openly (Vollmer et al., 2026), and the systematic review of ways to reduce hallucination that compares retrieval with the other methods (Basu & Huynh, 2026).
- Two studies of NotebookLM itself: one where it helped (Duran et al., 2026) and one where it was the wrong tool for an image task (Seki et al., 2026). Read them as a pair.
- The dentistry comparison whose phrase, “avoided fabricated references but still generated incorrect interpretations,” is this chapter in nine words (Hooshiar, 2025).
- “Lost in the middle,” the experiment on where in the context a document sits (N. F. Liu et al., 2024). Also in the foundations chapter.
- A two-day workshop for clinicians that limits citation generation to NotebookLM and then makes participants check every one. It is honest that whether this reduces fabrication was not tested (S.-W. Liu et al., 2026).
Talks
- Grounding AI in real sources: the written article this chapter is built on, with the conversation shown as raw text and the original shared-notebook exercise.
- Word embeddings as the basis for LLMs: what “similar meaning, nearby numbers” actually means, with an interactive embedding projector, for anyone who wants to understand step 2 of the pipeline in more detail. Also in the foundations chapter.
- Prompt engineering: ten rules with clinical examples, for students whose Step 2 questions came back vague. the prompting lab is built on it.