The Short Answer
ChatGPT retrieved a higher share of Cochrane-included studies than Claude Sonnet 5 and Gemini 3.1 Pro, but overall recall was only partial. On average, one response retrieved 39.2% of included studies and cited 5.0% of excluded studies.
In practice, model choice and who the chatbot is prompted to be (patient vs clinician vs evidence-synthesis researcher) materially change which primary studies it pulls in, so the evidence picture can shift depending on configuration.
A key limitation is selective retrieval bias: sample size was the only independently significant predictor, suggesting chatbots may favor larger-sample trials rather than the most targeted studies experts would include.
On this page
- Introduction
- Why This Matters
- How the researchers tested study retrieval quality
- What the chatbots did with included vs excluded Cochrane studies
- Why “retrieval recall” is trickier than it sounds
- The role of user framing: who the chatbot thinks you are
- The sample-size effect: why bigger trials keep getting pulled in
- What we should do differently with AI medical citations
- Key Takeaways
AI chatbots vs experts: can they retrieve clinical studies?
Introduction
If you’ve ever asked an AI chatbot a medical question and then wondered, “Okay, but where did that answer come from?”, you’re not alone. A big reason people trust modern chatbots is that many now try to cite clinical studies alongside their explanations. The big question is whether those citations line up with what experts would actually pick from the literature.
This new research (from the original paper) dives into exactly that: how well three recent general-purpose AI chatbots retrieve the same primary studies that appear as included evidence in real-world systematic reviews—specifically, Cochrane intervention reviews. Instead of focusing only on citation “hallucinations” (fake or incorrect references), the authors evaluate something more subtle: the quality of retrieval—did the model retrieve studies that systematic review authors truly considered relevant?
Why This Matters
This matters right now because chatbots are increasingly being used as “information first” tools by patients, clinicians, and researchers. And when the chatbot’s citations are off—even when they’re not outright fabricated—you can still end up with a distorted evidence picture. The problem isn’t just misinformation; it’s selective omission. Systematic reviews exist largely to reduce that selective thinking by experts who search, screen, and compare studies.
A practical scenario: imagine a clinician quickly checking whether a weight-loss medication has trial evidence in a patient population similar to theirs. If the chatbot cites trials that systematic review authors would have excluded, or misses key trials, it can subtly skew clinical judgment. The clinician might not notice the mismatch if they trust the citation list at face value—especially if the chatbot provides a confident narrative.
Compared to earlier AI evaluation work that often reported low retrieval recall, this study helps explain why the recall differs. Prior work frequently emphasized fabricated citations or cited irrelevant studies. Here, the takeaway is more nuanced: performance varies by model choice, user role framing, and—most importantly—sample size of the studies. That last part is a big deal: it suggests chatbots may systematically gravitate toward “bigger trial” evidence, not necessarily the most targeted or nuanced studies.
How the researchers tested study retrieval quality
The authors used the Cochrane Database of Systematic Reviews—specifically, intervention reviews from Issues 6 and 7 of 2026. They selected 20 review questions spanning many clinical areas (drug/biologic interventions, surgical or procedural techniques, bedside care/management, nutrition supplements, rehab/physical treatments, behavioral interventions, and telehealth/service delivery). They excluded protocols and non-intervention topics like diagnosis and prognosis.
For each review question, the researchers asked the chatbots to answer with primary clinical citations (original studies like trials), not secondary sources (like guidelines or narrative reviews). Then they compared the chatbot-cited studies to the corresponding Cochrane review’s sets of included and excluded studies.
The chatbot lineup: three modern general-purpose models
The experiment focused on three consumer-facing LLM chatbots (all with reasoning-enabled configurations and web search capabilities). The study evaluated:
| Model (reasoning configuration) | What’s notable in the setup |
|---|---|
Claude Sonnet 5 (medium effort; thinking enabled) |
Knowledge cutoff January 2026; reasoning enabled |
Gemini 3.1 Pro (extended thinking) |
Knowledge cutoff Jan 31, 2025; reasoning enabled |
GPT-5.5 (high reasoning) |
Knowledge cutoff Dec 1, 2025; reasoning enabled |
Each response was generated independently in a fresh web session (no chat history), collected during mid-to-late July 2026.
User role framing: patient, clinician, or evidence-synthesis researcher
The authors didn’t just ask “one generic prompt.” They rewrote the same core question so the model believed the user was:
- Patient
- Clinician
- Evidence-synthesis researcher
The wording was kept consistent except for role-specific framing at the beginning of each prompt. This is important because it tests whether the chatbot’s evidence selection shifts when it thinks it’s talking to different audiences.
Replication and scale: 720 total responses
To account for variability, each chatbot–role–question prompt was run four independent times. With 20 Cochrane review questions and three roles and three chatbots, the total number of responses was reported as 720.
That scale matters because it reduces the chance that results were driven by one lucky (or unlucky) run. The study also makes clear what it’s benchmarking against: Cochrane’s “included studies” and “excluded studies” lists, question-by-question.
What the chatbots did with included vs excluded Cochrane studies
The most direct performance metric was recall of included studies—how much of the Cochrane included-study set the chatbot managed to retrieve when it produced its answer and citations.
On average, across all 720 responses:
- A chatbot response retrieved 39.2% ± 29.8% of Cochrane included studies
- It cited only 5.0% ± 9.4% of Cochrane excluded studies
That precision-ish behavior (low excluded citations) might sound reassuring at first glance. But the recall number is the real warning light: even at its best, the model doesn’t “find what experts would” in any consistent or complete way. It retrieves some included evidence, but it misses a lot.
Recall varied massively—by model and by user role
The study reports that recall differed significantly by model and role.
Model comparison: GPT-5.5 led the pack
Average recall of Cochrane included studies was:
GPT-5.5: 63.1% ± 29.5%Claude Sonnet 5: 37.0% ± 23.8%Gemini 3.1 Pro: 17.3% ± 13.1%
The reported statistical result for the model effect is p = 2.0×10−5, with GPT-5.5 clearly ahead.
User role framing: evidence-synthesis researchers prompted higher recall
Average recall by user role:
- Evidence-synthesis researcher: 42.8% ± 30.8%
- Clinician: 38.6% ± 28.9%
- Patient: 36.1% ± 29.3%
This role effect was also statistically significant (p = 2.0×10−5).
So, role framing didn’t radically change the direction (all roles were imperfect), but it did shift recall in a measurable way.
The bias toward “bigger” trials
The authors then asked: what actually predicts whether the chatbot retrieves the included studies?
They controlled for factors like:
- publication year
- citations per year
- open-access status
After controlling for those, they found sample size was the only independently significant predictor of retrieval, reported as:
- odds ratio 1.80 per 1-unit increase in log sample size
- 95% CI 1.37–2.36
- p = 2.34×10−5
In plain language: when studies were larger (more participants), they were much more likely to be retrieved and cited by the chatbots.
Why “retrieval recall” is trickier than it sounds
It’s tempting to treat citation lists as a binary “right/wrong.” But retrieval recall is more like hunting for specific fingerprints at a crime scene. The chatbot may include some “obvious matches,” but fail to pull in the full set of evidence that experts concluded was relevant.
Here are a few ways that can happen—even without fabrication:
The chatbot may surface more prominent studies
Larger trials tend to be more widely discussed and easier to retrieve from indexed information. This aligns with the study’s sample-size bias.Systematic reviews use structured screening
Cochrane review authors follow specific inclusion/exclusion criteria. A chatbot answering in natural language might miss that nuance and instead gravitate toward studies that “fit” superficially.The model may not reliably translate evidence-selection into citations
Even if the model “knows” the topic, generating a correct set of primary citations is a different sub-task than producing an answer.
This is why the difference between citing excluded studies (5.0%) and missing included studies (39.2%) matters. A low fabrication rate doesn’t automatically mean the evidence set is complete or faithful.
The role of user framing: who the chatbot thinks you are
The role effect is subtle but meaningful: the same core question, framed as coming from an evidence-synthesis researcher, produced higher included-study recall than when framed as coming from a patient.
One way to interpret this (without over-claiming) is that role framing changes how the model “plans” its response. An evidence-synthesis researcher persona likely encourages:
- a more method-like tone
- more emphasis on trial-level evidence
- more structured retrieval behavior
Meanwhile, a patient framing may implicitly lead the chatbot toward simpler explanations and fewer detailed evidence hops. The clinician role sits in between.
Still, the results also show that none of the roles produce expert-level retrieval. Even the best role average still misses a large portion of Cochrane included studies.
If you want to connect back to earlier work: this result builds on the broader idea from previous AI research that prompt framing can shift outcomes, but here it specifically affects the study selection behavior, not just phrasing or confidence.
The sample-size effect: why bigger trials keep getting pulled in
This is the most actionable finding in the whole paper: once you control for year, citation rate, and open-access status, sample size is the only independent predictor.
That suggests a mechanism: chatbots may prefer studies that are more statistically “satisfying” and more likely to be widely referenced, which often correlates with larger trials. In other words, evidence retrieval isn’t neutral—it’s biased toward a certain kind of study.
Practical implications for real users
- Patients: If chatbots disproportionately cite larger trials, they may under-expose smaller-but-relevant evidence for specific subgroups or rare outcomes.
- Clinicians: Evidence for effect heterogeneity (who benefits, who doesn’t) often comes from smaller trials. A retrieval bias could hide those nuances.
- Researchers doing rapid scoping: If the chatbot is used as an initial evidence radar, it might miss key smaller studies that later matter for systematic review conclusions.
This is also a caution for tool-building: if systems are evaluated only on whether they produce “a citation,” they may score well while still failing to retrieve a balanced study set.
And again, you can find the overall experimental details and code/data availability at the repository linked in the paper: https://github.com/QingfangLiu/llm-evidence-retrieval-bias
What we should do differently with AI medical citations
The study’s core message is uncomfortable but useful: even with web search and reasoning-enabled configurations, chatbots don’t automatically replicate expert evidence selection. So how should you respond if you’re building workflows (or using chatbots)?
1) Treat citations from chatbots as a starting point, not the evidence set
The recall numbers (like the average 39.2% of included studies) mean that an AI citation list may look convincing while still being incomplete.
2) Expect systematic bias—not random errors
The strong relationship with sample size implies you may routinely get a “larger-trial skew.” That can distort uncertainty too, because smaller trials often contribute to different parts of the evidence landscape.
3) Evaluate models on retrieval quality, not just citation plausibility
A lot of earlier concerns focused on fabricated citations. This research shows another axis: retrieval coverage relative to expert-included studies.
If your goal is clinical or research decision support, you’ll want metrics that reflect what matters: “Did it retrieve the set that experts would have included?”
Key Takeaways
- Chatbots retrieve only part of the expert evidence set. On average, responses retrieved 39.2% of Cochrane included studies, while citing only 5.0% of Cochrane excluded studies.
- Model choice matters a lot. Mean included-study recall was 63.1% for
GPT-5.5, 37.0% forClaude Sonnet 5, and 17.3% forGemini 3.1 Pro(statistically significant). - User role framing changes retrieval. Evidence-synthesis researcher prompts produced higher recall (42.8%) than clinician (38.6%) or patient (36.1%).
- Sample size is the key predictor (after controls). Bigger trials were much more likely to be retrieved (odds ratio 1.80 per log sample size unit), suggesting evidence retrieval bias.
- Practical guidance: use chatbot citations as a leads-and-triage step, not as a substitute for systematic searching or expert screening—especially if smaller trials or subgroup evidence could be important.
- For the future: evaluations should include retrieval coverage against systematic review sets, not only “no fake citations” checks. That’s how we’ll measure whether AI can truly support evidence-based decisions.
If you want, tell me your audience (patients vs clinicians vs researchers) and I can rewrite this into a version tailored to how people would actually use chatbots in that setting.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions — arXiv
- Authors: Authors: Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu