AI chatbots vs experts: can they retrieve clinical studies?

When an AI chatbot cites clinical studies, do those match what experts would include? Research benchmarks Claude Sonnet 5, Gemini 3.1 Pro, and GPT-5.5 against Cochrane reviews to measure retrieval quality.
The finding Chatbots retrieve only a fraction of Cochrane-included primary studies, with recall varying by model and user role.
The method Researchers prompted LLMs with Cochrane intervention questions and benchmarked cited trials against the Cochrane included/excluded sets.
The caveat Retrieval performance is biased toward clinical trials with larger sample sizes, even when citations are not fabricated.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

ChatGPT retrieved a higher share of Cochrane-included studies than Claude Sonnet 5 and Gemini 3.1 Pro, but overall recall was only partial. On average, one response retrieved 39.2% of included studies and cited 5.0% of excluded studies.

In practice, model choice and who the chatbot is prompted to be (patient vs clinician vs evidence-synthesis researcher) materially change which primary studies it pulls in, so the evidence picture can shift depending on configuration.

A key limitation is selective retrieval bias: sample size was the only independently significant predictor, suggesting chatbots may favor larger-sample trials rather than the most targeted studies experts would include.

AI chatbots vs experts: can they retrieve clinical studies?

Introduction

If you’ve ever asked an AI chatbot a medical question and then wondered, “Okay, but where did that answer come from?”, you’re not alone. A big reason people trust modern chatbots is that many now try to cite clinical studies alongside their explanations. The big question is whether those citations line up with what experts would actually pick from the literature.

This new research (from the original paper) dives into exactly that: how well three recent general-purpose AI chatbots retrieve the same primary studies that appear as included evidence in real-world systematic reviews—specifically, Cochrane intervention reviews. Instead of focusing only on citation “hallucinations” (fake or incorrect references), the authors evaluate something more subtle: the quality of retrieval—did the model retrieve studies that systematic review authors truly considered relevant?

Why This Matters

This matters right now because chatbots are increasingly being used as “information first” tools by patients, clinicians, and researchers. And when the chatbot’s citations are off—even when they’re not outright fabricated—you can still end up with a distorted evidence picture. The problem isn’t just misinformation; it’s selective omission. Systematic reviews exist largely to reduce that selective thinking by experts who search, screen, and compare studies.

A practical scenario: imagine a clinician quickly checking whether a weight-loss medication has trial evidence in a patient population similar to theirs. If the chatbot cites trials that systematic review authors would have excluded, or misses key trials, it can subtly skew clinical judgment. The clinician might not notice the mismatch if they trust the citation list at face value—especially if the chatbot provides a confident narrative.

Compared to earlier AI evaluation work that often reported low retrieval recall, this study helps explain why the recall differs. Prior work frequently emphasized fabricated citations or cited irrelevant studies. Here, the takeaway is more nuanced: performance varies by model choice, user role framing, and—most importantly—sample size of the studies. That last part is a big deal: it suggests chatbots may systematically gravitate toward “bigger trial” evidence, not necessarily the most targeted or nuanced studies.

How the researchers tested study retrieval quality

The authors used the Cochrane Database of Systematic Reviews—specifically, intervention reviews from Issues 6 and 7 of 2026. They selected 20 review questions spanning many clinical areas (drug/biologic interventions, surgical or procedural techniques, bedside care/management, nutrition supplements, rehab/physical treatments, behavioral interventions, and telehealth/service delivery). They excluded protocols and non-intervention topics like diagnosis and prognosis.

For each review question, the researchers asked the chatbots to answer with primary clinical citations (original studies like trials), not secondary sources (like guidelines or narrative reviews). Then they compared the chatbot-cited studies to the corresponding Cochrane review’s sets of included and excluded studies.

The chatbot lineup: three modern general-purpose models

The experiment focused on three consumer-facing LLM chatbots (all with reasoning-enabled configurations and web search capabilities). The study evaluated:

Model (reasoning configuration) What’s notable in the setup
Claude Sonnet 5 (medium effort; thinking enabled) Knowledge cutoff January 2026; reasoning enabled
Gemini 3.1 Pro (extended thinking) Knowledge cutoff Jan 31, 2025; reasoning enabled
GPT-5.5 (high reasoning) Knowledge cutoff Dec 1, 2025; reasoning enabled

Each response was generated independently in a fresh web session (no chat history), collected during mid-to-late July 2026.

User role framing: patient, clinician, or evidence-synthesis researcher

The authors didn’t just ask “one generic prompt.” They rewrote the same core question so the model believed the user was:

  • Patient
  • Clinician
  • Evidence-synthesis researcher

The wording was kept consistent except for role-specific framing at the beginning of each prompt. This is important because it tests whether the chatbot’s evidence selection shifts when it thinks it’s talking to different audiences.

Replication and scale: 720 total responses

To account for variability, each chatbot–role–question prompt was run four independent times. With 20 Cochrane review questions and three roles and three chatbots, the total number of responses was reported as 720.

That scale matters because it reduces the chance that results were driven by one lucky (or unlucky) run. The study also makes clear what it’s benchmarking against: Cochrane’s “included studies” and “excluded studies” lists, question-by-question.

What the chatbots did with included vs excluded Cochrane studies

The most direct performance metric was recall of included studies—how much of the Cochrane included-study set the chatbot managed to retrieve when it produced its answer and citations.

On average, across all 720 responses:

  • A chatbot response retrieved 39.2% ± 29.8% of Cochrane included studies
  • It cited only 5.0% ± 9.4% of Cochrane excluded studies

That precision-ish behavior (low excluded citations) might sound reassuring at first glance. But the recall number is the real warning light: even at its best, the model doesn’t “find what experts would” in any consistent or complete way. It retrieves some included evidence, but it misses a lot.

Recall varied massively—by model and by user role

The study reports that recall differed significantly by model and role.

Model comparison: GPT-5.5 led the pack

Average recall of Cochrane included studies was:

  • GPT-5.5: 63.1% ± 29.5%
  • Claude Sonnet 5: 37.0% ± 23.8%
  • Gemini 3.1 Pro: 17.3% ± 13.1%

The reported statistical result for the model effect is p = 2.0×10−5, with GPT-5.5 clearly ahead.

User role framing: evidence-synthesis researchers prompted higher recall

Average recall by user role:

  • Evidence-synthesis researcher: 42.8% ± 30.8%
  • Clinician: 38.6% ± 28.9%
  • Patient: 36.1% ± 29.3%

This role effect was also statistically significant (p = 2.0×10−5).

So, role framing didn’t radically change the direction (all roles were imperfect), but it did shift recall in a measurable way.

The bias toward “bigger” trials

The authors then asked: what actually predicts whether the chatbot retrieves the included studies?

They controlled for factors like:

  • publication year
  • citations per year
  • open-access status

After controlling for those, they found sample size was the only independently significant predictor of retrieval, reported as:

  • odds ratio 1.80 per 1-unit increase in log sample size
  • 95% CI 1.37–2.36
  • p = 2.34×10−5

In plain language: when studies were larger (more participants), they were much more likely to be retrieved and cited by the chatbots.

Why “retrieval recall” is trickier than it sounds

It’s tempting to treat citation lists as a binary “right/wrong.” But retrieval recall is more like hunting for specific fingerprints at a crime scene. The chatbot may include some “obvious matches,” but fail to pull in the full set of evidence that experts concluded was relevant.

Here are a few ways that can happen—even without fabrication:

  1. The chatbot may surface more prominent studies
    Larger trials tend to be more widely discussed and easier to retrieve from indexed information. This aligns with the study’s sample-size bias.

  2. Systematic reviews use structured screening
    Cochrane review authors follow specific inclusion/exclusion criteria. A chatbot answering in natural language might miss that nuance and instead gravitate toward studies that “fit” superficially.

  3. The model may not reliably translate evidence-selection into citations
    Even if the model “knows” the topic, generating a correct set of primary citations is a different sub-task than producing an answer.

This is why the difference between citing excluded studies (5.0%) and missing included studies (39.2%) matters. A low fabrication rate doesn’t automatically mean the evidence set is complete or faithful.

The role of user framing: who the chatbot thinks you are

The role effect is subtle but meaningful: the same core question, framed as coming from an evidence-synthesis researcher, produced higher included-study recall than when framed as coming from a patient.

One way to interpret this (without over-claiming) is that role framing changes how the model “plans” its response. An evidence-synthesis researcher persona likely encourages:

  • a more method-like tone
  • more emphasis on trial-level evidence
  • more structured retrieval behavior

Meanwhile, a patient framing may implicitly lead the chatbot toward simpler explanations and fewer detailed evidence hops. The clinician role sits in between.

Still, the results also show that none of the roles produce expert-level retrieval. Even the best role average still misses a large portion of Cochrane included studies.

If you want to connect back to earlier work: this result builds on the broader idea from previous AI research that prompt framing can shift outcomes, but here it specifically affects the study selection behavior, not just phrasing or confidence.

The sample-size effect: why bigger trials keep getting pulled in

This is the most actionable finding in the whole paper: once you control for year, citation rate, and open-access status, sample size is the only independent predictor.

That suggests a mechanism: chatbots may prefer studies that are more statistically “satisfying” and more likely to be widely referenced, which often correlates with larger trials. In other words, evidence retrieval isn’t neutral—it’s biased toward a certain kind of study.

Practical implications for real users

  • Patients: If chatbots disproportionately cite larger trials, they may under-expose smaller-but-relevant evidence for specific subgroups or rare outcomes.
  • Clinicians: Evidence for effect heterogeneity (who benefits, who doesn’t) often comes from smaller trials. A retrieval bias could hide those nuances.
  • Researchers doing rapid scoping: If the chatbot is used as an initial evidence radar, it might miss key smaller studies that later matter for systematic review conclusions.

This is also a caution for tool-building: if systems are evaluated only on whether they produce “a citation,” they may score well while still failing to retrieve a balanced study set.

And again, you can find the overall experimental details and code/data availability at the repository linked in the paper: https://github.com/QingfangLiu/llm-evidence-retrieval-bias

What we should do differently with AI medical citations

The study’s core message is uncomfortable but useful: even with web search and reasoning-enabled configurations, chatbots don’t automatically replicate expert evidence selection. So how should you respond if you’re building workflows (or using chatbots)?

1) Treat citations from chatbots as a starting point, not the evidence set

The recall numbers (like the average 39.2% of included studies) mean that an AI citation list may look convincing while still being incomplete.

2) Expect systematic bias—not random errors

The strong relationship with sample size implies you may routinely get a “larger-trial skew.” That can distort uncertainty too, because smaller trials often contribute to different parts of the evidence landscape.

3) Evaluate models on retrieval quality, not just citation plausibility

A lot of earlier concerns focused on fabricated citations. This research shows another axis: retrieval coverage relative to expert-included studies.

If your goal is clinical or research decision support, you’ll want metrics that reflect what matters: “Did it retrieve the set that experts would have included?”

Key Takeaways

  • Chatbots retrieve only part of the expert evidence set. On average, responses retrieved 39.2% of Cochrane included studies, while citing only 5.0% of Cochrane excluded studies.
  • Model choice matters a lot. Mean included-study recall was 63.1% for GPT-5.5, 37.0% for Claude Sonnet 5, and 17.3% for Gemini 3.1 Pro (statistically significant).
  • User role framing changes retrieval. Evidence-synthesis researcher prompts produced higher recall (42.8%) than clinician (38.6%) or patient (36.1%).
  • Sample size is the key predictor (after controls). Bigger trials were much more likely to be retrieved (odds ratio 1.80 per log sample size unit), suggesting evidence retrieval bias.
  • Practical guidance: use chatbot citations as a leads-and-triage step, not as a substitute for systematic searching or expert screening—especially if smaller trials or subgroup evidence could be important.
  • For the future: evaluations should include retrieval coverage against systematic review sets, not only “no fake citations” checks. That’s how we’ll measure whether AI can truly support evidence-based decisions.

If you want, tell me your audience (patients vs clinicians vs researchers) and I can rewrite this into a version tailored to how people would actually use chatbots in that setting.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Generative AI Isn’t Ready to Replace Stats Experts—But It Can Help

In-Context Privacy Learning for Chatbots (Just-in-Time Tools)

Are AI Chatbots More Empathetic Than Doctors? A Practical Look at the New Empathy Research

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.