The Short Answer
Mid-2025 LLMs show low overlap with expert-selected physics references (well under 6%) and often return incorrect citation metadata, so they are not a standalone literature-review replacement.
Use AI as a first-draft search assistant: generate candidate references, then have humans verify every entry—especially titles, authors, years, journals, DOIs, and links.
The main caveat is verification burden: fabricated references and metadata mismatches are common enough that unreviewed AI citation lists can mislead early research.
On this page
- Why This Matters: The “Early Search” Problem in Physics Is Getting Worse
- How the Study Was Set Up: 8 Expert Projects, Parallel Human vs AI Searches
- What the Researchers Measured: Overlap and Relevance Were Low for Mid-2025 Models
- The Big Gotcha: Mid-2025 AI References Had Lots of Metadata Errors
- What Improved in 2026: ChatGPT Pro 5.5 Looked Perfect in a One-Project Spot Test
- Practical Implications: How to Use LLMs for Literature Search Without Getting Burned
- Key Takeaways
AI for Physics Literature Reviews: What Works (and What Fails)
AI tools are starting to show up everywhere in research workflows—but one of the most dangerous places to “wing it” is the literature review. If you miss a key paper early, you might accidentally reinvent something already done, or quietly build on the wrong assumptions. New research from this arXiv paper takes a hard, controlled look at whether large language models (LLMs) can actually help physicists, astrophysicists, and cosmologists find the right references.
The study (based on 8 expert-conceived projects) compares human literature searches against AI-generated candidate reference lists using mid-2025 models like ChatGPT-4o, ChatGPT Deep Research, and Gemini. It then stress-tests the reliability of what the AI returns—especially for “hallucinations,” where models either fabricate nonexistent papers or mix up metadata for real ones. Spoiler: the AI is not yet a replacement for experts, but it may become a useful collaborator—if you verify what it gives you.
Why This Matters: The “Early Search” Problem in Physics Is Getting Worse
Physics moves fast, but literature gets faster. The number of papers published per day is rising dramatically, and most researchers don’t start with exhaustive reviews. Early on, you’re often doing a fast “does this idea already exist?” scan—especially when you’re new to a subfield or joining a new project. That’s exactly when missing citations is most costly, because you haven’t yet invested time in deep validation.
What’s interesting about this research is that it doesn’t just ask “can LLMs write plausible citations?” It measures something more operational: how much useful overlap you get between expert-picked references and AI-picked ones. The results show surprisingly low overlap for the mid-2025 models (the paper reports overlap well under 6%), which means the AI isn’t “quietly completing” the expert bibliography on its own. Instead, it tends to produce a different set—sometimes adding relevant coverage, sometimes drifting into irrelevant territory.
There’s also a real-world scenario where this matters immediately: suppose you’re a graduate student tasked with a rapid literature scan before you even have a full working hypothesis. You could use an LLM to generate a candidate list, but this paper suggests you must treat it like a first-draft search assistant, not like an authority. The practical workflow becomes: AI suggests, humans verify, and—crucially—you look for metadata errors and fabricated entries.
Finally, this work builds on earlier findings that LLMs can hallucinate references (especially older chat-style systems). But this study goes further by separating two failure modes—fabrications vs metadata mismatches—and by showing that stronger “tool-use” modes can drastically reduce hallucination risk. In other words: the direction of travel is toward safer AI-assisted searching, but the paper makes it clear that the full story is about how the model searches and checks, not only what it “knows.”
How the Study Was Set Up: 8 Expert Projects, Parallel Human vs AI Searches
To evaluate AI as a literature-review helper in physics-adjacent fields, the authors designed a controlled comparison across eight projects spanning physics, astrophysics, and cosmology. Each project had a clear background and goal, and the key part is that humans and AI were asked to do the same kind of task in parallel.
The eight projects the researchers used
Each project represents a frontier research area and had an expert who prepared context and did a human literature search. The AI prompters generated model outputs independently using standardized prompts grounded in that context. The projects included topics like:
- AGN – MaNGA: AGN Duty Cycle (identifying fading/variable active galactic nuclei in survey data)
- LBG – galaxy–dark matter halo connection of Lyman-break galaxies
- IA – intrinsic alignments in varying environments (weak lensing contaminant modeling)
- AR – debris emergence on laser-ablated sub-wavelength shapes (manufacturing physics)
- RG – radio galaxies at CMB frequencies with
HalfDome-based modeling - GW – environments of gravitational-wave black hole binaries using cross-correlation with lensing maps
- PTA – pulsar timing array forecasts for deviations from GR
- SU2 – Massive Yang–Mills theory (inflation scenarios with self-interacting massive vectors)
The task: “choose up to 50 papers per project”
For each project, the human expert selected up to 50 important references, grouped into four buckets:
- Review papers (past 15 years)
- Highly cited papers (e.g., foundational proposals/discoveries)
- Recent papers (past 10 years)
- Other papers (e.g., key survey/collaboration papers like Planck/WAMP/DES/WISE/etc.)
In parallel, the AI systems produced candidate references also capped at 50 per model per project. Each AI reference included fields like title, authors, year, journal, DOI, and link—and the AI was instructed to prioritize ADS, INSPIRE, and arXiv links when available.
The models compared (mid-2025) and one extra check (2026)
The study focused mainly on mid-2025 models:
ChatGPT-4oChatGPT Deep Research(a tool-augmented mode)Gemini
Then the authors ran a small “close to submission” spot test for one project using a newer model:
ChatGPT Pro 5.5(dated July 2nd, 2026)
If you want the headline takeaway: model choice and tool-use matter a lot.
What the Researchers Measured: Overlap and Relevance Were Low for Mid-2025 Models
The first key question was about completeness and overlap: did AI retrieve the same important references humans picked?
The authors computed the overlap in two ways:
1. Comparing human-picked vs AI-picked papers and classifying whether each paper was relevant.
2. Comparing AI candidate references that were not found by a human, then verifying whether they were real and whether their metadata was correct.
A reality check: overlap between humans and AI was small
Across all projects, the human experts selected 194 papers total. Across the three mid-2025 LLM runs, the AI produced 701 candidate references in total (split as ChatGPT-4o: 219, ChatGPT Deep Research: 157, Gemini: 325).
When comparing “papers found by humans vs found by AI,” the paper reports that overlap is small (<<6%) for the mid-2025 models. That’s the opposite of what you’d hope if you wanted AI to replicate expert search competence.
Overlap aside, AI sometimes finds relevant papers humans missed
Even though overlap was small, the researchers also found that AI produced a non-negligible fraction of relevant papers that humans did not select.
From the paper’s reported breakdown:
- For ChatGPT-4o and ChatGPT Deep Research, the fraction of papers found by AI but missed by humans was smaller than the reverse (humans found but AI missed).
- For Gemini, the asymmetry was smaller (the paper notes a difference of only 5% between those two fractions).
So, the AI isn’t useless—it’s just not “matching” expert search coverage. It behaves more like a parallel searcher whose results only partially intersect.
Irrelevance was model-dependent (Gemini drifted more)
The paper also breaks down the references AI generated that humans judged irrelevant. Based on the authors’ reported rates:
- Gemini had a 31% irrelevance rate in the “AI-only” set.
- ChatGPT-4o had 21%.
- ChatGPT Deep Research was much lower at 6%.
This aligns with the intuition that “better search with verification” gives fewer junk citations.
The comparison as a quick summary table (mid-2025 models)
| Model (mid-2025) | Total AI refs (all projects) | Irrelevant rate among AI-only refs | Notes from study framing |
|---|---|---|---|
ChatGPT-4o |
219 | 21% | Produces some relevant candidates, but also lots of irrelevant ones |
ChatGPT Deep Research |
157 | 6% | Most reliable of the mid-2025 set, lowest irrelevance |
Gemini |
325 | 31% | Highest volume but also highest irrelevance |
The Big Gotcha: Mid-2025 AI References Had Lots of Metadata Errors
The second major question was reliability: are the references AI returns accurate, or are they hallucinated?
The authors explicitly separate hallucinations into two types:
- Fabrications: the reference points to a nonexistent paper.
- Metadata mismatches: the paper exists, but some fields are wrong (title, first author, year, journal, DOI, or link).
This is a crucial distinction because a metadata mismatch can be just as damaging as a fabrication if you unknowingly cite the wrong version.
What happened when they verified the AI-only references
The authors verified all AI-generated references that were not found by a human, for a total of 641 references.
The results were stark:
- 211 (33%) were perfect (all fields matched)
- 22 (3%) were fabrications (nonexistent papers)
- 408 (64%) were metadata mismatches (real papers, but at least one field wrong)
That means that even when the AI didn’t fully invent something, the majority of the time it still got some detail wrong.
Which fields were most often wrong?
For the mismatches where the true paper could be resolved using DOI or a working link (the paper discusses 399 such cases), the mismatch distribution was mostly:
- First author: 306
- Title: 181
- Year: 105
- Journal: 124
And mismatches frequently involved two fields (the paper notes 206 cases where two of the four fields were wrong, in the DOI/link-resolved subset).
Model-by-model reliability in plain terms
The paper reports the key model ordering like this:
- ChatGPT Deep Research performed best: it produced no fabrications and had 78% perfect references in the AI-only set.
- Gemini produced the most fabrications (21) and the most mismatches.
- ChatGPT-4o produced only a single fabrication, but still had a very large fraction of metadata mismatches (76%).
So if you’re using AI for literature searching, “does it fabricate?” is not enough—metadata correctness is the real battlefield.
What Improved in 2026: ChatGPT Pro 5.5 Looked Perfect in a One-Project Spot Test
The authors didn’t run a full sweep of ChatGPT Pro 5.5 across all eight projects (that would require a lot of extra coordination and time). Instead, they did a focused comparison on one project: SU2 (Massive Yang–Mills theory).
Here’s what they did:
- Used the same prompts and procedure
- Kept a cap of 50 papers
- Compared ChatGPT Pro 5.5 against ChatGPT-4o for this project
Results: dramatic reduction in fabrication and mismatches (in this one test)
For SU2:
- ChatGPT Pro 5.5 generated 44 papers total
- ChatGPT-4o generated 37 substantially different papers (not counting near-duplicates as separate outcomes)
- Crucially: all ChatGPT Pro 5.5 references were perfect, with all fields matching title/author/year/journal.
The paper reports an overlap with the human expert selection of only 4 papers for the Pro model (compared to 4 overlaps for ChatGPT-4o as well), but the big difference is that the Pro model’s candidate citations were accurate.
Why this likely matters: it’s about tool-use + verification
The authors interpret this improvement as reflecting stronger tool use and verification in the ChatGPT Pro 5.5 system, rather than a magical change in “what the model understands.”
In other words: you may not just need a “smarter brain”—you need an AI workflow that checks its own output against sources.
Practical Implications: How to Use LLMs for Literature Search Without Getting Burned
So what should you do if you want to use AI in a real physics workflow tomorrow?
Think of an LLM literature search as two separate tasks:
1. Candidate generation (find “possible relevant” papers)
2. Citation verification (ensure those candidates are real and correctly described)
Based on the study, mid-2025 models can help more with the first task—finding candidates that sometimes humans missed—but you should assume the second task will often fail unless you use a tool-augmented mode or a newer system with stronger verification.
A simple “AI-assisted” workflow that matches the evidence
- Use AI to produce a candidate list, especially if you’re new to the subfield.
- Prefer modes that behave like search + verification (the paper finds
ChatGPT Deep Researchis far more reliable than standardChatGPT-4oin mid-2025). - Treat all AI-provided DOIs, titles, author lists, and years as untrusted until checked.
- If you’re building a reference list: verify at least DOI/link → then confirm title/first author/year/journal.
Why the low overlap isn’t necessarily bad
The small overlap (<6%) might sound like “AI is wrong.” But another interpretation is: AI searches differently than humans. Humans tend to move beyond keywords and incorporate foundational work and neighboring context. Meanwhile, AI appears to lean more keyword-driven in mid-2025.
That difference can be an advantage if you use AI as a parallel search generator—not as the final curator.
Where hallucination risk still matters
Even though fabrication rates were relatively low (3% fabrications among AI-only refs), the metadata mismatch rate was huge (64% for mid-2025). A fabrication is obvious once discovered, but metadata mismatches can be stealthy—especially when the paper exists and looks “close enough.”
This is why verification isn’t optional.
Key Takeaways
- Overlap is small: For mid-2025 models, the overlap between human-selected and AI-selected references was reported as <<6%, meaning AI doesn’t yet reproduce expert search coverage on its own.
- AI can still add value: Despite low overlap, AI generated a meaningful number of relevant papers that humans missed—so it can help broaden coverage.
- Model reliability varies dramatically:
- In the mid-2025 study,
ChatGPT Deep Researchproduced no fabrications among AI-only references and had the highest perfect rate (78%). Geminihad higher fabrication and mismatch rates and a 31% irrelevance rate among AI-only refs.ChatGPT-4ohad low fabrication (1 fabrication overall in the paper’s framing) but still had a large metadata mismatch problem.
- In the mid-2025 study,
- Mid-2025 citation accuracy is the main risk: Among 641 AI-only references, 64% were metadata mismatches and 3% were fabrications. So “real DOI but wrong details” is common.
ChatGPT Pro 5.5looked much better in a spot test: For the SU2 project, the paper reports allChatGPT Pro 5.5references were perfect (no mismatches or fabrications in that single-project evaluation).- Practical advice: Use LLMs to generate candidates, but verify citations—especially title, first author, year, journal, DOI, and link—before citing them in your own work.
- Future direction: The improvements likely come from stronger tool-use and verification, suggesting the best systems will be the ones that check against sources, not just generate text that sounds right.
If you’d like, I can also turn this into a “plug-and-play” checklist for verifying AI-generated references (DOI-first vs arXiv-first workflows), tailored for physics/astro papers.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review — arXiv
- Authors: Authors: Anamaria Hell, Kateryna Vovk, Veena Krishnaraj, Jia Liu, Kosuke Aizawa, Adrian E. Bayer, Linda Blot, Jessica Cowell, Suyog Garg, Jonathan Grée, Ben Hor