The Short Answer
Abstract-based prompting beats full-text prompting for accurate environmental paper citations, including matching identifiers like DOI and Scopus EIDs. Even when full text seems richer, the study found abstract prompts perform better for reference correctness.
Practically, prompt the LLM using abstracts when generating reference lists, then copy citations only after checking identifier fields (DOI/EID) and basic author/title matches—especially if you’re short on time for full validation.
Don’t assume accuracy for every generated reference: citation correctness varies by LLM platform and by rank within the output list, so lower-ranked citations are riskier than top-ranked ones.
On this page
- Introduction: LLMs + environmental literature search, but are the citations right?
- Why This Matters: Citation mistakes are easier to insert than to detect
- How the study tested LLM reference retrieval accuracy (and what “errors” meant)
- Abstract-only prompts beat full-text prompts: the main accuracy comparison
- Not all LLMs behave the same: platform-level performance differences are huge
- The “ranking position” trap: lower-listed references are riskier
- What fields were easiest vs hardest for LLMs to retrieve correctly
- So what should environmental scientists do with this?
- Key Takeaways
Generative Search Can Misquote Environmental Papers—Abstract vs Full Text
Introduction: LLMs + environmental literature search, but are the citations right?
If you’ve used an LLM to “find the key papers” in environmental science, you’ve probably noticed how fast it can spit out bibliographies. That speed is the whole appeal. But new research from Chen, Zhang & Zhang raises an uncomfortable question: when LLMs generate references, do they get the bibliographic details right—or do they quietly invent, guess, or mix things up?
This study zeroes in on a surprisingly practical detail: what you feed the LLM as the prompt. Do you get better reference retrieval when you prompt using an article’s abstract, or when you use the full text? The researchers tested multiple widely used LLM platforms and measured not just whether the references “sound relevant,” but whether identifiers like DOIs and Scopus EIDs actually match real records.
Based on new quantitative work in the environmental science domain, the key headline is: abstract-based prompting beats full-text prompting for citation accuracy, even though full text contains more information. The results also show that accuracy isn’t consistent—it varies by LLM platform, by journal, and even by where a reference appears in the model’s ranked output.
Why This Matters: Citation mistakes are easier to insert than to detect
Here’s the problem in plain terms: LLMs don’t just retrieve information—they produce structured outputs (like reference lists) that look “official” even when parts are wrong. In environmental science, where conclusions often depend on prior work (methods, datasets, causal claims, intervention comparisons), a fabricated or mismatched citation can ripple into reviews, grant proposals, and even downstream research decisions.
This is especially relevant right now because environmental research workflows are increasingly hybrid: people use LLMs for discovery, then rely on humans for validation. But validation usually happens after the citations get copied into documents. If the workflow doesn’t force a verification step (DOI/EID, link, author/title match), errors become “sticky.”
A real-world scenario you could apply today:
- You’re drafting an environmental policy brief or literature review.
- You paste an LLM prompt like “generate 10 key references related to this content…”
- You copy the references into your manuscript—even if you don’t open every DOI.
- Later, one wrong citation survives because it’s plausible, and the LLM already ranked it near the top.
The new research suggests something even more subtle: the lower the reference’s position in the LLM’s output list, the riskier it gets. That means “top-ranked papers” from the LLM are more trustworthy than the rest, but still not guaranteed.
Compared with earlier AI research that discussed hallucinated citations in general, this study is more operational: it measures error rates in a specific field (environmental science) and under a specific decision point people control constantly—abstract vs full-text prompting.
How the study tested LLM reference retrieval accuracy (and what “errors” meant)
The researchers pulled original research articles from five leading environmental science journals:
- Energy and Environmental Science
- Nature Sustainability
- Nature Climate Change
- Lancet Planetary Health
- Environmental Science and Technology
They selected papers from 2024 and 2025: five randomly selected original articles per journal per year. That gives 50 original articles total.
Then they ran literature retrieval in two prompt styles:
1. Abstract-only prompting
2. Full-text prompting (full text provided to the LLM after removing references)
They tested six free-tier LLM platforms (legacy models used as of August 2026 data collection):
- Claude (Sonnet 4.6)
- ChatGPT (GPT-5.3)
- Grok (Grok 4 or 4.2)
- DeepSeek (DeepSeek V3)
- Perplexity (Nemotron 3 Super)
- Gemini (Gemini 3.5 Flash)
For each original article and each LLM, the model generated 10 references, so the study analyzed:
- 500 references from abstract-only prompting
- 500 references from full-text prompting
Across both styles, that totals 6,000 retrieved references.
What “accuracy” meant: a multi-metric score, not vibes
A key strength here: they didn’t judge correctness by “looks right.” They validated each retrieved reference against multiple identifiers and also checked whether it was truly relevant to the index article.
They scored each retrieved reference with six metrics:
- Validity of bibliographic data (title/author/journal/year/volume/issue/first page)
- Google Scholar link
- DOI
- Scopus EID
- Relevance (whether the retrieved article was the index article or cited by it)
- Plus they used a “complete fabrication” rule (below)
They computed:
- Score ratio = total score / 6 (so it ranges from 0 to 1)
And they defined complete fabrication as:
- Total score = 0, meaning none of the validated metrics matched.
This lets them separate:
- “Partly correct but incomplete” from
- “Totally wrong / fabricated” cases.
Abstract-only prompts beat full-text prompts: the main accuracy comparison
The headline result is clear: abstract-only prompting improved retrieval accuracy compared to full-text prompting.
Here’s what they found across all platforms and all prompts:
| Input method | Mean score ratio | Complete fabrication rate |
|---|---|---|
| Abstract-only | 0.43 | 27.8% |
| Full text | 0.40 | 30.4% |
Statistically, the differences mattered:
- Score ratio: p < 0.001
- Complete fabrication rate: p = 0.027
Which parts improved with abstract-only?
The paper reports that abstract-only prompting improved:
- DOI accuracy
- Scopus EID accuracy
- complete fabrication rate
- overall score ratio
But it did not significantly improve:
- relevance score
- Google Scholar link score
That’s an important nuance. Abstracts may help the model get identifiers and bibliographic structure right—but that doesn’t automatically guarantee it will pick the most cited/most directly connected papers.
Why would abstract help more than full text?
The authors lay out a plausible explanation without pretending to know the exact mechanism. Think of it like a memory task:
- Full text is long—more “noise” competes for attention.
- The abstract is a curated summary—so relevant details are more concentrated.
They also connect this to a known “lost in the middle” effect in long-context models: important tokens can get buried when prompts get too long.
So instead of asking “more text is always better,” their evidence supports: task-aligned compression (abstracts) can make retrieval more reliable.
Not all LLMs behave the same: platform-level performance differences are huge
Even if abstract prompting helps, you still need to ask: which model are we talking about?
The researchers found large differences across LLM platforms in mean score ratio and fabrication rate.
Mean score ratio by platform (highest to lowest)
Claude: 0.64 (highest)Grok: 0.60ChatGPT: 0.40DeepSeek: 0.39Perplexity: 0.32Gemini: 0.14 (lowest)
Complete fabrication rate by platform
Claude: 4% (lowest)Gemini: 71% (highest)
That spread is so large it changes how you should use these tools. If one platform is generating structured references that are correct only 1 time in ~7, then “just double-check the top 3” is not a safe strategy.
Abstract vs full text within each platform
When the authors compared abstract-only vs full-text within each platform:
- Grok: abstract-only higher score ratio (p = 0.03)
- DeepSeek: abstract-only higher score ratio (p = 0.001)
- For the other models, the abstract/full-text difference wasn’t statistically significant in that within-model comparison.
So the abstract advantage is strongest in certain models—but the overall cross-platform advantage still held.
The “ranking position” trap: lower-listed references are riskier
One of the most practically useful findings is about output sequence—where a reference appears in the model’s list.
They reported that:
- reference position was associated with retrieval accuracy
- lower positions corresponded to:
- lower score ratios
- higher odds of complete fabrication
In regression terms (multilevel mixed-effects analysis), output sequence had an adjusted association of score ratio:
- coefficient about -0.03 with p < 0.001
and the logistic model showed that output sequence was also independently linked with complete fabrication.
What this suggests about model behavior
The authors interpret this as evidence that the model likely ranks retrieved references internally, then generates the output list in that ranked order. If so, your best move is not to trust the entire list equally—it’s to treat the list as tiered, like you would with search results.
But importantly: even top-ranked references are not guaranteed to be correct, so validation is still necessary.
What fields were easiest vs hardest for LLMs to retrieve correctly
The study breaks down component-level performance (how often each metric matched).
Across all retrieved references, the best accuracies were:
- Source title: 66.7%
- Volume: 63.1%
- Author: 60.5%
- DOI: 58.2%
- Year: 68.5% (reported among highest)
- Google Scholar link: 42.3% (lower)
The weakest were:
- Relevance score: 25.4%
- Scopus EID: 12.3% (very low)
So even when bibliographic structure (year/volume/title/author) looks solid, some identifiers—especially EID—are much harder for LLMs to get right.
And that aligns with what many researchers experience in practice: DOI errors are annoying but often correctable; EID mismatches can be harder because fewer people routinely check them.
So what should environmental scientists do with this?
The paper’s limitations matter: it used free-tier models, two prompt templates, and a small sample of 50 original articles (10 references each, across 6 platforms, across 2 prompt types). Still, the consistency of the main pattern—abstract prompting > full-text prompting—and the size of platform differences make the findings actionable.
If you want a “today” workflow that reduces citation errors:
- Prefer abstract-based prompting over full-text prompting when your goal is bibliographic retrieval accuracy.
(This study’s evidence shows a measurable improvement in score ratio and lower fabrication rate.) - Validate DOI at minimum for any reference you intend to cite. This is one of the more frequently correct fields (58.2% overall accuracy).
- Treat output ranking position as a risk signal:
- check the top references carefully,
- and don’t assume later ones are even approximately accurate.
- If your platform is Gemini or Perplexity, assume more errors and budget more time for verification.
- Gemini’s complete fabrication rate of 71% is a flashing warning light.
And if you’re building automated literature review pipelines (or using LLMs for evidence synthesis), this study supports the idea that retrieval should be paired with verification, ideally DOI/EID checks and link validation. The authors also point toward retrieval-augmented generation (RAG) style systems as a likely improvement direction.
Key Takeaways
- Abstract-only prompting beats full-text prompting for LLM-assisted reference retrieval in environmental science:
- score ratio 0.43 vs 0.40
- complete fabrication rate 27.8% vs 30.4%
- Accuracy is moderate and inconsistent:
- mean score ratio across both modes was 0.42 (range 0–1)
- LLM platforms vary wildly:
Claudehad the best performance (0.64, 4% fabrication)Geminiwas the worst (0.14, 71% fabrication)
- Reference output position matters:
- lower-listed references had lower accuracy and higher fabrication odds
- Some fields are much harder than others:
- DOI was fairly recoverable (58.2% match rate),
- Scopus EID was very low (12.3% match rate)
- Practical advice: use abstract prompts, validate DOIs, and treat later references as higher risk, even if they look relevant.
If you want, tell me which LLMs you’re currently using and what your typical literature workflow looks like (manual reading vs copy-paste vs RAG tooling). I can suggest a tighter prompt + verification checklist tailored to your setup.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- Errors of LLM-Assisted Literature Retrieval in Environmental Science: A Comparison Study of Abstract versus Full-text Based Prompts — arXiv
- Authors: Authors: Yanjun Chen, Yongfeng Zhang, Lanjing Zhang