Generative Search Can Misquote Environmental Papers—Abstract vs Full Text

Using an LLM to generate “key references” for environmental science can introduce citation errors. A new comparison of abstract vs full-text prompting shows abstract prompts improve DOI/Scopus EID accuracy.
The finding Abstract-only prompting produces more accurate DOI/Scopus EID matches than full-text prompting for environmental references generated by LLMs.
Why it matters Misquoted citations are easy to insert and hard to detect, so identifier mismatches can silently propagate into reviews and decisions.
What to do Use abstract prompts for faster discovery, but verify DOI/EID and author/title details before finalizing any bibliography.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

Abstract-based prompting beats full-text prompting for accurate environmental paper citations, including matching identifiers like DOI and Scopus EIDs. Even when full text seems richer, the study found abstract prompts perform better for reference correctness.

Practically, prompt the LLM using abstracts when generating reference lists, then copy citations only after checking identifier fields (DOI/EID) and basic author/title matches—especially if you’re short on time for full validation.

Don’t assume accuracy for every generated reference: citation correctness varies by LLM platform and by rank within the output list, so lower-ranked citations are riskier than top-ranked ones.

Generative Search Can Misquote Environmental Papers—Abstract vs Full Text

Introduction: LLMs + environmental literature search, but are the citations right?

If you’ve used an LLM to “find the key papers” in environmental science, you’ve probably noticed how fast it can spit out bibliographies. That speed is the whole appeal. But new research from Chen, Zhang & Zhang raises an uncomfortable question: when LLMs generate references, do they get the bibliographic details right—or do they quietly invent, guess, or mix things up?

This study zeroes in on a surprisingly practical detail: what you feed the LLM as the prompt. Do you get better reference retrieval when you prompt using an article’s abstract, or when you use the full text? The researchers tested multiple widely used LLM platforms and measured not just whether the references “sound relevant,” but whether identifiers like DOIs and Scopus EIDs actually match real records.

Based on new quantitative work in the environmental science domain, the key headline is: abstract-based prompting beats full-text prompting for citation accuracy, even though full text contains more information. The results also show that accuracy isn’t consistent—it varies by LLM platform, by journal, and even by where a reference appears in the model’s ranked output.

Why This Matters: Citation mistakes are easier to insert than to detect

Here’s the problem in plain terms: LLMs don’t just retrieve information—they produce structured outputs (like reference lists) that look “official” even when parts are wrong. In environmental science, where conclusions often depend on prior work (methods, datasets, causal claims, intervention comparisons), a fabricated or mismatched citation can ripple into reviews, grant proposals, and even downstream research decisions.

This is especially relevant right now because environmental research workflows are increasingly hybrid: people use LLMs for discovery, then rely on humans for validation. But validation usually happens after the citations get copied into documents. If the workflow doesn’t force a verification step (DOI/EID, link, author/title match), errors become “sticky.”

A real-world scenario you could apply today:
- You’re drafting an environmental policy brief or literature review.
- You paste an LLM prompt like “generate 10 key references related to this content…”
- You copy the references into your manuscript—even if you don’t open every DOI.
- Later, one wrong citation survives because it’s plausible, and the LLM already ranked it near the top.

The new research suggests something even more subtle: the lower the reference’s position in the LLM’s output list, the riskier it gets. That means “top-ranked papers” from the LLM are more trustworthy than the rest, but still not guaranteed.

Compared with earlier AI research that discussed hallucinated citations in general, this study is more operational: it measures error rates in a specific field (environmental science) and under a specific decision point people control constantly—abstract vs full-text prompting.

How the study tested LLM reference retrieval accuracy (and what “errors” meant)

The researchers pulled original research articles from five leading environmental science journals:
- Energy and Environmental Science
- Nature Sustainability
- Nature Climate Change
- Lancet Planetary Health
- Environmental Science and Technology

They selected papers from 2024 and 2025: five randomly selected original articles per journal per year. That gives 50 original articles total.

Then they ran literature retrieval in two prompt styles:
1. Abstract-only prompting
2. Full-text prompting (full text provided to the LLM after removing references)

They tested six free-tier LLM platforms (legacy models used as of August 2026 data collection):
- Claude (Sonnet 4.6)
- ChatGPT (GPT-5.3)
- Grok (Grok 4 or 4.2)
- DeepSeek (DeepSeek V3)
- Perplexity (Nemotron 3 Super)
- Gemini (Gemini 3.5 Flash)

For each original article and each LLM, the model generated 10 references, so the study analyzed:
- 500 references from abstract-only prompting
- 500 references from full-text prompting
Across both styles, that totals 6,000 retrieved references.

What “accuracy” meant: a multi-metric score, not vibes

A key strength here: they didn’t judge correctness by “looks right.” They validated each retrieved reference against multiple identifiers and also checked whether it was truly relevant to the index article.

They scored each retrieved reference with six metrics:
- Validity of bibliographic data (title/author/journal/year/volume/issue/first page)
- Google Scholar link
- DOI
- Scopus EID
- Relevance (whether the retrieved article was the index article or cited by it)
- Plus they used a “complete fabrication” rule (below)

They computed:
- Score ratio = total score / 6 (so it ranges from 0 to 1)

And they defined complete fabrication as:
- Total score = 0, meaning none of the validated metrics matched.

This lets them separate:
- “Partly correct but incomplete” from
- “Totally wrong / fabricated” cases.

Abstract-only prompts beat full-text prompts: the main accuracy comparison

The headline result is clear: abstract-only prompting improved retrieval accuracy compared to full-text prompting.

Here’s what they found across all platforms and all prompts:

Input method Mean score ratio Complete fabrication rate
Abstract-only 0.43 27.8%
Full text 0.40 30.4%

Statistically, the differences mattered:
- Score ratio: p < 0.001
- Complete fabrication rate: p = 0.027

Which parts improved with abstract-only?

The paper reports that abstract-only prompting improved:
- DOI accuracy
- Scopus EID accuracy
- complete fabrication rate
- overall score ratio

But it did not significantly improve:
- relevance score
- Google Scholar link score

That’s an important nuance. Abstracts may help the model get identifiers and bibliographic structure right—but that doesn’t automatically guarantee it will pick the most cited/most directly connected papers.

Why would abstract help more than full text?

The authors lay out a plausible explanation without pretending to know the exact mechanism. Think of it like a memory task:
- Full text is long—more “noise” competes for attention.
- The abstract is a curated summary—so relevant details are more concentrated.

They also connect this to a known “lost in the middle” effect in long-context models: important tokens can get buried when prompts get too long.

So instead of asking “more text is always better,” their evidence supports: task-aligned compression (abstracts) can make retrieval more reliable.

Not all LLMs behave the same: platform-level performance differences are huge

Even if abstract prompting helps, you still need to ask: which model are we talking about?

The researchers found large differences across LLM platforms in mean score ratio and fabrication rate.

Mean score ratio by platform (highest to lowest)

  • Claude: 0.64 (highest)
  • Grok: 0.60
  • ChatGPT: 0.40
  • DeepSeek: 0.39
  • Perplexity: 0.32
  • Gemini: 0.14 (lowest)

Complete fabrication rate by platform

  • Claude: 4% (lowest)
  • Gemini: 71% (highest)

That spread is so large it changes how you should use these tools. If one platform is generating structured references that are correct only 1 time in ~7, then “just double-check the top 3” is not a safe strategy.

Abstract vs full text within each platform

When the authors compared abstract-only vs full-text within each platform:
- Grok: abstract-only higher score ratio (p = 0.03)
- DeepSeek: abstract-only higher score ratio (p = 0.001)
- For the other models, the abstract/full-text difference wasn’t statistically significant in that within-model comparison.

So the abstract advantage is strongest in certain models—but the overall cross-platform advantage still held.

The “ranking position” trap: lower-listed references are riskier

One of the most practically useful findings is about output sequence—where a reference appears in the model’s list.

They reported that:
- reference position was associated with retrieval accuracy
- lower positions corresponded to:
- lower score ratios
- higher odds of complete fabrication

In regression terms (multilevel mixed-effects analysis), output sequence had an adjusted association of score ratio:
- coefficient about -0.03 with p < 0.001
and the logistic model showed that output sequence was also independently linked with complete fabrication.

What this suggests about model behavior

The authors interpret this as evidence that the model likely ranks retrieved references internally, then generates the output list in that ranked order. If so, your best move is not to trust the entire list equally—it’s to treat the list as tiered, like you would with search results.

But importantly: even top-ranked references are not guaranteed to be correct, so validation is still necessary.

What fields were easiest vs hardest for LLMs to retrieve correctly

The study breaks down component-level performance (how often each metric matched).

Across all retrieved references, the best accuracies were:
- Source title: 66.7%
- Volume: 63.1%
- Author: 60.5%
- DOI: 58.2%
- Year: 68.5% (reported among highest)
- Google Scholar link: 42.3% (lower)

The weakest were:
- Relevance score: 25.4%
- Scopus EID: 12.3% (very low)

So even when bibliographic structure (year/volume/title/author) looks solid, some identifiers—especially EID—are much harder for LLMs to get right.

And that aligns with what many researchers experience in practice: DOI errors are annoying but often correctable; EID mismatches can be harder because fewer people routinely check them.

So what should environmental scientists do with this?

The paper’s limitations matter: it used free-tier models, two prompt templates, and a small sample of 50 original articles (10 references each, across 6 platforms, across 2 prompt types). Still, the consistency of the main pattern—abstract prompting > full-text prompting—and the size of platform differences make the findings actionable.

If you want a “today” workflow that reduces citation errors:

  1. Prefer abstract-based prompting over full-text prompting when your goal is bibliographic retrieval accuracy.
    (This study’s evidence shows a measurable improvement in score ratio and lower fabrication rate.)
  2. Validate DOI at minimum for any reference you intend to cite. This is one of the more frequently correct fields (58.2% overall accuracy).
  3. Treat output ranking position as a risk signal:
    • check the top references carefully,
    • and don’t assume later ones are even approximately accurate.
  4. If your platform is Gemini or Perplexity, assume more errors and budget more time for verification.
    • Gemini’s complete fabrication rate of 71% is a flashing warning light.

And if you’re building automated literature review pipelines (or using LLMs for evidence synthesis), this study supports the idea that retrieval should be paired with verification, ideally DOI/EID checks and link validation. The authors also point toward retrieval-augmented generation (RAG) style systems as a likely improvement direction.

Key Takeaways

  • Abstract-only prompting beats full-text prompting for LLM-assisted reference retrieval in environmental science:
    • score ratio 0.43 vs 0.40
    • complete fabrication rate 27.8% vs 30.4%
  • Accuracy is moderate and inconsistent:
    • mean score ratio across both modes was 0.42 (range 0–1)
  • LLM platforms vary wildly:
    • Claude had the best performance (0.64, 4% fabrication)
    • Gemini was the worst (0.14, 71% fabrication)
  • Reference output position matters:
    • lower-listed references had lower accuracy and higher fabrication odds
  • Some fields are much harder than others:
    • DOI was fairly recoverable (58.2% match rate),
    • Scopus EID was very low (12.3% match rate)
  • Practical advice: use abstract prompts, validate DOIs, and treat later references as higher risk, even if they look relevant.

If you want, tell me which LLMs you’re currently using and what your typical literature workflow looks like (manual reading vs copy-paste vs RAG tooling). I can suggest a tighter prompt + verification checklist tailored to your setup.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Generative AI and Stack Overflow: Why “More Answers” ≠ Equal Recognition

Generative AI in VS Code Issues: What Developers Actually Complain About

LLM Mental Health Bias: LGBTQIA+ Identity Changes Context, Not Help

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime • 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.