AI for Grant Proposals: Testing Bias in Scientific Planning

Researchers are using LLMs to draft grant plans, but the real risk is bias in how proposals get scored. This study tests human and AI reviewers on one-page physics, astrophysics, and cosmology plans—and finds AI reviewers favor AI-written work.
The finding Human reviewers rated human- and AI-written plans similarly overall, while AI reviewers preferred AI-written proposals by about one point.
The method Researchers generated one-page plans for the same eight physics/astrophysics/cosmology projects using humans vs three LLMs, then scored them blindly.
The caveat If AI systems are used for evaluation, they may introduce systematic preference for AI-authored writing, so bias testing is required.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

LLM-written one-page scientific project plans can score similarly to human-written plans with human reviewers, but AI reviewers systematically rate AI-written proposals about one point higher on a 1–5 scale.

So what: if you use LLMs for proposal drafting or evaluation, you should test for evaluator bias—especially in AI-based scoring—to avoid creating feedback loops that reward AI-style writing over realistic feasibility.

Caveat: the study uses a controlled setup (eight projects, 32 proposals, fixed inputs, and specific models/rubrics), so you should run bias checks in your own domain and review workflow before broad deployment.

AI for Grant Proposals: Testing Bias in Scientific Planning

Introduction: LLMs are writing science proposals—so who’s actually judging?

If you’ve been around grant season lately, you’ve probably noticed a quiet shift: researchers are increasingly using large language models (LLMs) to draft sections, polish wording, sketch timelines, and generally “get the proposal moving.” That’s especially tempting in physics, astrophysics, and cosmology, where projects can be complex and writing a coherent plan takes real effort.

But here’s the catch: a model can write a convincing plan, while the review process might still be influenced by cues that have nothing to do with scientific quality. The new research behind this post digs directly into that problem—specifically, how well LLMs handle scientific project planning and proposal evaluation. This blog is based on new research from the original paper: AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation.

The study tests a very practical question: when an AI drafts a one-page research plan, do humans (and AIs) rate it as good science—or do they accidentally reward “AI-style” writing? And can reviewers reliably tell whether a proposal was written by a human or an LLM?

Why This Matters: Bias in the “writing” stage can turn into bias in funding

This matters right now because proposal evaluation isn’t just about the content—it’s also about how reviewers interpret clarity, structure, feasibility, and risk. If an LLM tends to produce highly readable, tightly structured text, reviewers may subconsciously treat that as stronger planning, even when feasibility is overstated. In other words, “looks like a good plan” can start to overpower “is a realistic plan.”

Imagine a real-world scenario: a funding agency uses AI tools to help triage, summarize, or even (in some workflows) provide rubric scores. If the AI evaluator has a systematic preference for AI-written proposals—as this study finds—then the system could create a feedback loop: AI-assisted applicants get rewarded, which then incentivizes more AI-style writing. Human applicants who prefer to write more cautiously, with more idiosyncratic detail, could be disadvantaged.

This also builds on earlier AI research. Previous work has found LLMs can generate research ideas and produce helpful feedback, sometimes even with quality differences or reviewer disagreements. This paper complements that by isolating the planning document itself and running a blinded, controlled comparison where both human and AI evaluators score the same fixed project goals. That’s the part that turns it from “interesting” into “actionable,” especially for agencies designing review protocols.

What the Researchers Actually Compared: Human vs LLM proposals, scored blindly

The study focuses on one specific stage of the workflow: project planning, where you draft a compact proposal specifying the scientific goal, the methodology, and the resources/timeline needed to actually carry out the work.

The setup: 8 projects → 32 proposals → 6 evaluators

The researchers assembled a controlled corpus of eight expert-conceived research projects in physics, astrophysics, and cosmology. For each project, they started with the same inputs:
- a title
- a background paragraph
- a goal paragraph

Those three elements were provided verbatim to both the human planner and the AI system. The only thing that changed was who wrote the one-page plan and which model produced it.

Then they produced:
- 1 human-written proposal per project (8 total)
- 3 AI-written proposals per project (24 total)

So the final dataset was 32 proposals total.

Finally, every proposal was evaluated (blindly) by:
- 4 human reviewers
- 2 AI reviewers (LLMs)

Who did what (and why it helps isolate bias)

A key design choice is that the AI proposals were generated in a way that mirrors real use—but still stays controlled.

  • Human project planners (experts): wrote the plan without AI assistance.
  • AI prompters: used non-expert students to generate AI proposals. Each prompter took the fixed project title/background/goal and prompted three different assistants to produce three proposals per project, with a fixed template and prompt.
  • Human reviewers: judged authorship (human vs AI) and then scored quality using a rubric.
  • AI reviewers: later scored the same anonymized proposals using a similar process.

To reduce superficial “tells,” proposals followed an identical one-page template and were anonymized. Human-written proposals also got a light grammar/spelling pass to prevent obvious mistakes from making them easier to detect as human.

The Writing Format: A one-page template that tests real planning, not grant-length performance

Each proposal was about one page and followed the same structure with four sections:
1. Project Title
2. Background (1 sentence)
3. Goal (1 sentence)
4. Methodology split into no more than five major steps/phases, each with approximate completion time

That template intentionally caps length to keep the evaluation workload manageable. Importantly, the study evaluates planning quality, not idea novelty. All the underlying scientific ideas were expert-conceived and held fixed across human and AI versions—so differences in scores come from how the plan is expressed, structured, and made feasible (or not).

This is a smart constraint because it targets the exact failure mode you’d worry about in proposal review: even with a good idea, plans can become overconfident, overly generic, or poorly risk-managed.

And yes, this study is part of a broader line of work: the authors note a companion paper focusing on literature review ability (their Paper I). Here, the focus is the planning stage, which is where many proposal evaluation bottlenecks actually happen.

Authorship Detection: Humans miss AI signatures more often than you’d expect

The first major analysis asks: how reliably can evaluators tell whether a proposal was human-written or AI-written? This is important because detecting authorship can shape scoring—sometimes consciously, sometimes not.

Origin classification accuracy (the headline numbers)

Across the four human reviewers:
- Correctly identified AI-written proposals: 79%
- Correctly identified human-written proposals: 72%

Not perfect, but well above chance. And importantly, humans sometimes mistook human proposals for AI slightly more often than the other way around.

Meanwhile, the two most capable AI reviewers were essentially perfect on this small dataset:
- Claude Opus 4.8: 100%
- ChatGPT Pro 5.5: 100%

There was also an intermediate model (Claude Sonnet 4.6) that was less reliable on AI-written proposals (details below), suggesting that the “perfect” detection isn’t universal across model tiers.

Model comparison (authors’ observed performance)

Evaluator Correctly identifies human-written Correctly identifies AI-written
Human reviewers (avg) 72% 79%
Claude Opus 4.8 (AI reviewer) 100% 100%
ChatGPT Pro 5.5 (AI reviewer) 100% 100%
Claude Sonnet 4.6 (AI reviewer) 100% 67%

What cues drove the judgments?

This part is fascinating because it reveals that people (and models) are looking at style and structure cues that often correlate with AI writing.

Human reviewers often flagged proposals as AI-written when they looked:
- “too clean”
- “very tight” or template-like (e.g., five-step structure)
- missing citations
- using unrealistic/round-number timelines
- listing irrelevant or nonsensical tools/methods
- leaning on ML buzzwords (including LLM-adjacent phrasing)

They marked proposals as human when they included:
- specific author–year citations (and sometimes self-citations)
- concrete parameter ranges or observational priors
- more idiosyncratic jargon
- first-person language (“we”)
- insider domain markers (examples cited in the paper include references that hint at specific instruments/institutes)

Human judgments were inconsistent: different reviewers treated the same feature differently, and some used heuristic shortcuts (like guessing “AI” if they noticed patterns).

The AI reviewers, however, claimed to rely on a consistent feature set tied to repeatable structure and generic deliverables (like “public white paper” or “community code release”), and they never misclassified in this experiment.

Proposal Quality Scores: AI writing wasn’t rated worse—AI reviewing was the twist

Now for the core question: do LLM-generated plans score worse? The surprising answer is: not from humans.

What human reviewers thought

Human reviewers rated proposals in a narrow band:
- Human-written average: about 3.52 / 5
- AI-written average: about 3.51 / 5

So, to human reviewers, AI-written plans were about equal overall to expert-written ones—sometimes even slightly better for some models (e.g., one AI variant scored highest among authors).

But human scores were noisy because reviewers disagreed substantially. That matters: noise in human evaluation can hide biases, but it can also make it harder to detect systematic effects.

What AI reviewers did differently

The AI reviewers showed a clear pro-AI bias.

They scored AI-written proposals about one point higher on the 1–5 scale compared with human-written ones:
- Claude Opus 4.8: AI authors averaged 4.47 vs human 3.59
- ChatGPT Pro 5.5: AI authors averaged 4.41 vs human 3.47

And this held up in additional variants (including a rate-first/origin-after prompting order, and using a smaller model). The bias seems tied to AI-style evaluation rather than ordering effects.

Where the bias showed up most (by rubric aspect)

The rubric has four aspects scored from 1 to 5:
1. Clarity and structure
2. Appropriateness of methods to the scientific goal
3. Resource and tool planning
4. Feasibility, timeline, and risk awareness

The pro-AI gap was strongest for:
- Clarity and Structure (AI-written near ~4.9–5.0, human around ~3.75)
- Resource and Tool Planning (humans scored lowest by AI reviewers)

It was weaker for:
- Appropriateness of Methods to Scientific Goal (human-written was actually near the top in this aspect)

And interestingly, Feasibility/Timeline/Risk was consistently the weakest dimension for AI-written proposals across evaluators—AI plans tended to be less realistic or overly rounded in timelines. But even there, AI reviewers still scored AI-written proposals higher overall.

The Hidden Twist: Bias depends on how strong the human plan is

One of the most important insights is that the average “humans rate AI and human plans similarly” can hide a lot.

Across the eight projects:
- human reviewers preferred human-written proposals in 5 out of 8
- AI-written proposals scored higher in the remaining 3

And the size/direction of that gap wasn’t random. It was tightly linked to how strong the human plan was:
- the authors report a strong anti-correlation between human-vs-AI score differences and human proposal quality, with Pearson r = −0.95

Meaning: when the expert’s plan was strong and detailed, reviewers tended to favor it. When the expert’s plan was weaker, the AI plan often caught up or surpassed it.

The AI reviewers showed a similar shape but shifted upward by about a consistent offset—i.e., they preferred AI-written proposals in every project, but the magnitude still depended on human proposal quality.

This suggests a practical takeaway: AI writing may be most helpful (or most competitive) when human drafts are underdeveloped. But it also means AI evaluation systems may reward polished structure even when feasibility details are shaky—creating an evaluation bias that is amplified when human submissions vary in baseline quality.

What This Means for Anyone Using LLMs in Proposal Prep or Review

If you’re a researcher, these results don’t say “don’t use LLMs.” They say “use them carefully and don’t outsource judgment to an AI reviewer without safeguards.”

Here are practical implications you can apply today:

For applicants (proposal writers)

  • Don’t assume AI-written ≈ better science. Human reviewers in this study rated overall quality about the same, but AI reviewers showed a strong preference for AI-style plans.
  • Treat feasibility and risk realism as non-negotiable. The weakest rubric aspect for AI-written proposals was feasibility/timeline/risk awareness, even though AI evaluators still leaned positive overall.
  • Avoid “generic template overreach.” Several cues that triggered suspicion (or different scoring) were template-like structure, generic deliverables, and rounded timelines.

For reviewers and agencies (review system designers)

  • If AI is scoring proposals, bias control is mandatory. The AI reviewers were consistently pro-AI, and the effect was about ~1 point overall on a five-point rubric.
  • Authorship detection may not be reliable—so design for that. Humans misclassified often enough that any process relying on “guessing AI authorship” would be unstable.
  • Rubric design matters. Bias was strongest in clarity/structure and resource/tool planning—areas where LLMs shine at sounding coherent. Agencies may need calibration steps or blind procedures that actively check for feasibility inflation.

The authors also connect their findings to real policies: major funders (like NIH, NSF, and ERC) already restrict or forbid generative-AI usage in parts of peer review. This paper provides empirical motivation for why those restrictions exist—especially when evaluative judgments might be swayed by “AI style” rather than scientific merit.

Key Takeaways

  • LLMs can draft scientific project plans that are comparable to expert-written one-page proposals in the eyes of human reviewers (overall human vs AI scores were essentially equal).
  • Humans aren’t great at authorship detection: human reviewers correctly identified AI-written proposals 79% of the time and human-written proposals 72%.
  • Top AI reviewers were perfect at detection in this small study (Claude Opus 4.8 and ChatGPT Pro 5.5 scored 100% accuracy), but that doesn’t mean detection is broadly reliable across bigger and messier settings.
  • The real risk is evaluator bias: AI reviewers scored AI-written proposals about ~1 point higher on a 1–5 rubric and showed a systematic preference for AI-style plans.
  • Bias concentrates in certain rubric aspects, especially clarity/structure and resource/tool planning; AI plans still tend to be weaker on feasibility/timeline/risk awareness.
  • Project-to-project variation is huge: human reviewers favored human-written plans in 5 of 8 projects, and the AI advantage depended strongly on how strong the human plan already was.
  • For funding systems: if AI tools are used in proposal scoring, expect pro-AI bias unless safeguards are built into the workflow (and policy restrictions make more sense than they might seem at first glance).

If you want, tell me how you’re planning to use LLMs—proposal drafting, internal triage, or evaluation. I can suggest a checklist tailored to that exact workflow so you get the productivity benefits without stepping into the bias traps this study highlights.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

LLM Bias Testing with Psychology-Grade Prompts: What Works

Sex Bias in AI Clinician Reasoning: How Large Language Models Mirror Medical Stereotypes

Tool-Augmented AI Agents for Wireless Network Planning: Small Models, Big Impact

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.