The Short Answer
ChatGPT’s research quality scoring shows a slightly male-favourable pattern in many REF 2021 areas, though the differences are usually small and not consistent everywhere. The study compares female-first-author vs male-first-author papers using prompts with author information withheld.
If you’re using AI to pre-score or triage research, this means “no names in the prompt” doesn’t guarantee neutrality—gender-associated patterns can still arise via indirect links to fields, topics, or writing patterns. Plan for human validation to avoid uneven downstream effort.
The effects are not uniform across all Units of Assessment, and the paper examines explanations such as abstract complexity (Flesch–Kincaid Grade) without showing a neat account of the pattern. Treat the findings as a caution about second-order effects, not as proof of direct bias from author names.
On this page
- Introduction: Is ChatGPT’s “quality” scoring getting gendered?
- What the study found, in plain terms
- Why This Matters (Right Now)
- What the Researchers Actually Measured: REF proxy scores vs ChatGPT “quality” ranks
- The main results: Where male-first-author papers got higher ChatGPT quality scores
- Comparing ChatGPT vs REF proxy: How the gender pattern changes depending on the metric
- Why it might be happening: Citations and abstract complexity (but not in a simple way)
- The solo-author check: Reducing authorship-role ambiguity
- Key Takeaways
ChatGPT vs Gender in Research Scores: What the Data Show
Introduction: Is ChatGPT’s “quality” scoring getting gendered?
If you’re using AI tools to help with research evaluation, here’s a question that matters more than it sounds: does ChatGPT score research quality differently by gender? New research based on the REF 2021 journal-paper submissions digs into this—and it’s not just about “direct bias” from names in prompts. The study by Kayvan Kousha and Mike Thelwall (from the original paper) checks whether first-author gender is associated with higher or lower ChatGPT research quality scores.
The setup is clever (and important). ChatGPT was given titles and abstracts only, with author information withheld to avoid obvious name-based effects. Still, the researchers test whether scores differ systematically between female-first-author and male-first-author papers. Their dataset is huge: 89,744 UK REF 2021 journal articles, spread across 34 REF Units of Assessment (UoAs).
What the study found, in plain terms
Overall, the pattern is slightly male-favourable for ChatGPT in many areas—especially in health, science, and engineering-related UoAs—but the differences are usually small and not consistent everywhere. In most UoAs, the direction of gender differences is similar for both ChatGPT scores and the REF “quality” proxy the study could access. The researchers also compare relative advantage, not just raw score direction, by converting both systems into percentile ranks within each UoA.
They also look for an explanation: maybe ChatGPT is responding to writing style. The paper tests abstract complexity (using the Flesch–Kincaid Grade) and finds that abstract complexity differences don’t neatly explain the ChatGPT gender pattern. They then explore related factors like citation impact, plus a check on solo-authored papers to reduce ambiguity in interpreting “first author = lead contribution.”
All of this adds up to a clear caution: even when gender cues are removed from the prompt, AI-assisted research evaluation can still show gender-associated patterns—likely driven by indirect factors tied to fields, topics, methods, and publishing structures.
Why This Matters (Right Now)
This matters right now because many research systems are under pressure to speed up evaluation, triage outputs, and standardise judgments. If AI starts “helpfully” pre-scoring papers, people may assume the tool is neutral—especially if it never sees names. But this paper suggests neutrality-by-default may not be guaranteed.
A real-world scenario you can imagine today
Picture a UK-style environment where a university (or a funding body) wants to identify which papers are “most likely high-impact” before sending them to expensive expert review. Suppose an internal workflow uses ChatGPT to prioritise manuscripts or sort them into bands. Even if author gender is removed, the model’s preferences might still line up with the demographics of different subfields. That could lead to uneven downstream effort: more human time going to papers that ChatGPT ranks higher in certain areas, reinforcing the evaluation pipeline rather than correcting it.
How it builds on earlier AI fairness research
There’s a broader research thread on AI bias in evaluation. Other experiments have shown gender effects depending on what metadata the model receives and how prompts are constructed (including studies where changing author metadata changes LLM judgments). This paper builds on that idea by testing a tougher variant: no author info in the prompt—only title and abstract. So the message is stronger: even without “name-based” bias, gender differences may surface through second-order correlations (field/topic/method/journal context).
In short: the significance isn’t just “bias exists,” but bias can appear even when obvious bias signals are removed. That’s the kind of failure mode institutions find hardest to detect—and hardest to undo later.
What the Researchers Actually Measured: REF proxy scores vs ChatGPT “quality” ranks
The core idea is simple: both systems aim to reflect research quality, but they do it differently.
ChatGPT scoring setup (what went into the model)
From a prior large-scale dataset (described in the study), ChatGPT provided a quality score for each REF 2021 journal article based on title + abstract. The scoring used a zero-shot prompt adapted from REF reviewer criteria—asking the model to assess originality, significance, and rigour, then assign a score on the REF 1* to 4* scale.
Crucially, because REF destroyed individual article scores, the study didn’t have access to exact “human expert” outcomes per article. Instead it used a proxy.
REF proxy quality measure (what the paper could compare against)
The researchers used an imperfect but available proxy: the average REF departmental score for the submitting department(s), derived from public departmental profiles.
That creates two issues the authors explicitly discuss:
1. Loss of individual variation: article-level differences get dampened because everyone is compared to a departmental average.
2. Department-level confounding: if some departments specialise in types of work that ChatGPT scores differently, gender-linked patterns could arise indirectly.
Why they compared ranks instead of raw scores
Even if both systems use the same nominal scale (1* to 4*), their score distributions differ. Prior work found ChatGPT scores have different means and smaller variance than human review.
So the paper’s primary comparison uses a rank-based method:
- rank articles within each UoA by ChatGPT score
- rank articles within each UoA by REF proxy score
- compute a rank gain: ChatGPT percentile rank - REF percentile rank
This matters because rank comparisons reduce distortions from “ChatGPT score scale vs REF proxy scale” mismatches.
Gender variable: “first author” (and why it’s not perfect)
The study uses first-author gender inferred from first names using a UK/USA gendered-name list (excluding ambiguous or gender-neutral names). They also run a solo-authored check where first author is unavoidably the only author—helping interpret gender roles more cleanly.
The main results: Where male-first-author papers got higher ChatGPT quality scores
The paper reports two related questions:
- RQ1: Do mean ChatGPT scores differ between female and male first-authored papers?
- RQ2: Does ChatGPT gain relative to REF differ between female and male first-authored papers?
Gender differences show up in many UoAs—but not uniformly
Across the 34 UoAs, the researchers found that for most fields, the direction of gender differences matched between:
- mean ChatGPT scores, and
- mean REF proxy scores
That’s not comforting: it suggests the overall system isn’t “flat” with respect to gender even in the comparison proxy. But the key twist is how ChatGPT differs from REF.
The study notes that the average female-minus-male difference was generally larger for ChatGPT than for REF—especially across several STEM-related UoAs. The authors are careful here: you can’t simply compare magnitudes between systems because the score distributions may differ.
Importantly: statistical clarity is limited in places
They also point out that in some UoAs where the direction varies, at least one of the two 95% confidence intervals contains 0, meaning there isn’t clear evidence for a consistent difference there.
So the pattern is best described as:
- often male-favourable
- stronger in certain UoAs
- usually small
- not universal
Rank-gain analysis: male-first-author papers gained more often (than female)
When looking at ChatGPT rank gain relative to REF proxy, male-first-author papers had higher mean ranks in 23 of 34 UoAs. Female-first-author papers had higher mean ranks in the remaining 11 UoAs.
Then comes the “are differences real?” layer:
- Mann–Whitney tests found statistically significant differences in 13 UoAs
- 11 were male-favourable
- only 2 were female-favourable: Clinical Medicine and Area Studies
The overall takeaway from the rank-gain results is: ChatGPT’s relative scoring advantage (over the REF proxy) is more often favourable to male first authors, particularly in health, science, and engineering-related areas.
Comparing ChatGPT vs REF proxy: How the gender pattern changes depending on the metric
To make the comparison concrete, here’s what the paper’s analysis implies in practice:
| What you compare | How it’s measured | Gender pattern frequency (overall) | Where it’s most noticeable |
|---|---|---|---|
| Mean ChatGPT scores vs female/male first author | Female minus male differences in mean ChatGPT scores by UoA | Direction often aligns with REF proxy direction; commonly male-favourable in many UoAs | STEM-adjacent areas; especially health, science, engineering-related UoAs |
| ChatGPT rank gain relative to REF proxy | ChatGPT percentile rank - REF percentile rank, averaged; tested with Mann–Whitney |
Male-favourable in 23/34 UoAs; significant male-favourable in 11/13 significant cases | Especially most STEM-related UoAs |
| Does “relative advantage” reproduce REF differences? | Check whether ChatGPT gender differences exceed REF proxy patterns | Often ChatGPT shows stronger gender separation than the proxy | Again, mainly in STEM-related UoAs |
And one more nuance the authors emphasise: some differences may be driven by the departmental averaging process used to compute the REF proxy. If departments differ systematically by field or topic—and those field patterns differ by author gender—then you can see “gender-associated” effects that aren’t direct gender bias.
Why it might be happening: Citations and abstract complexity (but not in a simple way)
A big part of this paper is asking: if ChatGPT is acting differently, what is it responding to? The authors test two plausible “indirect mechanisms.”
Mechanism 1: Citation impact differences by gender
The researchers check normalised citation impact (NLCS 2024). Across UoAs, the female-minus-male difference is negative in 21 of 34 UoAs and positive in 13. The average UoA-level difference is small and negative: about -0.023, with median around -0.019.
That implies: in many areas—again, often including health and STEM—male first-authored papers show slightly higher normalised citation impact.
But it doesn’t perfectly match the ChatGPT rank-gain pattern:
- some social science and humanities UoAs show higher citation impact for female first authors
- yet ChatGPT’s relative advantage in those areas is often still male-favourable or mixed
So citations may mimic part of the pattern, but not fully explain it. The authors also point out a possible methodological issue: the previous study that found a small female gain used article-level REF scores converted into bibliometric REF-like measures. This paper uses a departmental proxy, which could flip the observed direction after aggregation.
Mechanism 2: Abstract complexity (Flesch–Kincaid Grade)
Next: maybe ChatGPT rewards certain kinds of writing. The study tests abstract complexity using Flesch–Kincaid Grade:
- higher values = more complex / less readable
- abstract readability data came from a prior REF2021 abstract study
The result: the gender pattern for abstract complexity is mixed.
- female-minus-male differences are positive in 20 of 34 UoAs and negative in 14
- in panels where writing-style differences would help explain the ChatGPT result, the pattern still doesn’t neatly line up
Most importantly, the paper concludes abstract complexity is not a likely cause of the male-favourable ChatGPT pattern.
So if you were hoping the explanation would be “ChatGPT just likes more complex abstracts written by men,” the data doesn’t support that clean story.
The solo-author check: Reducing authorship-role ambiguity
Because first-author gender is messy in team science (other authors may contribute leadership in some fields), the authors run a solo-authored paper analysis. In solo authorship, first author = only author, so gender attribution is clearer.
A key result:
- In Panel C, solo-author results didn’t show a clear consistent male- or female-favourable pattern.
- In Panel D, the story shifts: in the first-author analysis, six of ten UoAs were female-favourable; in solo-author Panel D, nine of ten UoAs were female-favourable (with only English Language and Literature showing a tiny male-favourable difference).
However, none of the UoA-level differences in solo-author Panel D were statistically significant, and the authors caution that direction can be affected by sample sizes. Still, the solo-author pattern suggests: the first-author effects may be entangled with team authorship structure and field differences, not just “gender in the abstract.”
Key Takeaways
- Yes, there’s a pattern: In REF 2021 journal articles, male first-authored papers more often receive slightly higher ChatGPT research quality scores and more favourable rank gain relative to a REF proxy, especially in health, science, and engineering-related UoAs.
- Differences are usually small and not universal: Many UoAs show mixed results, and confidence intervals often allow the possibility of no difference in specific fields.
- Rank-based comparison matters: The study uses percentile ranks within each UoA to handle different score distributions between ChatGPT and the REF proxy.
- Abstract complexity doesn’t explain it: Gender differences in Flesch–Kincaid abstract complexity are not a good single explanation for the ChatGPT advantage pattern.
- Citations may partly overlap—but not fully: Normalised citation impact shows a slight male advantage in many UoAs, but it doesn’t perfectly match ChatGPT’s gender-related rank gains.
- Departmental REF proxy is a caution flag: Because REF individual article scores were unavailable, the comparison uses a department-level average proxy, which can introduce indirect effects tied to field and publishing structure.
- Solo-author results complicate the story: When authorship ambiguity is reduced (solo papers), the direction of gender-related effects in some panels shifts, suggesting the first-author findings may partly reflect team structures and field-specific authorship patterns.
If you’re using ChatGPT or similar LLMs in evaluation workflows, this paper supports a practical rule: don’t assume “no names in the prompt” means “no gender-related effects.” Instead, check for demographic-associated patterns—especially in field-sensitive STEM contexts—and treat AI scoring as a tool to support, not replace, careful human oversight.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- Does ChatGPT score research quality differently by gender? — arXiv
- Authors: Authors: Kayvan Kousha, Mike Thelwall