Generative AI in College Isn’t Inflating Grades (At Least Not Yet)

Students worry GenAI can replace learning—and colleges worry about grade inflation. New research on a large U.S. university finds no significant differential grade or satisfaction effects from ChatGPT availability (after COVID is modeled), tempering the substitution alarm.
The finding ChatGPT availability produced no significant differential grade inflation or satisfaction harm once COVID-19 effects were modeled.
The method The paper estimates “GenAI susceptibility” from syllabi assessment types and uses a differences-in-differences approach across pre- and post-ChatGPT periods.
The nuance It shows interest effects only under an assumption about transient COVID impacts, so results are not uniformly present across satisfaction measures.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

Generative AI availability via ChatGPT did not show a significant differential effect on grades or satisfaction-related course evaluations in the study after accounting for COVID-19 confounds. The results temper—but do not dismiss—the substitution-inflation worry.

For practitioners, this suggests colleges shouldn’t assume blanket grade inflation is already occurring; assessment and policy work can focus more directly on preserving learning signals and integrity rather than reacting only to fear of automatic score jumps.

A key caveat is that the study’s conclusions depend on its setting, time window, and how it measures “GenAI susceptibility” from syllabi and models COVID effects; other institutions or different assessment designs could still show different patterns.

Generative AI in College Isn’t Inflating Grades (At Least Not Yet)

Introduction: The “GenAI substitution” worry is real—but new evidence is cooler than you might expect

Students are using generative AI (GenAI) in higher education at massive scale, and that’s naturally triggering a serious question: are students using AI to boost grades without learning? A new research study (based on https://arxiv.org/abs/2607.21534) tackles this head-on by looking at how grades and student satisfaction changed after the public release of ChatGPT in November 2022.

The core concern behind this research is often called the “GenAI substitution hypothesis.” The idea is straightforward: if GenAI can write essays or solve homework in a way that’s hard to distinguish from student work, then students might “substitute” AI for their own cognitive effort—especially in courses where a lot of the grade comes from take-home assignments, essays, and other work that isn’t watched in real time. If that’s happening, you’d expect bigger grade jumps in more AI-susceptible courses, and possibly a drop in satisfaction (like feeling less understanding or less interest).

Here’s the surprising part: using administrative data from a large U.S. public university (2015–2025), the authors find no significant differential effect of ChatGPT availability on grades overall, or on satisfaction-related course evaluations, once they account for the huge confounding shock of COVID-19. They also find no evidence that grade effects disproportionately help lower-performing students. In short, this is new research that tempers—not dismisses—the most alarming version of the substitution story.

Why This Matters: Colleges need answers now, not after the next wave of policy headlines

This research is significant right now because universities are still deciding what to do in the middle of an ongoing behavior shift. There’s a temptation to jump straight to one of two extremes: either “AI breaks everything” or “AI is just another tool.” But real decision-making needs something more precise: Do grades stop reflecting learning? And if they do, in which kinds of courses, and for which students?

A concrete scenario: imagine a department that’s redesigning assessments next semester. They might wonder whether to move more weight to in-class exams, redesign take-homes to require process evidence, or create AI-transparent policies. If GenAI substitution were strongly happening, you’d expect take-home-heavy sections to show a measurable jump in grades after ChatGPT—and students might report feeling less engaged or less understanding. This study suggests the answer is not “automatically yes,” at least in the setting and time window examined. That means assessment redesign can focus on learning quality and integrity, but doesn’t have to be driven solely by fear of blanket grade inflation.

This also builds on earlier AI-in-education findings, including quasi-experimental studies that did report grade increases in more AI-compatible courses. What’s notable here is that the authors use a careful course exposure measure from syllabi, anchor susceptibility to a pre-ChatGPT, pre-COVID baseline, and explicitly model COVID-19 effects as either transient or persistent. In other words: they aren’t just asking “did grades change?” They’re asking “did grades change differentially in AI-susceptible courses, after separating out COVID?” That’s the kind of design that changes how seriously you should take either alarmist or dismissive narratives.

How the study measures “GenAI susceptibility” without guessing: Syllabi → assessment types → grade weight

A big reason this question is hard is measurement. “AI availability” is vague. Some students use GenAI in every course; some instructors allow it; some assessments are more automatable than others. So this paper approaches the problem by measuring course susceptibility—how easily a student could use GenAI to complete graded work without being directly observed.

What counts as “GenAI-susceptible” vs “non-susceptible” work?

The authors classify assessment types based on whether they require independent completion without instructor observation and could plausibly be produced using GenAI (at least in principle). In their setup:

  • Susceptible assessments include things like:

    • open-book / open-notes or take-home exams
    • homework
    • papers
    • projects
  • Less susceptible assessments include things that require live human performance or proctoring, such as:

    • in-person closed-book exams
    • presentations
    • live skills demonstrations

They also handle ambiguous cases (like “unknown exams”) and special cases (like projects with both presentation and non-presentation components) by splitting weight where needed. The outcome of all this is a single number per course offering: the share of the final grade determined by susceptible assessments, which they call Susceptibility (a value from 0 to 1).

The scale of the dataset: big enough to matter

The research uses:
- 156,135 unique students
- 87,936 course offerings
- across 2015–2025
- from a university with extensive syllabus and administrative records

Even more impressive: the authors scrape and process 36,357 syllabi for offerings between 2015 and 2025 (out of 44,876 syllabi total scraped), excluding a small number with sensitive content. They also use course evaluation data (available as digital records starting in 2016) to measure student satisfaction proxies.

The LLM pipeline: validated against human judgment

Syllabi are messy PDFs full of text, so the authors use a language-model pipeline to reconstruct:
- what assessment categories exist, and
- how much each category is weighted in the final grade.

To make this credible, they don’t just eyeball results. They validate the pipeline against 525 human-consensus-labeled syllabi, with two trained coders per syllabus and consensus resolution. For the continuous Susceptibility measure, they report mean absolute error of 0.063 on the 0–1 scale. That’s not perfect—but it’s accurate enough to study patterns across tens of thousands of course offerings.

If you want to trust the conclusions, you need to trust the measurement—and this is one of the study’s strongest methodological efforts.

What the researchers actually test: Did ChatGPT change outcomes more in susceptible courses?

Once they’ve measured Susceptibility, the key question becomes: What happened after ChatGPT, and did it happen disproportionately in courses that were more susceptible?

The policy shock: ChatGPT’s public release in Nov 2022

The authors treat ChatGPT’s public release as a major unexpected shock and use it as the time break.

Then they run a difference-in-differences style analysis comparing outcomes in:
- more susceptible vs less susceptible courses
- before vs after the ChatGPT release

The design has to wrestle with COVID-19

COVID-19 is a huge problem in education research because it didn’t politely “finish” before ChatGPT arrived. The paper therefore uses a dual-shock framework:
- one shock for COVID-era disruption (2010s turns into 2020s chaos)
- another shock for ChatGPT availability

They test outcomes under two assumptions about COVID:
1. COVID effects are transient (they fade after 2022)
2. COVID effects are persistent (they carry into the post-ChatGPT period)

This lets the authors give a bounded estimate: you can’t know the truth of COVID’s lingering impact, but you can avoid pretending it didn’t exist.

Why they anchor Susceptibility to 2019 (this matters a lot)

A clever detail: they anchor each course’s Susceptibility using the 2019 offering rather than measuring it repeatedly each year.

Their reasoning is practical:
- If instructors change grading policies after ChatGPT arrives, then a time-varying measure could accidentally “bake in” the response to AI—making the effect harder to interpret.
- Anchoring to a clean pre-AI, pre-COVID-ish baseline also reduces contamination.

This choice ends up being central to whether positive “AI inflation” effects show up or vanish in robustness checks.

The results: Grades, withdrawals, failures, and satisfaction don’t show substitution-driven jumps

The paper’s headline is basically: once COVID is separated out, ChatGPT doesn’t appear to inflate grades or satisfaction in susceptible courses.

Average grades: tiny and statistically insignificant (under the conservative COVID model)

The main grade outcome is a term-level final grade mapped to GPA-like values (A+ up to E with A+ = 4.0 and E = 0.0). They then estimate the differential effect of Susceptibility after ChatGPT.

Under their preferred approach (dual-shock design, with a conservative interpretation—details depend on the COVID assumption), the effect is:

  • Comparing fully susceptible (Susceptibility=1) vs fully non-susceptible (Susceptibility=0), the post-ChatGPT minus pre-COVID difference is 0.03 grade points (SE = 0.052, not significant).
  • If COVID effects are treated as persistent, they estimate a difference of 0.045 grade points (SE = 0.026, p < 0.1).

So the estimated range is [0.030, 0.045] grade points, and it’s economically small and not statistically significant at the 5% level.

They also report that formal “parallel pre-trends” tests for grades fail in parts of the analysis (meaning you have to be cautious about strict causal interpretation of the null). Still, the authors argue the null is supported by:
- the overall flatness visually in event-study plots, and
- heterogeneity analyses that don’t show where effects would concentrate if substitution were happening.

Grade distribution shifts: at most a shaky “more A’s” story, but the floor (failures) doesn’t rise

If GenAI substitution were driving grade inflation, you’d expect it to show up in the grade distribution—especially near failure thresholds.

They check thresholds like:
- probability of at least an A, B, C, D, and
- probability of being below D (non-passing)

Their results are:
- No robust evidence of grade increases at the bottom.
- There’s some descriptive evidence of an increase in the share of A grades after ChatGPT, but the authors treat it as unstable because parallel pre-trends fail for the A threshold.

Crucially:
- Passing margins and failure rates don’t meaningfully change in the conservative estimates.

Withdrawals and failures: no differential changes after ChatGPT

They examine:
- withdrawal rate and
- failure rate (defined as not receiving D- or higher)

Findings:
- Withdrawal changes are small and insignificant (conservative estimate 0.010 with SE = 0.006, not significant).
- Failure effects are also null (conservative estimate about −0.002, SE = 0.002, not significant).

This is important because one of the clearest substitution-alarm signatures would be fewer failures and fewer withdrawals in susceptible courses.

Do weaker students benefit more? The study finds no “help for the bottom” pattern

A common worry (and one suggested by some work in workplaces) is that GenAI might especially help students who previously struggled.

The authors build a prior performance proxy (AbilityRank) based on first-term GPA rank within cohorts, residualized against course means to reduce confounding from course difficulty/selection.

Then they test terciles (bottom/middle/top) separately.

Result:
- GenAI availability shows no statistically significant grade effect within any tercile.
- The point estimates are small and flat across prior ability.
- There’s no sign that the bottom of the distribution disproportionately benefits.

Course evaluations (“satisfaction” proxies): no robust changes once COVID is handled carefully

The authors use course evaluations to proxy satisfaction and engagement. Specifically, they analyze median responses for:
- self-reported understanding (understand)
- subject interest (interest)
- relative workload (workload) — with coding adjusted so interpretation is consistent

Their key message:
- Under the conservative specification (treating COVID as persistent), they find no significant changes in understanding, interest, or relative workload in more susceptible courses.

They do find something that looks like:
- interest increases and relative workload decreases
…but that only appears under the transient COVID assumption, and the authors present it as not robust.

So overall: no strong evidence that GenAI reduces satisfaction (at least not in the aggregate outcomes available here).

How this study compares to earlier work: why previous findings might differ

Two earlier quasi-experimental studies are mentioned in the paper—both finding evidence consistent with substitution (grades rising more in AI-compatible/susceptible courses). This paper disagrees on the main causal outcome, and the differences are plausibly explained by measurement and modeling choices.

Here’s a simplified comparison of the broad approaches, based on how the authors describe limitations and improvements:

Study (as discussed in this paper) Exposure / susceptibility measure Key design choice highlighted Main conclusion (relative to substitution)
Hausman et al. (Israeli university) “AI-compatible” defined so much of final grade comes from take-home-type assessments Uses a model that (per this paper) may allow courses to switch treatment groups and misses some fixed effects structure Finds more susceptible courses get higher grades
Chirikov (public U.S. university) Task composition via writing/coding task shares from syllabi Restricts to fall offerings with courses appearing every year; focuses on certain subsets; measurement validated mostly by spot checks Finds grade increases in exposed courses
This study (large Midwest U.S. university) LLM-reconstructed grade-weighted susceptibility from syllabi, anchored to 2019 Uses a dual-shock approach for COVID vs ChatGPT and validates LLM measurement against 525 human-labeled syllabi Finds no robust differential effects on grades or satisfaction after accounting for COVID

The big takeaway isn’t “previous studies were wrong.” It’s that substitution isn’t showing up consistently across settings and designs, and that the treatment definition and COVID modeling can flip the story.

Practical implications: What should universities do if the worst-case grade inflation signal isn’t showing up (yet)?

This doesn’t mean universities should ignore assessment integrity. It means you might need a more targeted response than “swap everything to proctored exams immediately.”

1) Keep monitoring the “grade signal,” not just cheating headlines

The study suggests that in this university, during these early post-ChatGPT years, grades didn’t inflate more in susceptible courses after COVID effects are separated out. That’s reassuring—but it’s not a permanent law of nature.

Implementation idea:
- track outcome changes by course assessment structure over time,
- watch for shifts as GenAI capabilities evolve and as students’ norms mature.

2) If you redesign assessments, anchor changes to learning goals—not panic

For courses heavy on essays and take-homes, instructors can still redesign to protect learning:
- require drafts or structured outlines,
- include in-course oral checks for key understanding,
- design prompts that reflect class-specific material or iterative thinking.

And importantly: if grades aren’t inflating already (in aggregate), these redesigns are still justified as learning improvements, not just compliance theater.

3) Don’t assume “lower-performing students get saved” automatically

Because this study finds no differential grade lift by prior ability, it challenges the simple “GenAI levels the playing field” narrative. That doesn’t mean GenAI can’t help. It means this university’s grades didn’t show equalizing effects during this period.

So support programs (tutoring, study guidance, AI literacy) may matter more than simply allowing tools.

Key Takeaways

  • Main finding: After accounting for COVID-19 disruption, the introduction of ChatGPT shows no robust differential effect on final grades in courses with higher Susceptibility to GenAI (estimated range about 0.030 to 0.045 grade points, not significant at 5%).
  • No evidence of a grade “floor” effect: Withdrawal and failure rates also show null differential changes in more susceptible courses.
  • No “help for weaker students” pattern: Grade effects do not concentrate among students with lower prior academic performance (AbilityRank terciles are flat).
  • Satisfaction proxies don’t drop: Course evaluation medians for understanding, interest, and relative workload show no robust negative movement once COVID modeling is treated conservatively.
  • A potentially important nuance: Some earlier signals like “more A’s” appear descriptive and unstable (pre-trend issues), not cleanly causal.
  • What this means for course policy now: You can be more confident that, at least in this setting and timeframe, GenAI hasn’t automatically destroyed the grade signal. But continued monitoring and learning-centered assessment design still matter.
  • What to watch next: As GenAI becomes more embedded and as students/instructors adjust policies, the balance could change—so institutions should measure outcomes over time, especially for assignment structures most vulnerable to offloading.

If you’d like, I can also turn this into a “what course designers should change next semester” checklist based strictly on the assessment categories the paper treats as GenAI-susceptible.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Generative AI Evidence Half-Lives: When “Current” Claims Drift

Title: Hallucination as a Feature: How Creative Thinking Shapes Generative AI

Hidden Links in Synthetic Data: A Subtle Poison That Tricks Generative Models

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.