The Short Answer
Students generally accept ChatGPT’s feedback for revision, but they often reject AI as the final grading authority when they’re told it produced the score. The key finding is that usefulness and grading legitimacy are judged separately.
Practically, you can use transparent AI feedback to help students improve writing, but you should explicitly preserve human instructor authority over evaluative decisions—ideally with student reflection after grading.
A major nuance is that disclosure doesn’t automatically increase trust in AI grading; it shifts students’ reasoning toward fairness, limits, and the role of instructors, so the assessment design must address authority.
On this page
- Introduction: The real question isn’t “Is AI feedback good?”—it’s “Who gets to decide?”
- Why This Matters (Right Now): Transparent AI testing classroom dynamics—not just writing
- How the Study Was Set Up: From handwritten work to AI scores—then reflection
- What Students Actually Said About AI Feedback: Useful comments, but not full authority
- Feedback Utility vs Evaluative Authority: The split students keep making (shown clearly in the data)
- What “Transparency” Actually Does: It pushes students into reflection about ethics and legitimacy
- The Paper’s 3-Principle Ethical Framework: Transparency, human mediation, reflective practice
- Key Takeaways
- Key Takeaways
AI Grades Feel Different When Students Know It’s ChatGPT
New research (from arXiv:2609.05346) digs into a surprising question: what happens to student trust and thinking when they’re told a machine (specifically ChatGPT) graded their writing?
Introduction: The real question isn’t “Is AI feedback good?”—it’s “Who gets to decide?”
If you’ve used AI in education—even lightly—you’ve probably noticed the same pattern: students will often say the feedback is useful… but still push back on the idea of AI being the “final judge.” That tension is exactly what this new research explores.
In the study behind arXiv:2609.05346, Rayed AlGhamdi looks at how undergraduate students in a technical communication course interpret transparent AI-assisted writing assessment—meaning: students are explicitly told that ChatGPT generated both the score and the feedback. This matters because most previous work either focuses on outcomes (did grades improve?) or on “blinded” conditions (students didn’t know where the feedback came from). Here, the source is not hidden.
So the study doesn’t just ask whether AI can produce helpful comments. It asks: once students know the evaluator is AI, how do they separate helpful feedback from legitimate grading authority—and what role do they expect the human instructor to play?
The answer, based on qualitative reflections from 13 computing students, is nuanced. Students generally accept AI feedback for surface-level revision (grammar, organization, clarity), but they repeatedly insist that human instructors should retain evaluative authority. The research also proposes a practical ethical framework built on transparency, human mediation, and reflective practice.
Why This Matters (Right Now): Transparent AI testing classroom dynamics—not just writing
Right now, universities are stuck between two bad options: banning AI (usually unrealistic) or letting it silently influence assessment (often ethically awkward). This research points to a third path: use AI openly, but design the assessment conversation so students can reason about authority.
Here’s the real-world scenario where this matters today: imagine a department rolling out AI-assisted feedback for large writing-heavy courses (technical writing, lab reports, proposals). If students receive AI-generated scores without being told, you might get short-term satisfaction (“the comments are clear!”) but long-term problems: mistrust, academic integrity disputes, and students gaming the system because they don’t understand how decisions are made.
On the other hand, if you disclose AI involvement and ask students to reflect after grading, the study suggests something more constructive happens. Students don’t automatically reject AI; instead, they become more thoughtful about limits, fairness, and why instructors matter. That’s a big shift from treating AI like a hidden grading tool and it builds on previous trust-and-feedback research by showing that source transparency changes the kind of thinking students do—not just what they think of the feedback.
Finally, this research builds on earlier AI feedback studies by zooming in on a distinction that’s often blurred: many studies talk about “trust in AI feedback,” but AlGhamdi’s data suggests students are actually making two separate judgments:
1) Is this feedback useful?
2) Should this system be allowed to grade me?
That split is the heart of what’s actionable right now.
How the Study Was Set Up: From handwritten work to AI scores—then reflection
This wasn’t a lab-only experiment where students just clicked buttons and moved on. It was embedded in a real undergraduate course, which is important because students were responding in an authentic educational context.
What students did (the course sequence)
- Students completed an in-class handwritten writing task (13 of 19 students submitted reflections usable for analysis).
- Submissions were scanned and evaluated by
ChatGPTusing a rubric-based prompt aligned with course objectives. - Students were then explicitly informed that
ChatGPTgenerated both the feedback and the score. - After that, students wrote a short pen-and-paper reflection answering three questions:
- Whether they agreed/disagreed with the feedback and score (and why)
- What they learned/noticed about their writing from the evaluation
- Their opinion about using
ChatGPTto evaluate writing and provide feedback
Who participated?
- Sample: 13 male undergraduate computing students
- Course context: an undergraduate technical communication course at a Saudi public university
- Task nature: handwritten in-class responses (used to preserve authorship integrity—since students could otherwise use GenAI before submitting)
Why handwritten writing matters here
The study deliberately used handwritten, in-class writing to avoid a common problem in AI assessment research: you don’t know whether the submitted text is student-generated or AI-assisted. In this design, the authenticity of the assessed writing was clearer, which lets the analysis focus on perceptions of AI grading itself.
For readers interested in the original details, this overall method and design are described in the paper at arXiv:2609.05346.
What Students Actually Said About AI Feedback: Useful comments, but not full authority
The findings come from inductive thematic analysis of the students’ reflections. Four recurring themes appeared:
- Perceived usefulness of feedback
- Awareness of AI’s contextual and pedagogical limitations
- Conditional trust (feedback utility ≠ evaluative authority)
- Reflection on the institutional/pedagogical role of the human instructor
Theme 1: Students found AI feedback genuinely helpful
All 13 participants described the ChatGPT feedback as useful—especially because it pointed to specific issues (like grammar or disorganized paragraphs) instead of vague advice like “needs improvement.”
They repeatedly described the feedback as:
- “clear”
- “direct”
- “well-organized”
- “objective”
And some directly connected the feedback to learning outcomes, mentioning it helped them:
- understand mistakes (especially grammar/sentence structure)
- organize their writing better
- gain confidence about their writing
A key nuance: the value students reported wasn’t just emotional reassurance. It was practical—the feedback seemed actionable, mainly for surface-level revision tasks.
Theme 2: Students noticed limitations quickly (especially interpretation)
Even though feedback was useful, students didn’t treat AI as infallible. They raised limitations in two broad categories:
Technical limitations
- ChatGPT might misread handwriting when the submission is scanned.
- One student explicitly noted spelling mistakes the AI attributed to the student, suggesting the model might be reacting to handwriting rather than actual text.
Contextual/pedagogical limitations
- The AI didn’t “know” the institution’s grading norms in the way the instructor does.
- Students also noticed that AI feedback can be overly positive, even if it’s still somewhat usable.
This is where the transparency seems to matter: when students knew AI was evaluating them, they didn’t just accept the output—they evaluated the conditions under which the output could be wrong.
Theme 3: Students separated “usefulness” from “authority” (this is the headline insight)
This is the most conceptually important theme in the paper.
Students expressed a repeated pattern that sounds simple but is actually powerful:
“The feedback is useful, but the grading authority should stay human.”
They treated these as separate judgments, not opposite ends of one trust spectrum. So even students who agreed with the content of the feedback could still insist the score should be checked by the instructor.
The paper describes this as a distinction between:
- feedback utility (Is this helpful?)
- evaluative authority (Should this system be allowed to decide my grade?)
Theme 4: The instructor’s role is legitimate, dialogic, and institutionally necessary
Finally, many students broadened their reflections beyond AI’s accuracy and into what education means.
They argued that human instructors matter because they:
- know individual students better
- can explain concepts in personally appropriate ways
- can treat grading as a dialogical process (not a one-way machine verdict)
- represent institutional purpose (“If AI grades everything, why do we need university?”)
- serve as a fairness safeguard—especially because AI errors could affect grades
In other words: students didn’t reject AI. They defended the human reason-for-existence of assessment.
Feedback Utility vs Evaluative Authority: The split students keep making (shown clearly in the data)
A helpful way to see this is as a “same event, different question” problem.
The research suggests students didn’t just ask, “Is AI right?” They asked at least two different questions—sometimes in the same sentence.
Here’s a comparison of what students accepted versus what they resisted:
| Aspect of AI in assessment | What students generally liked | What students generally rejected |
|---|---|---|
ChatGPT feedback content |
Clear, specific guidance; helpful for grammar/organization/clarity | Not perfect; can misread handwriting; may miss context |
ChatGPT grading authority |
Useful as “initial feedback” or a learning support tool | Should not be the final decision-maker on grades |
| Need for human instructor | Instructor can validate, contextualize, and safeguard fairness | AI alone can’t provide accountability or dialogic teaching |
This distinction helps reconcile why students in AI assessment studies can simultaneously sound both positive and skeptical. They’re not contradicting themselves—they’re partitioning the evaluation system into different roles.
This is also where the paper’s comparison to the author’s earlier blinded study becomes interesting: in a previous version of this same course work (students were unaware AI produced feedback), students’ reflections focused more on feedback content rather than authority. In the transparent condition, students talked more about the legitimacy of who gets to grade.
What “Transparency” Actually Does: It pushes students into reflection about ethics and legitimacy
Transparency here isn’t just a compliance checkbox. In this study, telling students “ChatGPT graded your work” changed the nature of their thinking.
Transparency shifted the focus after the fact
Most writing-assessment discussions emphasize revision during drafting. This study placed reflection after students received feedback and scores. That timing mattered. Students began asking systemic questions like:
- “Who should decide my grade?”
- “What role should AI play in assessment?”
- “What happens if AI makes mistakes that affect grades?”
So transparency plus post-assessment reflection appears to redirect attention from “How do I fix my paragraph?” to “How does evaluation work, and is it fair?”
A simple classroom practice emerges from this
One practical implication is almost embarrassingly easy to implement: a short reflection prompt after AI-based scoring.
Students can be asked to write (briefly) about:
- whether they agree with AI feedback,
- what limits they think exist,
- and whether they accept AI as an evaluator.
And crucially, this doesn’t require extra technology. The paper notes the reflective format is low-cost and replicable.
The Paper’s 3-Principle Ethical Framework: Transparency, human mediation, reflective practice
Based on these themes, the paper proposes a framework for ethical GenAI integration in writing assessment with three principles:
Transparency
Tell students clearly that AI generated the feedback and score.Human mediation
Keep instructors responsible for evaluative decisions (scores/authority), using AI as a support tool.Reflective practice
Prompt students to reflect critically, especially after receiving AI-mediated evaluation.
If you’re designing an AI-assisted writing workflow, this framework has practical bite:
- Transparency helps students become critical rather than passive.
- Human mediation preserves accountability and fairness.
- Reflection turns AI output into a learning opportunity about assessment itself, not just writing mechanics.
And the paper ties this back naturally to its findings from arXiv:2609.05346: students accepted AI as useful feedback while consistently positioning the human instructor as the appropriate authority over grading decisions.
Key Takeaways
Key Takeaways
- AI feedback can be perceived as clearly useful: in this study, all 13 students described
ChatGPTfeedback as helpful, especially for surface-level writing issues like grammar and organization. - Students actively noticed AI limitations: they pointed out technical risks (e.g., misreading handwriting from scans) and contextual gaps (e.g., not matching institutional grading norms).
- The biggest insight is the split between roles: students separated feedback utility (usefulness) from evaluative authority (who should grade). They could like the feedback while rejecting AI as the final grader.
- Transparency changes student thinking: when students knew AI generated the score, their reflections extended beyond text quality to questions of trust, fairness, and the legitimacy of assessment.
- Human instructor role stays central: students repeatedly argued that instructors matter because they understand students better, enable dialogic teaching, and safeguard fairness when AI could err.
- Practical implementation idea: if you use AI-generated feedback in a course, consider pairing it with a short post-assessment reflection prompt to guide students toward critical, ethical engagement.
- Future direction: this study involved 13 male students in one course context, so broader research is needed—but the framework (transparency + human mediation + reflective practice) is immediately actionable.
If you want, I can also turn the paper’s proposed 3-principle framework into a ready-to-use policy checklist for instructors (what to disclose, what the human does, and what students reflect on).
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education — arXiv
- Authors: Authors: Rayed AlGhamdi