GPT-assisted Writing Detection with Interpretable Stylometry: A New Step in Academic Integrity

Detect GPT-assisted writing with evidence that’s visible in the submitted text. New research uses interpretable stylometric features and SHAP explanations—without “black-box” detectors—then tests generalization to unseen students.
The finding Interpretable stylometric signals can detect GPT-assisted writing with ROC-AUC 0.87 and F1 0.842 on a held-out test set.
The method A Random Forest classifier uses eight text-derived stylometric features, explained post-hoc with SHAP to show which lexical and grammatical traits matter.
The caveat The approach has measurable error rates, so high scores should inform human review rather than serve as definitive evidence of wrongdoing.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

GPT-assisted writing can be detected from the submitted text using interpretable stylometric features, reaching held-out ROC-AUC 0.87 and F1 0.842 with a Random Forest classifier. SHAP shows lexical and grammatical characteristics drive predictions.

For instructors, this enables a decision-support workflow: when a submission scores high, reviewers can inspect the specific text features (e.g., lexical richness or grammatical ratios) rather than rely on a black-box detector.

It’s not standalone proof of misconduct—false positives and false negatives remain—so results should be treated as evidence to review alongside context such as drafts or writing history.

GPT-assisted Writing Detection with Interpretable Stylometry: A New Step in Academic Integrity
Detecting AI writing using human-readable text signals

Introduction

If you’ve ever tried to figure out whether an essay is truly a student’s work—or subtly “borrowed” help from ChatGPT—you’re not alone. The hard part isn’t just spotting AI use. It’s getting evidence that’s both accurate enough to be useful and transparent enough to be defensible when grades and discipline are on the line.

New research from Kumar, Siddiqui, and Fuchsberger tackles exactly this problem: detecting GPT-assisted writing using interpretable stylometric features pulled straight from the final submitted text. Instead of relying on black-box detectors or proprietary model internals, the study asks a simpler, more teacher-friendly question: Can we detect AI-assisted writing using measurable writing-style signals that are directly visible in the text itself? And importantly, does it work on writing from new students it hasn’t seen before?

The results are promising. Using data from 9,090 participants—each providing both independently authored and ChatGPT-assisted responses (with paraphrasing instructions)—the best model (a Random Forest) achieved a held-out ROC-AUC of 0.87 and an F1-score of 0.842. It also provides post-hoc explanations via SHAP, pointing to lexical and grammatical characteristics as the strongest drivers. The catch? The false-flag risk isn’t zero, and the authors are clear that this shouldn’t be used as standalone proof of misconduct.

Why This Matters

This research is significant right now because institutions are stuck in a tough spot: they need scalable detection (for large classes, lots of submissions), but they also need procedurally fair evidence. Many current approaches are either (a) opaque (can’t explain why), (b) brittle (break under paraphrasing or stylistic variation), or (c) require access to generation signals (which you don’t have once a student submits an essay).

What makes this work feel “actionable” is the format of the evidence. Instead of saying “my model thinks this is AI-written,” it focuses on things instructors can actually inspect in principle: lexical diversity, vocabulary richness, the rate of rare words, and grammatical ratios. That doesn’t automatically make the decision fair—but it moves you from “trust the detector” toward “review the linguistic signals.”

A concrete scenario where this could be applied today: imagine a university’s writing center or academic integrity team running a decision-support workflow. When a submission triggers a high suspicion score, the reviewer isn’t stuck with a single opaque number—they can also see which text features drove the score (for instance, unusual patterns in Hapax Ratio, noun/adjective/verb ratios, or word entropy). This gives the reviewer something to cross-check against context: the student’s prior drafts, known writing proficiency, revision history (if available), and the plausibility of stylistic shifts.

How does this build on previous AI detection research? Earlier methods often use statistical patterns learned from language models, generation-probability tricks, or watermark-style signatures. The study you’re reading explicitly positions itself as a more text-intrinsic, transparent alternative, extending prior stylometry-style ideas by (1) using a sliding-window approach, (2) evaluating on unseen participants with a strict “no data leakage” user split, and (3) applying SHAP to interpret feature contributions. It’s not trying to replace instructor judgment—it’s trying to give it a sturdier footing than “trust the black box.”

What the Researchers Actually Measured: Interpretable Stylometric Features

The core idea is straightforward: even if GPT can mimic writing style, it may still leave patterns in how language is used. Stylometry captures those patterns without caring about the essay’s topic.

The study uses nine stylometric features computed from each text window, including:
- NN: total number of tokens
- |V|: number of unique tokens (vocabulary size)
- Hapax-based measures (including Hapax Ratio): how often rare words appear
- Type-Token Ratio (TTR): lexical diversity proxy
- Word Entropy: information-theoretic measure of token-frequency dispersion
- Non-stopword Ratio (NSR): proportion of content words
- Grammatical ratios based on part-of-speech tags (noun/verb/adjective/adverb ratios)
- Sentence-Length Variability: variation in sentence lengths
- Plus supporting sentence stats (mean and variability derived from sentence token lengths)

Why sliding windows matter (and why it’s not just “more data”)

Instead of computing features on the whole document, the authors use a sliding window protocol: 250 words per window with a 125-word stride (so 50% overlap). That gives multiple samples per document and helps reduce issues caused by variable length.

Here’s the analogy: think of the essay as a long road trip. If you only measure the whole trip’s distance, a detour early or late might be hidden. Windowing gives you several “checkpoints” along the route, so the model learns whether GPT-like signals appear consistently—not just in one region.

But there’s an important statistical caveat: windows from the same participant aren’t independent samples. The authors respect this by keeping all windows for a participant inside the same training/validation/test partition.

GPT-assisted here means paraphrased GPT output—not copy-paste

The dataset comes from the IIITD-BU (Paraphrased) collection described in the authors’ cited work (Mehta et al., 2026). Participants responded to prompts in two sessions:
1. Independent writing (no LLM use)
2. Assisted writing where students could use LLMs like ChatGPT, but were required to paraphrase what the model produced before typing

Copy-and-paste and grammar tools were disabled in that original setup. So in this study, “GPT-assisted” means: the student typed content that originated from GPT but was paraphrased manually.

That restriction matters because it makes the detection task more realistic: paraphrasing is exactly what people attempt to do to “remove” obvious AI fingerprints.

How the Study Was Set Up: Preventing Leaky Detection

Detection studies are notoriously easy to accidentally “cheat” on. A classic problem is data leakage: if the same person appears in both training and testing, the detector may learn that person’s idiosyncratic style rather than GPT-related signals.

This paper goes out of its way to avoid that.

The dataset and the user-split protocol

The dataset contains 180 samples per participant condition, but more precisely, the study builds a balanced set where each of 90 participants contributes:
- one independently authored response and
- one GPT-assisted response

That’s the raw “pairing,” but the study’s crucial number for ML evaluation is the participant count in the split.

They use 9,090 participants total for evaluation (as described in the abstract and methodology framing), partitioned into:
- 7,272 training participants (144,144 samples/windows)
- 1,818 held-out participants (3,636 samples/windows: 1,818 independent + 1,818 GPT-assisted)

All windows from a participant stay in the same split.

Windowing turned documents into many samples

After cleaning the final text, each document is split into windows:
- Most documents longer than 250 words produced 250-word windows
- Documents shorter than 250 words became a single full-document window

Across the dataset, the study reports 722 windows, including:
- 397 independently authored windows
- 325 GPT-assisted windows

(And for the held-out participants, they describe 127 windows across 3,636 documents—the paper provides these counts as part of the window-level dataset framing.)

Model outputs: window probabilities → document median

Each classifier outputs a probability for each window being GPT-assisted. To convert those window scores into a document decision, the authors aggregate using a median-probability rule:

  • Document score = median of window probabilities
  • Classify as GPT-assisted if the median probability ≥ 0.50

Median aggregation is a smart choice here: it reduces the impact of a single “weird” window that might look unusually GPT-like (or unusually human).

Which Classifiers Worked Best—and What the Numbers Say

The paper compares eight machine learning classifiers trained on the nine stylometric features:
- Logistic Regression
- Linear Discriminant Analysis (LDA)
- Quadratic Discriminant Analysis (QDA)
- Naive Bayes
- Random Forest
- Linear Support Vector Machine (linear SVM)
- K-Nearest Neighbors (kNN)
- Multilayer Perceptron (MLP)

They tune hyperparameters using Bayesian optimization on the training partition only, with careful cross-validation that keeps participant data intact.

Performance comparison (validation and test)

Validation showed multiple methods performing in the same general ballpark, but Random Forest came out on top and also ranked first in an additional nested-validation check.

Then comes the real test: the held-out evaluation on unseen participants.

Metric (held-out test set, 36 documents) Best model (Random Forest)
ROC-AUC 0.870
F1-score 0.842
False Positive Rate (independent flagged as GPT-assisted) 22.2%
False Negative Rate (GPT-assisted missed as independent) 11.1%

A key point: the paper’s held-out evaluation described 36 documents coming from 18 held-out participants (so each side has 18 documents: independently authored and GPT-assisted). That’s a small test set, and the paper reports wide confidence intervals, meaning uncertainty is non-trivial.

What those error rates mean in real academic settings

They treat GPT-assisted as the positive class. On the held-out set:
- 4 independent documents were incorrectly classified as GPT-assisted
→ False Positive Rate = 22.2%
- 2 GPT-assisted documents were incorrectly classified as independent
→ False Negative Rate = 11.1%

False positives are the big concern in academic integrity because they can lead to unfair accusations. Even if the model is fairly strong, a ~1 in 5 false-positive rate on this small held-out set is too risky for anything but decision support.

The authors explicitly position the system as a tool for contextual review, not a verdict machine.

What the Model Looked At: SHAP Explanations in Plain English

Here’s where the paper gets especially interesting: it doesn’t just report performance—it investigates which features were driving the Random Forest predictions.

Using SHAP analysis, they identify the most influential stylometric features globally:
- Hapax Ratio (most influential)
- NSR (Non-stopword Ratio)
- Noun Ratio
- Adverb Ratio
- TTR (Type-Token Ratio)
- Smaller contributions from: Verb Ratio, Sentence-Length Variability, Word Entropy
- Adjective Ratio contributes least

Direction of effects: which writing style signals match GPT-assisted text?

In statistical comparisons (summarizing features per document by median across windows), several differences emerged:

GPT-assisted writing showed higher:
- TTR
- Hapax Ratio
- Word Entropy
- NSR
- Noun Ratio
- Verb Ratio
- Adjective Ratio

And independently authored writing showed higher:
- Adverb Ratio (and Sentence-Length Variability)

A particularly strong association signals:
- NSR and Noun Ratio are the strongest positive associations with GPT-assisted writing
- Adverb Ratio shows the strongest negative association

A helpful way to interpret SHAP plots

SHAP tells you, for a given text window, how each feature nudged the prediction up or down.

The paper provides a concrete example:
- For a GPT-assisted example, the selected window had predicted probability 0.947, with NSR and Hapax Ratio contributing most positively.
- For an independent example, the selected window probability was 0.077, where features like Hapax Ratio, Noun Ratio, NSR, and TTR shifted the prediction toward independent authorship.

The practical takeaway: the model’s “reasoning” isn’t mystical—it’s built on observable linguistic behavior, and SHAP can point instructors to the suspected causes.

Limitations You Should Not Ignore (Especially If You’re an Instructor)

This work is solid, but it’s not “solved.” The paper highlights multiple limitations that matter if you’re thinking about adoption.

The small held-out test set limits reliability

The reported held-out evaluation uses 36 documents from 18 held-out participants. The confidence intervals are wide:
- ROC-AUC CI: 0.747–0.981
- F1 CI: 0.737–0.944

That doesn’t mean the model is useless—it means you can’t confidently treat it as a standalone truth source.

The dataset is one specific context

The study uses participants from one course context (students in a Large Language Models course at IIIT Delhi) and one style of assistance (paraphrased GPT output). Other forms of AI use—brainstorming, outline generation, editing, tutoring, or “partial copy + rewrite”—might produce different stylometric patterns.

Order effects and demographic coverage weren’t evaluated

Participants wrote independently first, then later wrote with GPT assistance. That session order could influence writing behavior in subtle ways.

And the study reports no fairness analysis across demographic groups or language backgrounds. Stylometry can interact with language proficiency and writing conventions, so we should be cautious.

Paraphrasing can blur signals—but also changes them

While the study includes paraphrasing by design, real-world students may paraphrase differently. They might also combine AI with other editing behaviors. Since writing patterns can shift over time and across models, external validation is crucial.

Key Takeaways

  • Stylometry can detect GPT-assisted writing using only the submitted text. This study shows measurable signal using interpretable features (lexical diversity, vocabulary richness, and grammatical ratios).
  • The best model, Random Forest, reached ROC-AUC = 0.87 and F1 = 0.842 on a held-out set of 36 documents from 18 participants.
  • Error rates matter: false positives were 22.2% and false negatives 11.1% on that held-out evaluation—too high for automated punishment, but useful for review triage.
  • SHAP explanations point to specific drivers, especially Hapax Ratio, NSR, Noun Ratio, Adverb Ratio, and TTR.
  • The authors’ framing is the right one: use detection as decision support, not as standalone evidence of academic misconduct.
  • Future work should test across more institutions, disciplines, languages, AI assistance styles, and larger external datasets—because current results are anchored to one course and one paraphrasing setup.

If you’re thinking about how this could show up in real academic integrity processes, the biggest lesson is that transparent signals are more defensible than opaque scores—and this research is a concrete step in that direction, even though it’s not the final answer.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Unpacking AI's Role in Academic Integrity: Can AI Help Identify Flaws in Scholarly Writing?

Unpacking Human Touch in AI Writing: A Fresh Look at Academic Integrity

A Measurement-Error Checklist for AI-Assisted Literature Reviews

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime • 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.