The Short Answer
GPT-assisted writing can be detected from the submitted text using interpretable stylometric features, reaching held-out ROC-AUC 0.87 and F1 0.842 with a Random Forest classifier. SHAP shows lexical and grammatical characteristics drive predictions.
For instructors, this enables a decision-support workflow: when a submission scores high, reviewers can inspect the specific text features (e.g., lexical richness or grammatical ratios) rather than rely on a black-box detector.
It’s not standalone proof of misconduct—false positives and false negatives remain—so results should be treated as evidence to review alongside context such as drafts or writing history.
On this page
- Introduction
- Why This Matters
- What the Researchers Actually Measured: Interpretable Stylometric Features
- How the Study Was Set Up: Preventing Leaky Detection
- Which Classifiers Worked Best—and What the Numbers Say
- What the Model Looked At: SHAP Explanations in Plain English
- Limitations You Should Not Ignore (Especially If You’re an Instructor)
- Key Takeaways
GPT-assisted Writing Detection with Interpretable Stylometry: A New Step in Academic Integrity
Detecting AI writing using human-readable text signals
Introduction
If you’ve ever tried to figure out whether an essay is truly a student’s work—or subtly “borrowed” help from ChatGPT—you’re not alone. The hard part isn’t just spotting AI use. It’s getting evidence that’s both accurate enough to be useful and transparent enough to be defensible when grades and discipline are on the line.
New research from Kumar, Siddiqui, and Fuchsberger tackles exactly this problem: detecting GPT-assisted writing using interpretable stylometric features pulled straight from the final submitted text. Instead of relying on black-box detectors or proprietary model internals, the study asks a simpler, more teacher-friendly question: Can we detect AI-assisted writing using measurable writing-style signals that are directly visible in the text itself? And importantly, does it work on writing from new students it hasn’t seen before?
The results are promising. Using data from 9,090 participants—each providing both independently authored and ChatGPT-assisted responses (with paraphrasing instructions)—the best model (a Random Forest) achieved a held-out ROC-AUC of 0.87 and an F1-score of 0.842. It also provides post-hoc explanations via SHAP, pointing to lexical and grammatical characteristics as the strongest drivers. The catch? The false-flag risk isn’t zero, and the authors are clear that this shouldn’t be used as standalone proof of misconduct.
Why This Matters
This research is significant right now because institutions are stuck in a tough spot: they need scalable detection (for large classes, lots of submissions), but they also need procedurally fair evidence. Many current approaches are either (a) opaque (can’t explain why), (b) brittle (break under paraphrasing or stylistic variation), or (c) require access to generation signals (which you don’t have once a student submits an essay).
What makes this work feel “actionable” is the format of the evidence. Instead of saying “my model thinks this is AI-written,” it focuses on things instructors can actually inspect in principle: lexical diversity, vocabulary richness, the rate of rare words, and grammatical ratios. That doesn’t automatically make the decision fair—but it moves you from “trust the detector” toward “review the linguistic signals.”
A concrete scenario where this could be applied today: imagine a university’s writing center or academic integrity team running a decision-support workflow. When a submission triggers a high suspicion score, the reviewer isn’t stuck with a single opaque number—they can also see which text features drove the score (for instance, unusual patterns in Hapax Ratio, noun/adjective/verb ratios, or word entropy). This gives the reviewer something to cross-check against context: the student’s prior drafts, known writing proficiency, revision history (if available), and the plausibility of stylistic shifts.
How does this build on previous AI detection research? Earlier methods often use statistical patterns learned from language models, generation-probability tricks, or watermark-style signatures. The study you’re reading explicitly positions itself as a more text-intrinsic, transparent alternative, extending prior stylometry-style ideas by (1) using a sliding-window approach, (2) evaluating on unseen participants with a strict “no data leakage” user split, and (3) applying SHAP to interpret feature contributions. It’s not trying to replace instructor judgment—it’s trying to give it a sturdier footing than “trust the black box.”
What the Researchers Actually Measured: Interpretable Stylometric Features
The core idea is straightforward: even if GPT can mimic writing style, it may still leave patterns in how language is used. Stylometry captures those patterns without caring about the essay’s topic.
The study uses nine stylometric features computed from each text window, including:
- NN: total number of tokens
- |V|: number of unique tokens (vocabulary size)
- Hapax-based measures (including Hapax Ratio): how often rare words appear
- Type-Token Ratio (TTR): lexical diversity proxy
- Word Entropy: information-theoretic measure of token-frequency dispersion
- Non-stopword Ratio (NSR): proportion of content words
- Grammatical ratios based on part-of-speech tags (noun/verb/adjective/adverb ratios)
- Sentence-Length Variability: variation in sentence lengths
- Plus supporting sentence stats (mean and variability derived from sentence token lengths)
Why sliding windows matter (and why it’s not just “more data”)
Instead of computing features on the whole document, the authors use a sliding window protocol: 250 words per window with a 125-word stride (so 50% overlap). That gives multiple samples per document and helps reduce issues caused by variable length.
Here’s the analogy: think of the essay as a long road trip. If you only measure the whole trip’s distance, a detour early or late might be hidden. Windowing gives you several “checkpoints” along the route, so the model learns whether GPT-like signals appear consistently—not just in one region.
But there’s an important statistical caveat: windows from the same participant aren’t independent samples. The authors respect this by keeping all windows for a participant inside the same training/validation/test partition.
GPT-assisted here means paraphrased GPT output—not copy-paste
The dataset comes from the IIITD-BU (Paraphrased) collection described in the authors’ cited work (Mehta et al., 2026). Participants responded to prompts in two sessions:
1. Independent writing (no LLM use)
2. Assisted writing where students could use LLMs like ChatGPT, but were required to paraphrase what the model produced before typing
Copy-and-paste and grammar tools were disabled in that original setup. So in this study, “GPT-assisted” means: the student typed content that originated from GPT but was paraphrased manually.
That restriction matters because it makes the detection task more realistic: paraphrasing is exactly what people attempt to do to “remove” obvious AI fingerprints.
How the Study Was Set Up: Preventing Leaky Detection
Detection studies are notoriously easy to accidentally “cheat” on. A classic problem is data leakage: if the same person appears in both training and testing, the detector may learn that person’s idiosyncratic style rather than GPT-related signals.
This paper goes out of its way to avoid that.
The dataset and the user-split protocol
The dataset contains 180 samples per participant condition, but more precisely, the study builds a balanced set where each of 90 participants contributes:
- one independently authored response and
- one GPT-assisted response
That’s the raw “pairing,” but the study’s crucial number for ML evaluation is the participant count in the split.
They use 9,090 participants total for evaluation (as described in the abstract and methodology framing), partitioned into:
- 7,272 training participants (144,144 samples/windows)
- 1,818 held-out participants (3,636 samples/windows: 1,818 independent + 1,818 GPT-assisted)
All windows from a participant stay in the same split.
Windowing turned documents into many samples
After cleaning the final text, each document is split into windows:
- Most documents longer than 250 words produced 250-word windows
- Documents shorter than 250 words became a single full-document window
Across the dataset, the study reports 722 windows, including:
- 397 independently authored windows
- 325 GPT-assisted windows
(And for the held-out participants, they describe 127 windows across 3,636 documents—the paper provides these counts as part of the window-level dataset framing.)
Model outputs: window probabilities → document median
Each classifier outputs a probability for each window being GPT-assisted. To convert those window scores into a document decision, the authors aggregate using a median-probability rule:
- Document score = median of window probabilities
- Classify as GPT-assisted if the median probability ≥ 0.50
Median aggregation is a smart choice here: it reduces the impact of a single “weird” window that might look unusually GPT-like (or unusually human).
Which Classifiers Worked Best—and What the Numbers Say
The paper compares eight machine learning classifiers trained on the nine stylometric features:
- Logistic Regression
- Linear Discriminant Analysis (LDA)
- Quadratic Discriminant Analysis (QDA)
- Naive Bayes
- Random Forest
- Linear Support Vector Machine (linear SVM)
- K-Nearest Neighbors (kNN)
- Multilayer Perceptron (MLP)
They tune hyperparameters using Bayesian optimization on the training partition only, with careful cross-validation that keeps participant data intact.
Performance comparison (validation and test)
Validation showed multiple methods performing in the same general ballpark, but Random Forest came out on top and also ranked first in an additional nested-validation check.
Then comes the real test: the held-out evaluation on unseen participants.
| Metric (held-out test set, 36 documents) | Best model (Random Forest) |
|---|---|
| ROC-AUC | 0.870 |
| F1-score | 0.842 |
| False Positive Rate (independent flagged as GPT-assisted) | 22.2% |
| False Negative Rate (GPT-assisted missed as independent) | 11.1% |
A key point: the paper’s held-out evaluation described 36 documents coming from 18 held-out participants (so each side has 18 documents: independently authored and GPT-assisted). That’s a small test set, and the paper reports wide confidence intervals, meaning uncertainty is non-trivial.
What those error rates mean in real academic settings
They treat GPT-assisted as the positive class. On the held-out set:
- 4 independent documents were incorrectly classified as GPT-assisted
→ False Positive Rate = 22.2%
- 2 GPT-assisted documents were incorrectly classified as independent
→ False Negative Rate = 11.1%
False positives are the big concern in academic integrity because they can lead to unfair accusations. Even if the model is fairly strong, a ~1 in 5 false-positive rate on this small held-out set is too risky for anything but decision support.
The authors explicitly position the system as a tool for contextual review, not a verdict machine.
What the Model Looked At: SHAP Explanations in Plain English
Here’s where the paper gets especially interesting: it doesn’t just report performance—it investigates which features were driving the Random Forest predictions.
Using SHAP analysis, they identify the most influential stylometric features globally:
- Hapax Ratio (most influential)
- NSR (Non-stopword Ratio)
- Noun Ratio
- Adverb Ratio
- TTR (Type-Token Ratio)
- Smaller contributions from: Verb Ratio, Sentence-Length Variability, Word Entropy
- Adjective Ratio contributes least
Direction of effects: which writing style signals match GPT-assisted text?
In statistical comparisons (summarizing features per document by median across windows), several differences emerged:
GPT-assisted writing showed higher:
- TTR
- Hapax Ratio
- Word Entropy
- NSR
- Noun Ratio
- Verb Ratio
- Adjective Ratio
And independently authored writing showed higher:
- Adverb Ratio (and Sentence-Length Variability)
A particularly strong association signals:
- NSR and Noun Ratio are the strongest positive associations with GPT-assisted writing
- Adverb Ratio shows the strongest negative association
A helpful way to interpret SHAP plots
SHAP tells you, for a given text window, how each feature nudged the prediction up or down.
The paper provides a concrete example:
- For a GPT-assisted example, the selected window had predicted probability 0.947, with NSR and Hapax Ratio contributing most positively.
- For an independent example, the selected window probability was 0.077, where features like Hapax Ratio, Noun Ratio, NSR, and TTR shifted the prediction toward independent authorship.
The practical takeaway: the model’s “reasoning” isn’t mystical—it’s built on observable linguistic behavior, and SHAP can point instructors to the suspected causes.
Limitations You Should Not Ignore (Especially If You’re an Instructor)
This work is solid, but it’s not “solved.” The paper highlights multiple limitations that matter if you’re thinking about adoption.
The small held-out test set limits reliability
The reported held-out evaluation uses 36 documents from 18 held-out participants. The confidence intervals are wide:
- ROC-AUC CI: 0.747–0.981
- F1 CI: 0.737–0.944
That doesn’t mean the model is useless—it means you can’t confidently treat it as a standalone truth source.
The dataset is one specific context
The study uses participants from one course context (students in a Large Language Models course at IIIT Delhi) and one style of assistance (paraphrased GPT output). Other forms of AI use—brainstorming, outline generation, editing, tutoring, or “partial copy + rewrite”—might produce different stylometric patterns.
Order effects and demographic coverage weren’t evaluated
Participants wrote independently first, then later wrote with GPT assistance. That session order could influence writing behavior in subtle ways.
And the study reports no fairness analysis across demographic groups or language backgrounds. Stylometry can interact with language proficiency and writing conventions, so we should be cautious.
Paraphrasing can blur signals—but also changes them
While the study includes paraphrasing by design, real-world students may paraphrase differently. They might also combine AI with other editing behaviors. Since writing patterns can shift over time and across models, external validation is crucial.
Key Takeaways
- Stylometry can detect GPT-assisted writing using only the submitted text. This study shows measurable signal using interpretable features (lexical diversity, vocabulary richness, and grammatical ratios).
- The best model,
Random Forest, reached ROC-AUC = 0.87 and F1 = 0.842 on a held-out set of 36 documents from 18 participants. - Error rates matter: false positives were 22.2% and false negatives 11.1% on that held-out evaluation—too high for automated punishment, but useful for review triage.
- SHAP explanations point to specific drivers, especially
Hapax Ratio,NSR,Noun Ratio,Adverb Ratio, andTTR. - The authors’ framing is the right one: use detection as decision support, not as standalone evidence of academic misconduct.
- Future work should test across more institutions, disciplines, languages, AI assistance styles, and larger external datasets—because current results are anchored to one course and one paraphrasing setup.
If you’re thinking about how this could show up in real academic integrity processes, the biggest lesson is that transparent signals are more defensible than opaque scores—and this research is a concrete step in that direction, even though it’s not the final answer.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- Detecting GPT-Assisted Writing Using Interpretable Stylometric Features — arXiv
- Authors: Authors: Rajesh Kumar, Nabeel Siddiqui, Alexander Fuchsberger