The Short Answer
SlopBench finds that under one fixed composite weighting, Kimi K2.6 scores the lowest “AI slop” (21.1) and Mistral Large scores the highest (40.6) across 18 models tested on 112 tasks. The study uses reader-checkable surface behaviors rather than relying on AI detectors.
Practically, model choice for writing assistants should include “writing feel” measures—especially opener repetition, paragraph rhythm, and fixed lexical constructions—because plausible, relevant text can still sound mass-produced.
Caveat: the ranking is not stable—under 500 random reweightings the same model does not keep the full order—so you should treat the composite as one view among many and consider reporting the four behaviors separately.
On this page
- Why This Matters: “Slop” Is a Real Product Problem Now
- What SlopBench Actually Measures (and Why That’s the Point)
- Model Ranking Results: Who Looks Least Sloppy, and How Stable Is It?
- Which Slop Behaviors Drive the Differences Between Models?
- Why Detectors and Crowds Aren’t Enough (and Sometimes Give No Signal)
- Key Takeaways
AI “Slop” Ranking Benchmark: SlopBench Tests 18 Models (New Research)
AI models don’t just fail in obvious ways like being wrong or unsafe. Sometimes they succeed—but they still sound mass-produced. You’ve seen it: the slow, generic opener, the same phrasing reused across attempts, the paragraph rhythm that feels oddly flat, and the “template-y” constructions that pop up again and again. That vibe is what the authors of new research call AI slop, and their question is refreshingly specific: which models are least likely to write that kind of stiff, repetitive prose?
This post is based on SlopBench, the benchmark from the original paper. The authors evaluate 18 language models across 112 hand-written writing tasks in four real-world-ish domains (email, essays, social posts, and workplace chat). They generate up to 10 samples per model per task, producing 19,928 outputs total, and then score “slop” using behaviors a reader could, in principle, sanity-check without needing fancy AI-detection tools.
The big result? Under one fixed composite weighting, Kimi K2.6 comes out lowest (21.1) and Mistral Large comes out highest (40.6)—but the ranking isn’t stable. Even more interesting: the paper shows that trying to “detect AI” or using one diversity metric doesn’t reliably tell you which model will sound sloppier. Instead, you need multiple surface-form measures.
Why This Matters: “Slop” Is a Real Product Problem Now
This work is significant right now because “quality” in many deployments has shifted from content truth to writing feel. If you’re using an AI assistant for customer emails, internal updates, or social posting drafts, being factually correct isn’t enough—people care whether the message sounds like a human who wrote it, not a system that generated it. And “slop” is exactly the kind of failure mode that can slip through standard evaluations because the output may be plausible, relevant, and safe.
A concrete scenario: imagine a team using an AI email client to respond to vendors. The assistant is supposed to help, not betray itself. Even if every draft is “good enough,” the organization could still look sloppy if the writing repeatedly uses the same greeting structure, the same opener frame across different replies, and the same mechanical paragraph pacing. In that world, a benchmark like SlopBench can help product teams decide which model family produces the most naturally varied drafts, not just which one scores best on correctness or preference tests.
Also, this research builds on earlier AI writing studies by separating different targets that people often mash together. Some work measures how writing assistants change complexity or lexical diversity; others study AI detector accuracy. SlopBench instead asks: when you repeatedly ask models to write in the same scenario, how often do they fall back into the same superficial patterns? That’s a different question—and it leads to different takeaways. (You can see their motivation and related work framing in the paper’s introduction: they emphasize that authorship detection doesn’t answer “how it reads,” and that style-quality needs different signals.)
What SlopBench Actually Measures (and Why That’s the Point)
SlopBench is designed around a simple idea: authorship detectors classify “machine-written,” but they don’t explain which models are sloppier in the way humans perceive. So SlopBench focuses on surface behaviors that correlate with the “mass-produced” feel—specifically ones that can vary across models in countable ways.
The benchmark setup: 18 models, 112 tasks, 19,928 outputs
SlopBench uses 112 English-language scenarios split into:
- 30 emails
- 27 essays
- 27 social posts
- 28 workplace chats
Each scenario specifies context, a goal, desired length band, and constraints (some tasks forbid things like a formal greeting, lists, engagement prompts, etc.). For each scenario, every model is sampled up to 10 times. Overall, that yields 19,928 generated outputs (out of an expected 20,160), meaning the run is very complete: 98.85% coverage.
The paper also notes a subtle but important detail: many models were accessed via resellers/proxies with route-specific default settings (sampling parameters, system prompt handling, quantization, etc.). So the scores reflect the served model + that route’s defaults, not just a “model family” in the abstract.
The four slop axes: length, opener repetition, paragraph rhythm, fixed lexical tells
SlopBench turns “slop” into a scorecard using four behaviors. Each axis is mapped to a 0–100 scale where higher means “more of the measured slop behavior.” And importantly, the authors treat the result as behavior, not “probability the text is AI.”
Here are the axes in plain language:
Length inflation (CC)
If a task specifies an allowed word-count band, models that routinely overshoot it get penalized. Think: padding. It’s computed using the scenario’s upper bound.Repeated openers (TT)
For each scenario, SlopBench checks the first five tokens of each sample and asks: what fraction of samples share the same exact opener? If a model keeps starting sentences the same way across attempts, that’s a classic “template” vibe.Paragraph rhythm flattening (RR)
For multi-paragraph outputs (mostly relevant to email), they measure how repetitive paragraph length is compared to a human baseline corpus. Flatter-than-human paragraph pacing can read as “stiff.”Fixed lexical constructions / “tells” (LL)
This one counts overused constructions (specific phrases or formatting patterns). The scoring is anchored against human writing corpora that were published before ChatGPT.
The authors also make a big point that two axes are domain-dependent:
- Paragraph-rhythm compares to a human corpus baseline, so it only works where the reference statistic exists.
- In workplace chat, the paragraph-rhythm signal ends up zero for every model, which means that axis contributes to the final score but doesn’t separate models at all.
The composite score is one weighting—NOT “the” truth
The headline ranking in the paper uses one fixed weighting of the axes. But the authors then do something more honest: they stress-test it with 500 random reweightings.
Because slop has no universally agreed definition, this matters. The paper treats the composite score as one reasonable definition among many—then checks how stable the ranks are.
Model Ranking Results: Who Looks Least Sloppy, and How Stable Is It?
Under the released composite formula, the authors report a fixed ordering:
- Lowest slop score:
Kimi K2.6at 21.1 - Highest slop score:
Mistral Largeat 40.6 - Only one endpoint rank survives their strongest stability check; everything in the middle is much less certain.
The headline ranking and the stability story
When they redraw the axis weights 500 times (sampling each axis weight uniformly from 0.10 to 0.45, then renormalizing), they get a striking pattern:
Mistral Largestays the highest in 97% of drawsKimi K2.6stays the lowest in 58% of draws- No draw preserves the full order of all 18 models
So the endpoints are somewhat robust; the internal ordering is basically sensitive to how you weight the different “slop behaviors.”
Where the ranking breaks: confidence intervals don’t fully separate models
They also compute rank uncertainty using a scenario-level bootstrap (resampling scenarios within each domain 500 times). Their rank intervals reveal that:
- Only
Mistral Largehas an interval that collapses to a single rank. Kimi K2.6spans a wide range (reported as 16 to 18 in the paper).- Several models overlap in uncertainty, so you get four main “tie groups” rather than a clean ordered list.
The key comparison: overall slop doesn’t generalize across domains
SlopBench also shows something that matters for real deployment: a model can be relatively clean in one writing domain and much sloppier in another.
For example, they report cases like:
DeepSeek V4 Proscores 13.5 in workplace chat but 40.4 on social postsGemini 3.5 Flashscores 18.9 on essays but 39.3 on emailKimi K2.6stays near the clean end in every domain and doesn’t “win” outright in any single domain
This is a reminder that “slop” isn’t one uniform failure mode. It’s a set of surface behaviors that may show up differently depending on the prompt style, length constraints, and expected formatting.
Snapshot: domain behavior makes ranks unstable
| Domain | What dominates the differences (high level) | Practical implication |
|---|---|---|
| All four axes separate models (email is where the signal is strongest) | Best domain for “slop-based” selection | |
| Essays | Lacks the paragraph-rhythm baseline, so one axis drops out | Rankings change because the scoring space changes |
| Social posts | Same issue as essays: less structure to anchor rhythm | Composite scores become less comparable |
| Workplace chat | Rhythm term returns zero for all models; structural differences dominate | Expect less separation from “rhythm” metrics |
Which Slop Behaviors Drive the Differences Between Models?
This is where SlopBench gets most actionable. Instead of saying “model X is best,” the paper also asks: which axis is responsible for each model’s score?
Repeated openers are the biggest lever
Across the evaluated set, opener repetition (TT) has the widest spread—reported as 24.8 to 91.2 across models. The next widest axes (paragraph rhythm flattening, lexical tells, length inflation) have substantially smaller spreads.
Because opener repetition is often tied to salutations and greeting frames, it’s not just “cosmetic.” It affects the first impression readers get.
A concrete email-specific finding: in 474 of 536 email model-scenario cells, the most common opener includes a greeting addressed to the recipient (e.g., “hi carol thanks for reaching”). That’s considered repeated opener behavior by the exact-match method (five-token window), even when the greeting is arguably normal email practice.
Length inflation and rhythm: contribute differently than you might expect
The paper finds that length inflation is driven mostly by a tail (a few long outputs), not by consistent small padding. In one example, they describe how DeepSeek V4 Pro gets pushed up by a small number of very long responses relative to the band.
Meanwhile, paragraph rhythm only separates models meaningfully in email. In workplace chat, the paragraph-rhythm baseline makes the axis effectively uninformative: the rhythm term is zero for every model, so it can’t help you distinguish quality there.
Lexical diversity: the “usual suspect” didn’t help here
Earlier work often claims AI writing is lexically “impoverished.” SlopBench tests lexical diversity using an MTLD measure and finds that:
- Every model exceeds the human corpus MTLD mean across 72 model-domain pairs
- In other words: models are more lexically diverse than their pre-ChatGPT references
- But this doesn’t mean outputs are less repetitive in the slop sense—models can diversify word choice while still reusing frames and constructions
This is consistent with the paper’s broader message: slop is about surface repetition and mechanical style cues, not generic “word choice variety.”
Why Detectors and Crowds Aren’t Enough (and Sometimes Give No Signal)
SlopBench goes beyond mechanical counting by checking whether other approaches agree with its slop ordering.
AI detector check (Pangram-style): saturated labeling gives no ranking power
The paper tested a commercial detector (Pangram’s browser scorer version 3.3.2) on 18 model texts spanning 65 to 544 words. Result: the detector gave 100% AI-generated with high confidence to all 18—no variance, no usable ranking signal.
This supports a central claim: even if a detector is good at separating human vs machine, it may fail at ranking degrees of “slop” within machine-generated text.
Crowd arena check: too noisy to confirm the ordering
They also compare to an arena-style human judging mechanism (a sequential Elo system with users choosing which output is sloppier). But the paper is frank: with only 2,914 model-games recorded across the 18 models, the correlation between mechanical slop and arena ratings is near zero and the uncertainty is huge.
The reported correlations are small and statistically inconclusive:
- Pearson correlation: -0.115 (95% CI includes negative-to-positive)
- Spearman correlation: -0.072 (also effectively ambiguous)
A key nuance: Elo ratings have inherent noise at this scale. The paper even simulates what the rating spread would look like under a null where all models are equally sloppy, and finds that much of the observed spread is consistent with that noise floor.
So the crowd and detector don’t “disprove” slopBench—but they also don’t validate it well enough to trust the ordering as a human psychological phenomenon.
A warning that affects interpretation
The authors also note a conflict with how their companion website reports a different blend of crowd and mechanical axes (with different normalization and weighting). That website’s board-relative scaling can reorder systems when new models join the leaderboard. In short: even where the same ingredients exist, how you scale them can distort conclusions.
Key Takeaways
- SlopBench focuses on measurable surface “slop,” not correctness, truthfulness, or coherence.
- It evaluates 18 models across 112 tasks and produces 19,928 outputs (nearly full coverage).
- The benchmark uses four interpretable axes: length inflation, repeated openers (exact first-five-token matches), paragraph rhythm flattening (anchored to human corpora), and overused lexical constructions.
- Under one fixed weighting,
Kimi K2.6is lowest (21.1) andMistral Largeis highest (40.6)—but:- With 500 random reweightings, no model ordering stays fully intact.
- Only
Mistral Large’s rank is tightly pinned; most of the table middle is statistically ambiguous.
- Opener repetition is the biggest driver of differences between models—suggesting a practical route for improving “naturalness” is reducing reusable greeting frames across attempts.
- AI detectors aren’t reliable for this job: in one check, Pangram labeled all tested model outputs as AI with no variance, so it can’t rank “slop.”
- Human crowd comparisons at this scale are too noisy to confirm the slop ordering; they’re useful as an exploratory signal, not as decisive validation.
- For real deployment (e.g., email drafting), treat SlopBench as a scorecard of behaviors, not a single “best model” verdict. Email is where the scoring is most complete; other domains have missing baselines or uninformative axes.
If you want, I can also turn this into a practical checklist for teams choosing models for “human-sounding” writing—based specifically on how to use SlopBench’s axes rather than relying on generic “quality” metrics.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- SlopBench: How Well Can We Rank Language Models by Slop? A Multi-Domain Benchmark of Repetitive AI Writing — arXiv
- Authors: Authors: Dhruv Roongta, Harsha Gaddipati, Anh Tuan Huynh