AI Help on Prolific Is Mostly Rare—But a Few Workers Do It a Lot

AI help on Prolific is mostly rare: only ~1% of submissions show AI assistance. But the behavior is concentrated, not uniform—and safeguards miss bounded responses. Here’s what the study found and why it matters for research integrity.
The finding AI help shows up in only ~1% of Prolific submissions, but it’s concentrated among a small slice of workers.
The method The researchers used direct observation by linking donated ChatGPT histories to Prolific submission records, alongside a worker survey.
The safeguard gap Many detectors miss bounded response formats, and about half of assisted submissions involved bounded responses.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

AI assistance on Prolific appears in only about 1% of submissions overall, even though AI use is concentrated among a small number of workers. The study also found no detectable increase over more than three years in their observation window.

For researchers, this means the threat is unlikely to be uniform across datasets—quality issues may cluster by participant and by specific task moments. You should audit and strengthen checks for the response formats that detectors may miss.

A key caveat is that safeguards centered on open-ended LLM detection can miss bounded formats (like ratings or multiple choice), where the study found about half of assisted submissions occurred. Also, approval rates were similar whether submissions were assisted or not, so payment/approval consequences alone may not deter prohibited use.

AI Help on Prolific Is Mostly Rare—But a Few Workers Do It a Lot

Online research platforms are the quiet engine behind a lot of modern science: psychology, economics, human behavior studies, and even parts of AI safety evaluation. The whole setup relies on a basic assumption—each submitted answer is produced by a human. But with generative AI now easy to access, that assumption gets shaky fast.

New research from Sehgal et al. on arXiv takes a big step toward answering the most important practical question: how often do people actually use AI while filling out online research tasks? Instead of relying on self-reports or imperfect AI-detection tools, the authors combine a large survey with a rare kind of “direct observation” by linking donated ChatGPT histories to the platform’s submission records.

And the results are a bit surprising. Using AI is fairly common among some workers, but it shows up in only a small fraction of submissions overall—about 1%. Even more, the behavior doesn’t appear to have exploded over time (within the study’s window). Still, there are clear weak spots in current safeguards—especially because AI-policing (like LLM detectors) often focuses only on open-ended responses, while real help includes other formats too.

Why This Matters: The Real Risk Isn’t “AI Everywhere”—It’s “AI Concentrated in the Wrong Places”

This matters right now because many platforms and researchers are reacting as if the threat is uniform—like if AI help rises, it will rise evenly across all tasks and all participants. But that’s not what this study finds.

Instead, the pattern looks more like a small number of “high-assistance” participants who consult AI frequently, with help appearing in tight bursts during specific parts of tasks. That means the risk isn’t “the average dataset is ruined.” It’s more like: a tiny slice of your data pipeline may become systematically compromised, depending on which kinds of tasks those participants get.

Here’s a scenario you might recognize: imagine a research team running lots of small surveys on attitudes and comprehension. If a platform’s detector mainly flags free-text responses, but many studies include bounded choices (ratings, yes/no, multiple choice), then AI assistance can slip through in exactly the form your quality checks aren’t watching. This paper shows that happens: half of assisted submissions involve bounded responses, which is important because many “AI detector” tools only cover open-ended formats.

Compared to earlier AI-research work, this study is different in a key way. Prior estimates often came from selected experiments or tasks that invited AI use, and many relied on self-report or detection—methods that don’t agree with each other very well. This paper instead uses a direct linkage strategy and covers 712,930 submissions across more than 127,000 studies. That scale makes the findings harder to dismiss as “just one task type” or “just one population.”

What the Researchers Actually Measured (And How They Did It Without Guessing)

Most discussions about AI assistance in online research are stuck between two imperfect options:

  • Self-report: participants tell you whether they used AI, but they may underreport (or just not remember accurately).
  • Detection: systems flag likely machine-generated content, but detectors can disagree and aren’t perfect.

Sehgal et al. used a hybrid approach:

Survey: How Many Workers Ever Used AI on a Study?

They surveyed 2,500 workers on Prolific (in the U.S., U.K., and Canada). Participants were randomly assigned to either:

  • a direct-question condition, or
  • an indirect list experiment designed to reduce the chance of underreporting.

Key results:
- Asked directly, 11.31% reported ever using AI (95% CI: 8.73%–14.29%).
- The list experiment estimated 14.24% (95% CI: 9.73%–18.70%).
- The estimates were close enough that anonymity didn’t produce clear evidence of “everyone is hiding it.”

Direct Observation: Who Used ChatGPT, and Did That Match Submissions?

Then comes the clever part.

The authors linked two sources:

  1. donated ChatGPT histories from workers who use ChatGPT weekly, and
  2. Prolific submission logs for those same workers.

They worked with 408 weekly ChatGPT users whose profiles were broadly similar to the survey sample, observing 712,930 submissions across tasks that occurred between Dec 2022 and Feb 2026, spanning more than 127,000 studies.

They counted a Prolific submission as “AI-assisted” when visible ChatGPT output clearly contributed to the participant’s response (even if it wasn’t strictly allowed by study instructions).

To make this work, they used an automated classifier that looks for temporally overlapping chat exchanges, and they validated that classifier against human labels (submission-level reliability reported as Cohen’s κ = 0.76 in validation).

The “Most Important” Result Up Front: AI Assistance Looks Small at the Submission Level

Across those 712,930 submissions, AI assistance appeared in only:

  • 1.00% of submissions (bootstrap 95% CI: 0.72%–1.32%)

So yes, AI help exists. But it’s not the dominant mode of participation.

How Often AI Help Shows Up: Many Workers Use It—Few Submissions Get Assisted

A useful way to think about this is: worker-level behavior ≠ submission-level behavior.

Worker-level: AI use is common among some participants

Among the observed donors:
- 68.4% had ever used ChatGPT assistance on a Prolific study.
- But even within that group, it wasn’t “all the time.”
- Only 13.5% of workers who used AI had assistance on at least 2% of their submissions.

So lots of workers dabble. Far fewer repeatedly incorporate AI help.

Submission-level: AI assistance is still rare

The overall assistant presence is about 1 in 100 submissions.

The paper also shows something important about how that assistance happens:

  • 40% of assisted submissions contained a single ChatGPT exchange.
  • In the others, the first and last AI messages spanned a median 29% of the study’s duration.

That temporal pattern suggests AI is being consulted for particular responses or steps, not continuously throughout the task.

AI Assistance Concentrates Heavily: 5% of Workers Account for Nearly Two-Thirds of Assisted Submissions

This is where the study becomes especially relevant for research quality—because concentration changes how you should think about detection and safeguards.

The concentration is extreme

Rank workers by number of AI-assisted submissions. Then:

  • The top 5% of workers (21 of 408) accounted for 64.1% of assisted submissions.
  • The top 10% accounted for 78.2%.

This is way more concentrated than you’d expect if AI help were spread evenly. The authors even quantify what happens if you remove those top workers:

  • With the top 5% excluded, the assisted-submission rate drops from 1.00% to 0.40%
  • That’s a 59.4% relative reduction.

Practical implication

If a platform can identify the “repeat assist” behavior, you can reduce most of the risk without trying to punish everyone or flagting every single submission.

This also suggests a strategy: don’t only treat AI assistance as a one-off event. Treat it as a behavioral pattern.

What Kind of Help Is It? Not Just Open-Ended Text—Bounded Formats Get Help Too

A lot of public discussion focuses on AI “writing.” But the authors find that AI assistance covers more than that.

Assisted submissions include both cognitive and subjective content

Among assisted submissions:

  • 55.1% included AI-supplied content for subjective/person-specific responses
    (attitudes, preferences, evaluations)
  • 52.2% included AI-supplied content for knowledge/cognitive responses
    (factual recall, reading comprehension, reasoning)
  • 11.9% included creative production

Also, assistance hits different response formats:

  • 52.6% included bounded responses
    (choices, ratings, exact answers)
  • 57.5% included free-form responses

Because multiple exchanges can occur in a single submission, those percentages can add up to more than 100%.

Why bounded responses matter for safeguards

This is a big deal: many platform checks are strongest for open-ended text, where it’s easier to identify AI-like prose.

But if AI help also covers bounded formats, then detectors that only watch free text will miss a lot of the actual behavior.

The authors note this directly: in their discussion they recommend safeguards should cover bounded responses too, because half of assisted submissions involve bounded formats—while Prolific’s optional LLM assistance detector is limited to open-ended questions.

Payment Consequences: Even When AI Use Was Prohibited, Workers Were Still Paid

If you’re hoping deterrence will solve the issue, this paper challenges that hope.

The authors examine whether AI-assisted submissions were rejected or approved for payment.

Nearly everything was approved

  • Assisted submissions were approved 99.6% of the time.
  • Approval rates were essentially the same whether the response was bounded or free-form:
    • 99.7% for bounded
    • 99.5% for free-form

Even stronger:
- In a subset where study instructions explicitly prohibited AI use, all 60 payment decisions were approved (with 100% approval among observed decisions).

No sign researchers punished AI assistance

Unaffected submissions were approved 99.70% of the time—statistically indistinguishable from assisted ones.

They also run regression analysis (restricted to approved/rejected outcomes) and find AI assistance wasn’t associated with rejection once they adjust for factors like reward, task duration, calendar time, and worker submission volume.

Practical implication

If platforms want AI assistance to decline, “allow and approve” won’t do it.

The deterrent mechanism currently appears weak: workers are financially rewarded even when the task says “don’t use AI.”

Does This Mean AI Threatens Online Research? The Data Says: Not Yet—But Platforms Aren’t Ready

The authors are careful not to overclaim. Their findings:

  • Do not support the idea that AI assistance is an existential threat to online research right now.
  • Do reveal serious gaps that would matter if the behavior increased.

Why “not yet”? Because the observed prevalence is low:

  • 1% of submissions show visible AI contribution
  • no detectable upward trend over more than 3 years

And also because when assistance happens, it’s temporally localized and concentrated among a small group of workers.

But “low prevalence” doesn’t equal “safe systems”

If the system’s defenses are narrow (e.g., focusing on open-ended detection) and the deterrent is weak (approval even when prohibited), then it’s plausible that an arms race is coming—where more workers use AI more often, or where assistance methods evolve.

The authors recommend two concrete safeguard improvements:

  1. Expand detection beyond free-text to include bounded response formats.
  2. Use worker-level quality signals, not just response-level checks—because assistance is highly concentrated among a small fraction of workers.

Limitations: This Study Doesn’t Cover Every AI Tool or Every Platform

No study like this can be perfectly universal, and the paper lists several limitations worth taking seriously:

  • They only link histories from ChatGPT; they didn’t capture other AI chatbots, even though many participants reported other services.
  • They observed weekly ChatGPT users and also selected a cohort of people who donated histories and returned for follow-up. That could bias upward or downward, so the exact “net” rate for all Prolific users might differ.
  • They observe evidence of ChatGPT being used for task-related help, but they cannot fully verify:
    • whether workers incorporated the output exactly as intended,
    • which study-specific policy violations occurred in every case,
    • or how the content was used in final submissions.

And finally:
- This analysis is about one platform (Prolific). Prolific is generally high quality, but results might differ elsewhere.

Still, for the debate about whether online research can continue functioning, this study provides some of the most direct evidence available.

Key Takeaways

  • AI assistance is rare at the submission level: only 1.00% of 712,930 observed Prolific submissions showed visible AI contribution.
  • Worker-level use is common, but repeated use is not: 68.4% of observed donors had used AI at least once, yet only 13.5% used it on ≥2% of their submissions.
  • Assistance is highly concentrated: the top 5% of workers accounted for 64.1% of assisted submissions.
  • It’s not just open-ended writing: around half of assisted submissions involved bounded responses (choices/ratings), not only free-form text.
  • Payment deterrence is basically missing: assisted submissions were approved for payment 99.6% of the time, including cases where instructions explicitly prohibited AI use.
  • Current safeguards look narrow: optional LLM detection focused on open-ended answers misses a meaningful portion of behavior.
  • The study doesn’t support “existential threat” claims—yet but it shows platforms aren’t well prepared if assistance becomes more widespread.

If you’re running research—especially surveys and tasks with lots of bounded responses—this paper is a wake-up call with a twist: the problem isn’t everywhere, but it’s structured. And structured problems are exactly the ones that smarter safeguards can catch early.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Monochrome Cable Tracing With Robot Help (TRACE)

Reddit Help-Seeking Didn’t Drop After ChatGPT—Here’s What Changed

LLM Mental Health Bias: LGBTQIA+ Identity Changes Context, Not Help

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime • 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.