The Short Answer
Accounts begin receiving ChatGPT ads roughly 14 days after account creation, and lower-income accounts are more likely to get ads (across race). Early ads skewed toward consumer goods and were clearly separated from the LLM’s response text.
So what: when you use ChatGPT for purchase- or skill-building intents, you may see sponsored prompts more often in lower-income contexts, making “friendly” recommendations feel like advice while still being ad-driven.
Caveat: this first phase reflects an early rollout audit using simulated accounts and prompts, so ad behavior may change as ads become more integrated into LLM chat interfaces.
On this page
- Why This Research Is Significant Right Now (And How It Shows Up in Your Life)
- What the Researchers Actually Measured: A Real-World Audit of ChatGPT Ads
- The Most Important Findings: Income Effects, Weak Race Signals, and How Quickly Ads Kick In
- What Kinds of Ads Showed Up: Consumer Goods, Retail, and Contextual Triggers
- The Real-World “So What”: Implications for Users, Platforms, and Regulators
- Key Takeaways
ChatGPT Ads Are Here: The First Real Measurements of Who Gets Targeted
ChatGPT ads are no longer a theoretical future—they’re showing up in user-facing interfaces right now. And new research from the arXiv paper “The Beginning of ChatGPT Ads” takes one of the earliest, most systematic looks at what those ads actually look like, who gets them, and how quickly they start appearing after an account is created.
The study is based on an “audit” of the ad system during the first rollout phase. The authors set up a large set of simulated accounts across different U.S. racial/ethnic and income groups, then asked realistic ChatGPT questions and recorded every ad they received—over thousands of conversations. If you’ve been wondering whether chatbots are about to become the next major ad surface, this paper gives the first empirical baseline.
And the headline findings are pretty striking: ad exposure begins roughly 14 days after account creation, lower-income accounts receive ads more often (across race), and early ads were mostly consumer goods—clearly separated from the model’s normal answers rather than smoothly blended in.
Why This Research Is Significant Right Now (And How It Shows Up in Your Life)
This matters right now because ChatGPT isn’t just “another website.” It’s an interface people treat like an information intermediary—someone they trust to help with decisions. When that intermediary starts injecting ads, even if they’re labeled “Sponsored,” the persuasive power changes. The chatbot style (friendly, coherent, tailored-sounding) makes recommendations feel less like marketing and more like advice.
Here’s a concrete scenario where this could apply today: imagine a student or new job seeker using ChatGPT at night to figure out “What training should I do to get an entry-level job?” The paper shows that ads were especially likely to appear for prompts that look like purchasable or skill-building intents—like education, fitness, cooking, and product-related topics. If you’re lower-income, the study found you’re more likely to see ads in general. That means your “help” chat could be nudging you toward specific brands or subscription products more often than someone in a higher income bracket.
Compared to earlier AI research, this study is doing something more urgent and operational: it’s not only asking whether models can produce biased content, or whether automated systems discriminate behind the scenes. It’s measuring ad delivery itself inside a closed chatbot experience where researchers can’t normally observe what users see. The result builds directly on prior auditing work using “sock puppets,” but it pushes that method into a brand-new ad frontier—LLM interfaces.
What the Researchers Actually Measured: A Real-World Audit of ChatGPT Ads
The paper’s core contribution is methodological as much as it is descriptive: it’s the first large-scale empirical audit of advertising content appearing in user-facing LLM chat interfaces.
The setup: 91 simulated accounts across race and income
The authors created 91 “sock puppet” accounts arranged in a 3×3 factorial design:
- Race/ethnicity cues: Black, Hispanic, White
- Income terciles: low, medium, high
- At least 9 accounts per demographic cell (so roughly equal coverage within each group)
To make those accounts feel geographically consistent (because ad systems often use location), the team used residential proxy routing plus geography-establishing prompts at the start of each chat session. Their proxy system targeted ZIP codes that map to demographic cells using U.S. Census data (median income thresholds around $57,800 and $77,800 for low/medium and medium/high).
How they prompted ChatGPT (and what they recorded)
They used 335 prompts taken from:
- OpenAI’s “How People Use ChatGPT” research-style categories (the paper cites Chatterji et al., 2025 within its method section),
- top-voted Reddit posts (after filtering for ChatGPT-like phrasing),
- and researcher-curated prompts on socially important topics (including health, mental health, politics—areas OpenAI said would be excluded from advertising early on).
Each day, every account received the same random subset of 30 prompts—plus up to 20 additional “extra” prompts that had previously elicited ads. That design choice increases the odds of observing ad behavior, but it also means the dataset is not a perfectly random slice of all user prompts.
The timeline: when ads start appearing
Data collection began in early February 2026, with accounts created in late January through mid-March. The authors report:
- No ads at first
- Ads were primarily observed starting March 8, 2026
- The study window for analysis is March 8–31, 2026
- Then ad volume dropped sharply after March 31, and the authors collected a small number of additional ads later, bringing totals to the full dataset
Scope: thousands of conversations and hundreds of advertisers
During the “ads rolling out” period, they had:
- 63,657 interactions in March 8–31
- 3,573 ads appearing in 5.61% of interactions
- Ads from 191 unique advertisers across 102 distinct prompts
Across the broader collection window (Feb 6–May 20), they report 3,602 total ads and 127,801 conversations.
They also released a public searchable archive (“ChatGPT Ad Library”) so other researchers can analyze what was captured:
https://emmalurie.github.io/chatgpt-ads-library/
The Most Important Findings: Income Effects, Weak Race Signals, and How Quickly Ads Kick In
This is the part most people will care about: who gets ads, how fast, and what kinds of ads they are.
Ad exposure starts about two weeks after account creation
The paper measures the median time from account activation to first ad exposure:
- Median: 14.0 days
- Mean: 17.6 days
- Strong right skew, with most first-ad events clustered around 13–16 days
Specifically, 62.1% of accounts that eventually received ads got their first ad within 14 days, and no account received an ad in the first 7 days.
That’s a meaningful signal: it suggests ChatGPT’s ad eligibility or ad-serving infrastructure depends on more than just “what prompts you type today.” It may involve system-level readiness, account maturity, or engagement scoring—even in the early rollout.
Lower-income accounts are more likely to receive ads
The clearest demographic result is about income, not race. The authors estimate ad exposure probability using logistic regression with the median household income associated with each account’s ZIP code. They find:
- Odds ratio of 0.98 per $1,000 increase in income
- p = 0.0438 (statistically significant)
So as income rises, the odds of receiving ads drop. Importantly, they ran robustness checks:
- Adding a quadratic term didn’t change the story (non-significant, OR ≈ 0.99)
- Removing three high-leverage accounts strengthened the effect (p = 0.0053)
They also show distributional patterns consistent with this not being driven by a couple of weird outliers: among accounts that received ads, the income values cluster around a modal value near $60,000, while unexposed accounts peak higher (around $80,000).
Race-based differences weren’t detected—though the study may not be powerful enough
The study reports no significant association between race/ethnicity cell and ad exposure or ad rate. The paper emphasizes caution: sample sizes are limited, so small effects could be missed.
Among the 53 accounts that received at least one confirmed ad in the March window, mean ad rates were roughly:
- Low-income: 27.4%
- Medium-income: 16.9%
- High-income: 19.6%
And within each race group, they calculate correlations between income and ad rates that trend negative but don’t hit statistical significance (example correlations cited include White ρ = −0.551 (p = 0.08), Hispanic ρ = −0.408 (p = 0.08), and Black ρ = −0.257 (p = 0.30)).
A key nuance: “no evidence of race differences” in a limited study isn’t the same thing as “no race targeting.” The paper’s own framing is responsible here: this is early rollout, with a truncated study period and underpowered demographic comparisons.
Ads stay “persistent” once they begin
Not only do ads start after about two weeks—but once ad serving starts, it’s not totally sporadic. The authors report that for exposed accounts measured from each account’s first ad date onward:
- Median ad rate: 24% of prompts
- Mean: 21.9%
- Interquartile range: 9.6% to 32.7%
- Most accounts weren’t constantly at 0% or 100%—they were in the middle, suggesting a sustained serving pattern
Even the “lightest” quartile accounts were receiving ads on about one-in-ten prompts.
The ad format: clearly separated from ChatGPT’s output (for now)
One of the more operational findings is about how the ads were displayed. In this first phase:
- ads appear as a demarcated unit at the foot of the conversation
- they’re clearly labeled and separated from the LLM response text
- the paper flags that this may change as ads integrate more deeply into chat interfaces
In other words, early on, it’s “ads appended after answers,” not “ads woven into the model’s reasoning.” That distinction matters for user perception and disclosure clarity—and it’s exactly the sort of thing that could change over time.
What Kinds of Ads Showed Up: Consumer Goods, Retail, and Contextual Triggers
Early ChatGPT ads didn’t look like an even mix of everything. They skewed heavily toward commercial categories that feel “natural fits” for recommendations.
Advertisers and sectors: retail dominates, but participation is broad
The 3,602 ads are spread across:
- 16 sectors
- 191 advertisers
The top sectors by volume were:
- Retail Trade: 1,057 ads
- Information: 843 ads
Combined, retail + information account for about 57% of impressions (and each has about 50–59 advertisers participating).
The paper also lists leading advertisers, including:
- Target (130 sessions)
- Top10.com (107)
- Preply (97)
- Advance Auto Parts (92)
The “top advertisers” mix big-box retail, auto parts, streaming, subscriptions, fitness brands, and vocational education—suggesting that advertisers aren’t only buying traditional “shopping” intent, but also lifestyle and training intent.
Some prompt-to-ad matches are obvious; others are oddly contextual
The authors point out that not every ad placement follows clean logic. Some prompt-ad pairings are straightforward—for example, questions like “best streaming service for sports?” leading to streaming brand ads.
But other matches are more surprising. They give an example of a recipe prompt (“step-by-step guide to make pad Thai”) yielding food delivery ads. That’s not totally irrational (delivery can follow recipe interest), but it does raise questions about how broadly ad systems interpret “helpful assistant” user intent.
Topic patterns: purchases and cooking get ads; excluded content mostly doesn’t
The paper categorizes prompts into topic labels and reports that:
- “Purchasable Products,” “Health, Fitness, Beauty & Self-Care,” and “Cooking & Recipes” elicited ads in roughly 10–14% of sessions
- other categories were far lower, sometimes below 3%
- an “OpenAI-Excluded” category (health, mental health, politics, per OpenAI’s stated approach) produced near-zero ad rates
Here’s a simplified view of the contrast the authors emphasize:
| Prompt topic cluster (as used in analysis) | Approx. ad rate reported |
|---|---|
| Purchasable Products | ~10–14% |
| Health, Fitness, Beauty & Self-Care | ~10–14% |
| Cooking & Recipes | ~10–14% |
| “Argument or Summary Generation” / reflection-type topics | <3% |
| OpenAI-excluded topics (health/mental health/politics) | ~0% |
They also cross-check this using an embedding-based clustering approach (k-means on prompt text embeddings). That robustness check produces a similar pattern:
- clusters tied to consumer product terms (“Nike, Adidas, running shoes”) show ad rates around 16%
- clusters centered on mental health and political content are close to 0%
That convergence matters because it reduces the likelihood that the topic result is just an artifact of human labeling.
If you want to see the exact ad-prompt relationships, this is one place the released archive is especially useful—since you can search by prompt text or advertiser name: the ChatGPT Ad Library.
The Real-World “So What”: Implications for Users, Platforms, and Regulators
This paper isn’t just reporting what ads existed; it’s raising alarms about what happens when ads enter a conversational layer.
Why income-based ad exposure is the big practical red flag
The strongest empirical pattern is income. Even without evidence of race-based delivery differences in this early dataset, income gradients can still be harmful in a “burden of exposure” sense: if lower-income users see ads more often, they may be more frequently exposed to persuasive content, sales funnels, and brand-driven guidance.
And because chatbots often feel “helpful,” it’s easier for users to underestimate persuasive intent—especially when ads are appended in a consistent location.
Why the “excluded topics” result is reassuring—but not a guarantee
The findings roughly match OpenAI’s stated intention to avoid ads on health, mental health, and politics prompts (at least at rollout start). That is good news.
But the paper also notes the uncertainty: real users will ask messy, borderline, or ambiguous questions, and edge cases might slip through over time—especially as the ad system learns what still “qualifies” and what advertisers push for.
This is why the study’s framing—baseline + ongoing monitoring—is so important. Early guardrails can be real, but they can also degrade as systems optimize and scale.
The methodology itself shows the stakes (and fragility) of auditing
A key background point: this is hard to study without direct access. ChatGPT’s closed environment doesn’t expose an API to see ads. The researchers used sock puppets routed through residential proxies tied to specific ZIP codes.
That works—but it’s expensive and fragile:
- They suspect the system detected inauthentic behavior around late March.
- Ad collection peaked and then dropped sharply around March 29, reaching zero by April 2.
So the dataset includes what likely represents “early rollout behavior before detection thresholds were tripped.” That’s exactly why a baseline matters: if future researchers can run more robust audits, comparisons over time become possible.
Key Takeaways
- Ads show up fast after account creation: median time to first ad was 14 days (no ads within the first 7 days).
- Lower-income accounts received ads more often: the odds of ad exposure decreased with higher median ZIP income (logistic regression OR 0.98 per $1,000 increase, p = 0.0438).
- Race-based differences weren’t detected here: no statistically significant race effect on ad delivery, but the study may be underpowered for small differences.
- Early ad content skewed toward consumer and retail: sectors like Retail Trade and Information made up about 57% of impressions; top advertisers included Target, Top10.com, Preply, and Advance Auto Parts.
- Context mattered more than pure demographics: prompts related to purchasable products, cooking, and health/fitness/beauty drew ads at about 10–14% of sessions, while excluded categories (health/mental health/politics) were near 0%.
- Ads were initially clearly separated from model output: in this phase, ads were appended as a demarcated, labeled element at the conversation footer—likely to change as integration deepens.
- A public searchable archive was released: the captured ad impressions are accessible via the ChatGPT Ad Library for transparency and future comparison.
If you want, I can also pull out a “reader-friendly map” of the study design (what each experimental lever likely corresponds to in real ad systems) and translate the archive fields into a simple guide for how to search it efficiently.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- The Beginning of ChatGPT Ads — arXiv
- Authors: Authors: Emma Lurie, Ro Encarnación, Sorelle A. Friedler, Danaé Metaxa