The Short Answer
AI shopping chatbot recommendations aren’t consistent: the recommended products and the displayed sources can change across repeated requests and across interfaces (chatbot vs API vs AI Overviews).
Practically, you shouldn’t treat one chatbot answer as a stable, verifiable buying verdict—compare across interfaces and repeated prompts if you’re relying on the evidence shown.
A key caveat is that isolated responses or API observations can misrepresent what consumers see, so audits must match consumer-facing conditions and the specific source layer displayed.
On this page
- Introduction: when you ask an AI for “the best phone”
- Why This Matters
- How the audit was actually set up: real shopping queries + multiple “ways in”
- What changed across systems: recommendations aren’t just different—they’re framed differently
- What sources you’re shown: the evidence trail is inconsistent across systems and time
- The interface vs API problem: auditors can’t treat APIs as a proxy for what consumers see
- Turning these findings into better audits (and better user behavior)
- Key Takeaways
AI Shopping With Chatbots Isn’t Consistent—Here’s What New Tests Found
Introduction: when you ask an AI for “the best phone”
If you’ve ever asked an AI chatbot for a product recommendation, you’ve probably assumed something pretty reasonable: that the advice is based on the best available info, and that if you ask again you’ll get roughly the same answer.
New research from the paper “If I Had to Buy Just ONE: Galaxy S26 Ultra”: Auditing AI-Generated Product Recommendations challenges that assumption in a concrete, measurable way. The authors ran an “AI audit” of popular systems using real shopping-style questions—then compared what users actually see across different interfaces, repeated requests, and even the provider APIs behind the scenes.
The punchline is simple but important: AI-mediated commercial advice isn’t a stable thing. The product, the way it’s framed (e.g., “my pick is…”), and the sources displayed can all shift—sometimes dramatically—depending on which system you use and how many times you ask.
Why This Matters
Right now, AI chatbots are becoming the “front door” to shopping decisions. Instead of reading reviews, you ask a chatbot, and it synthesizes the answer for you—often in a confident, personalized voice. That changes the psychology of buying: you’re not just gathering information anymore, you’re receiving what feels like a guided recommendation.
This research is significant today because regulators and platforms are increasingly under pressure to explain how these recommendation systems affect consumers. In the EU, for example, ChatGPT was designated a Very Large Online Search Engine under the Digital Services Act, which brings expectations around systemic risk and audits. If an AI assistant can recommend different products and show totally different sources across repeated requests, then “audit evidence” has to be designed differently than it would be for a traditional, deterministic website or a static ranking.
A practical scenario: imagine you’re car-shopping or upgrading your phone. You ask ChatGPT one day, get “my pick,” and feel good about it. Then you ask again later (or ask Gemini), and you get a different “best” product plus different cited sources. If those sources are your “evidence trail,” the instability means your ability to verify the recommendation is also unstable. In other words: the advice you act on may not be the advice an auditor sees, unless auditing accounts for how recommendations vary over repeated attempts and across layers of how sources are shown.
This work also builds on prior AI-auditing research (like audits of generative search quality and claim reliability), but it focuses on a missing gap: commercial advice specifically, using consumer-like queries, and comparing what’s visible in the actual user interface versus what’s accessible via APIs. That “interface vs API” distinction is a recurring issue in AI evaluation—and this paper makes it central to consumer protection.
How the audit was actually set up: real shopping queries + multiple “ways in”
A big reason audits can miss reality is that researchers often use artificial, templated prompts. This paper instead starts from 2,528 real commercial-advice queries (“ConsumerQ”) drawn from large chat datasets—so the questions resemble what people would actually type when asking for product help.
Then the authors compared outputs under multiple conditions:
ChatGPTconsumer interface (logged-out)ChatGPTAPIGoogle Geminiconsumer interface (logged-out)Gemini APIGoogle SearchAI Overviews (the “overview” block that appears in search results)
For analysis, the paper focuses on 117 physical-product queries, manually validated and constrained to avoid overly specific or geography-dependent phrasing. Each query was submitted three times under each condition, reshuffling request order and spacing requests by a median of a couple hours (with differences by provider).
That’s how they ended up with 1,755 observations, and 1,536 responses that actually produced recommendations.
To make the comparisons, the authors didn’t just eyeball answers—they annotated structured properties (like whether the system calls something “best,” or uses first-person preference) and also measured source overlap using a domain- and page-level overlap measure (Jaccard overlap).
If you want the intuition: the audit treats each system like a “shopping guide” that may be drawing from a changing mix of internal retrieval, web search, and formatting templates. The study then checks whether the guide’s final advice and displayed evidence stay consistent.
What changed across systems: recommendations aren’t just different—they’re framed differently
Let’s talk about what happens when the same shopping question hits different AI assistants.
The paper reports a striking difference in how confidently and personally the systems speak. In responses that recommend at least one product:
- ChatGPT expressed a first-person product preference in 79% of product-recommending responses
- Gemini did so in 7%
- AI Overviews did so in 2%
They also measured whether a product is labeled “best.” ChatGPT did this far more often than the others:
- ChatGPT: 73% labeled some product as “best”
- Gemini: 43%
- AI Overviews: 48%
Why should you care? Because this affects how persuasive the recommendation feels. If one assistant says “my pick is…” and another says “here are options,” you’re not evaluating the same kind of claim—even if the underlying products overlap.
Side-by-side comparison: framing choices across providers
| Behavior (product-recommending responses) | ChatGPT (interface) | Gemini (interface) | Google AI Overviews |
|---|---|---|---|
| First-person preference (“my pick…”, “I recommend…”) | 79% | 7% | 2% |
| Product labeled “best” | 73% | 43% | 48% |
| Asks clarifying questions (to narrow choice) | 82% | 94% | 97% |
| Products grouped by category | 71% | 65% | 44% |
There’s more. Even within the same system, repeated prompts can change the framing. When it comes to labeling a product as “best,” repeated requests flipped how a product is presented in roughly 27–40% of cases (depending on the specific behavior and system comparisons described in the paper).
And then there’s the important “what did it recommend?” part: product lists themselves can shift even if the query is unchanged. The study notes that naming differences (e.g., “iPhone 14” vs “Apple iPhone 14”) mean the reported overlap is a lower bound, but the pattern is still clear: recommendations vary.
How stable are repeated recommendations inside the same system?
The paper reports Jaccard overlap across repetitions (higher = more shared recommended products). Mean overlap values they cite include:
- AI Overviews: 0.421
- Gemini: 0.287
- ChatGPT: 0.178
So: repeated questions produced more consistent product sets for AI Overviews and Gemini than for ChatGPT—though all systems showed meaningful variability.
This is where the “one answer” myth breaks. The recommendation isn’t a single fixed output; it’s more like a frequently reassembled bundle.
What sources you’re shown: the evidence trail is inconsistent across systems and time
Now let’s follow the citations—the stuff you’d expect to help you verify the recommendation.
In many cases, the systems display sources with product recommendations, but not always—and even when they do, the overlap is low.
Among responses where sources were recoverable:
- ChatGPT displayed at least one source in 78.6% of interface observations
- Gemini displayed at least one source in 79.0%
- AI Overviews had substantive overview content in only 134 of 351 searches, and among those, sources appeared in 120 observations
When sources appear, systems often show a small number of domains in a single response:
- ChatGPT: median 3 domains
- Gemini: median 3 domains
- AI Overviews: median 5 domains
But domain overlap between systems for the same query and pass was extremely low. The key numbers:
- ChatGPT vs Gemini shared on average only 5.4% of displayed domains
- 76.7% of comparisons shared no domain at all
- ChatGPT vs AI Overviews mean domain overlap was 5.2%
- Gemini vs AI Overviews were closer at 9.8%, but still more than half of comparisons shared no domain
That’s not a “minor citation disagreement.” It’s a different evidence set.
Repeating the same question changes sources too
Inside a single system, sources also vary across repetitions:
- Mean domain overlap across repetitions:
- ChatGPT: 26.0%
- Gemini: 29.8%
- AI Overviews: 45.9%
So a single response captures only part of the sources a system may present over time. The paper even quantifies how source counts grow with repeated requests—e.g., for ChatGPT, the mean number of distinct domains observed per question increased from 2.65 after one request to 5.56 after three; similarly, distinct pages rose from 3.33 to 7.70.
Different “types” of sources appear depending on the system
The paper also explores what kinds of sites show up (editorial/product reviews, retailers, manufacturers, community sources, etc.). Even though the authors treat these as exploratory classifications, the high-level differences are still revealing:
- Editorial and product-review sources accounted for:
- 56.7% of displayed domains in ChatGPT
- 45.2% in Gemini
- For AI Overviews, the distribution was broader:
- editorial/product reviews 25.4%
- user/community 22.2%
- retailers/marketplaces 17.8%
- manufacturers/brands 16.1%
One practical implication: if you’re relying on “the chatbot’s sources” to understand why it recommended something, you can’t assume that sources are stable or identical across systems, let alone across repeated requests.
The interface vs API problem: auditors can’t treat APIs as a proxy for what consumers see
This is one of the most regulator-relevant sections of the paper.
In theory, APIs look perfect for auditing: you can fix parameters, run repeatable tests, and capture returned data cleanly. In practice, this paper shows APIs are not reliable proxies for consumer-facing outputs.
For identical questions submitted moments apart, the interface and API often differed in what domains (and pages) appeared.
The paper reports mean domain overlap between interface and API:
ChatGPTinterface ↔ API: 12.0%Geminiinterface ↔ API: 14.8%
And the mismatch can be total:
- ChatGPT: 60.9% of pairs shared no domain
- Gemini: 43.4% of pairs shared no domain
Side-by-side: interface vs API source overlap
| Provider | Mean domain overlap (interface vs API) | Mean exact-page overlap |
|---|---|---|
| ChatGPT | 12.0% | 4.8% |
| Gemini | 14.8% | 11.9% |
Even when sources appear, the kind of source layer exposed can differ. For ChatGPT, the API can report more about the web search pages and final citations; for Gemini, the API reports grounding sources linked to the answer, but not the separate “full retrieved set” the way you might expect.
So an auditor who checks “what the API cites” is often auditing a different slice of the information pipeline than what the consumer sees.
The paper also notes that the API sometimes shows sources more often than interfaces (by about 9.7 percentage points for ChatGPT and 8.1 for Gemini), but that didn’t translate into consistent source sets.
Why this matters in real life
Imagine you’re a consumer-protection watchdog investigating whether an AI assistant is cherry-picking certain review sites or retailers. If your audit uses only the API, you might miss the exact domains or source types that consumers are actually shown.
This study argues that audits should therefore:
- Use consumer-facing interfaces
- Observe repeated requests
- Treat the source layer as an important dimension (because the “sources” you can extract from an API are not necessarily the sources presented to the user)
If you want to connect this directly back to the paper’s framing, it’s not just “what answer is output”—it’s “what recommendation system repeatedly constructs for consumers,” including how it frames and which evidence it chooses to expose.
Turning these findings into better audits (and better user behavior)
So what should you do with all this, as either a researcher, regulator, or just a normal person trying not to get misled?
1) For auditors: sample repeat behavior, not single shots
A key message from the paper is that commercial advice is not deterministic. If you only run one request and call it “the system’s advice,” you’re collecting one snapshot from a moving target.
The authors literally show source and product drift across the three repetitions. For example, ChatGPT’s observed domains per query rose from 2.65 to 5.56 across three requests. That alone should tell any evaluator: single-run audits are underpowered.
2) For regulators/consumer-protection: treat interface and API as different “products”
This is the “stop assuming equivalence” conclusion. The paper demonstrates low interface–API overlap in both domains and pages.
So when compliance asks for evidence, the evidence collection method matters.
3) For users: treat AI recommendations as starting points, not verdicts
You can still use these tools—but you shouldn’t treat them like definitive purchasing authorities.
In practice:
- If the AI claims “this is the best,” try a second assistant or ask follow-ups.
- If citations matter to you, check whether the cited sources are consistent when you re-ask the question.
- Recognize that a “my pick” voice may be more persuasion than verification.
This isn’t about assuming the AI is “lying.” It’s about acknowledging that recommendations are dynamically assembled—and that the evidence trail may be incomplete or variable.
Key Takeaways
- AI shopping advice isn’t stable. Repeating the same product query can change both the recommended products and the displayed sources.
- The “voice” of recommendations differs sharply by system. ChatGPT used first-person preference in 79% of product recommendations, vs 7% for Gemini and 2% for AI Overviews.
- “Best” labels vary a lot. ChatGPT labeled products as “best” in 73% of responses, while Gemini did so in 43% and AI Overviews in 48%.
- Sources shown with recommendations barely overlap across providers. ChatGPT and Gemini shared only 5.4% of displayed domains on average for the same query/pass, and 76.7% of comparisons shared no domain at all.
- APIs don’t reliably proxy the consumer experience. Interface–API domain overlap averaged 12.0% for ChatGPT and 14.8% for Gemini; many pairs shared no domains.
- Audit recommendations like you’d audit a live system, not a static answer. Sample repeated requests, observe the consumer interface, and explicitly consider which “source layer” you’re measuring.
- For consumers: treat AI recommendations as suggestions to investigate—especially when the assistant speaks with certainty or personal endorsement.
If you’d like, I can also rewrite this in a “what to do next” checklist format for consumers or regulators, using the exact metrics from the paper as guidance.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- "If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations — arXiv
- Authors: Authors: Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante, Bram Rijsbosch, Gijs van Dijck, Anikó Hannák, Gerasimos Spanakis, Konrad Kollnig