The Short Answer
The audit found that AI mental health answers may include broadly authoritative citations, but citations are heavily concentrated and vary sharply across platforms—and non-English questions often get fewer citations and lower rates of language-appropriate routing. This means the “trust” signals users see can differ based on the product and language.
Practically, you should compare the same mental health question across platforms and languages, and explicitly prompt “List your Sources” to check whether the citation set changes meaningfully. Track citation counts, top domains, and whether non-English routing returns less evidence coverage.
A key caveat is that requesting sources shifted citation composition only modestly, so you can’t assume “List your Sources” fixes citation quality or coverage; you still need to verify claims against reputable sources.
On this page
- Why This Matters: Mental Health “Trust” Is Now a Platform Feature
- What the Researchers Actually Measured: Counting Every Citation Across Platforms, Languages, and Prompts
- The Citation Concentration Problem: A Few Domains End Up Defining the Information Environment
- Platform Differences: Same Median Citation Count, But Wildly Different Consistency and Source Preference
- Does Asking for Sources Help? “List your Sources” Changes Little in Volume, Somewhat in Composition
- Multilingual Routing Is Where the System’s Weakness Shows Up Hardest
- Small Glitches, Big Warning Signs: Lexical Confusion and Domain “Neighbor” Effects
- Conclusion: Audited Citations Are the New Baseline for Responsible Mental Health AI
- Key Takeaways
Citations That Gatekeep Mental Health Answers (and How to Audit Them)
Online mental health help is increasingly “conversational.” Instead of a classic search results page where you judge sources from a ranked list, you now get a single AI-crafted answer—often with a curated set of citations tucked underneath. That shift matters because it quietly moves source evaluation from the user to the platform. And until recently, we didn’t have a clear, evidence-based picture of what those platforms actually surface.
New research from the original paper audits the citation behavior of three widely used free AI products—ChatGPT, Perplexity, and Google AI Overview—across 1,140 mental health questions/answers, 15,942 recorded citations, and 1,713 unique domains. Even more interesting: the study doesn’t stop at English. It tests multilingual routing across higher-, medium-, and lower-resource languages (based on language availability), and it measures what changes (and what doesn’t) when users explicitly ask the system to “List your Sources.”
The punchline is both reassuring and alarming: the citations are often from broadly authoritative institutions, but they’re also heavily concentrated and inconsistent across platforms—and for many non-English queries, the systems surface fewer citations and route toward language-appropriate sources at dramatically lower rates.
Why This Matters: Mental Health “Trust” Is Now a Platform Feature
Here’s the real-world implication: when an AI system answers your mental health question, it’s not just answering. It’s curating who gets seen. In practice, that means two users asking the same kind of question might walk away with very different “evidence ecosystems,” depending on which product they used and what language they asked in.
This is especially important right now because mental health queries are often routine psychoeducation (symptoms, diagnosis, treatments, medication info)—not only high-acuity crisis scenarios. The research community has done excellent safety work on crisis handling, but that’s only part of the story. Most people won’t use AI only when things are at their worst. They’ll ask for guidance when they’re unsure, scared, or simply looking for a starting point.
A concrete scenario: imagine someone experiencing anxiety and searching conversationally for “symptoms of bipolar disorder” (a diagnosis/signs query). If the AI response leans heavily on a narrow set of institutional domains, that person may feel more confident—and may be more likely to accept the framing. But if the citations are sparse, inconsistent, or not localized for their language/country context, the “trust” the interface creates can become a shortcut around careful reading. That’s where audits like this become essential: they help regulators, product teams, and researchers ask, “Which sources are you really putting in front of people?”
This work builds on earlier evaluations that were often model-centric, while the user experience is product-centric. It also adds a multilingual lens to a space that’s still disproportionately anglocentric—exactly where disparities tend to widen in real deployments.
What the Researchers Actually Measured: Counting Every Citation Across Platforms, Languages, and Prompts
The study ran a multi-factor audit designed to capture how AI mental health information systems behave as citation gatekeepers. They focused on three free consumer products:
ChatGPT(free tier, logged out)Perplexity(free tier, logged out)Google AI Overview(via Google Search, logged out)
They varied four factors:
| Factor | Levels tested |
|---|---|
| Platform | 3 (ChatGPT, Perplexity, Google AI Overview) |
| Query language | 7 (English + Spanish, Japanese, Ukrainian, Hindi, Nepali, Twi) |
| Prompt condition | 2 (question only vs. question + List your Sources.) |
| Question set | 20 English questions + a subset of 3 translated questions per non-English language |
The question set: routine mental health behaviors, not just crisis
The authors created 20 English questions spanning eight functional categories (non-exclusive), including things like signs/symptoms, diagnosis, medication info, treatment effectiveness, treatment guidelines, and crisis resources. Examples include:
- Signs/symptoms: “What are the symptoms of bipolar disorder?”
- Diagnosis: “How is PTSD diagnosed?”
- Treatment effectiveness: “How effective is exposure therapy for OCD?”
- Medications: “What medications are available for depression?”
- Crisis resources: “What crisis resources are there for psychosis?”
- Meta trust: “What are some of the most trustworthy sources of mental health information?”
For multilingual testing, only 3 questions per language were translated: one symptom/diagnostic, one acute crisis, and one source-seeking meta question.
The citation capture: 15,942 citations mapped to 1,713 domains
For each response, trained annotators recorded citations from every channel they could find—inline citations, clickable links, and reference lists. Then the researchers standardized URLs to their registrable domains (e.g., nih.gov) and extracted signals for:
- whether the URL used a country-code top-level domain (ccTLD) (like
.jp), and - whether the URL included native-language signals (like localized paths or language parameters).
Overall scale:
- 1,140 responses
- 15,942 citations
- 1,713 unique domains
- 600 English responses with 10,972 citations to 1,038 unique domains
- 540 non-English responses with 4,970 citations to 870 unique domains
One notable practical detail: 83 responses (7.3%) contained no citations at all. And the no-citation rate varied a lot by platform.
A “source type” typology to classify what’s behind the links
They then categorized every cited domain into a nine-category typology (e.g., Government/public health, Academic/journal, Commercial health, Social/video, etc.) using a deterministic classifier that was validated against human coding:
- On a domain validation sample, Cohen’s κ was 0.88 between annotators.
- Agreement between the classifier and humans was 0.89 and 0.86 (per annotator vs classifier).
That high agreement matters: it means we can trust the distributional claims about what kinds of sources each product tends to surface.
The Citation Concentration Problem: A Few Domains End Up Defining the Information Environment
One of the most striking findings: citations weren’t just “generally credible.” They were concentrated.
Across English-question citations:
- Government sources and commercial health were the most cited source types, at roughly 22.6% each.
- Academic/journals came next at 21.6%.
- Then nonprofit/advocacy (13.3%) and nonprofit health systems (11.6%).
- Encyclopedias (e.g., wiki-style references) were smaller (3.9%), with social/video at 3.0%.
But the real power dynamic is in the domain concentration:
- The top 10 most-cited domains accounted for 43.6% of all English citations.
That’s a lot of influence for a short list of destinations.
Why this is risky even when sources are “good”
The paper argues an important point: even if the dominant sources are broadly authoritative, channeling users through a narrow set of institutions concentrates system-level risk. If a handful of pages have gaps, biases, outdated material, or a mismatch with a user’s context, that error gets amplified because the AI is effectively steering everyone toward the same citation core.
Think of it like a library recommendation system. Even if most recommended books are high quality, if you always point to the same shelves, you reduce discovery—and you reduce resilience to mistakes.
What changes by question type
Citation composition wasn’t identical across the kinds of questions being asked. The authors report intuitive shifts that have practical consequences:
- Treatment effectiveness questions drew the most from academic/journal sources (56.9%).
- Crisis questions leaned heavily toward nonprofit/advocacy (41.6%).
- Medication questions had the strongest tilt toward commercial health sources at 33.7%.
That last piece is especially important for real users: medication answers may feel “clinical,” but the citation mix can include commercial incentives and branded content more than people realize.
Platform Differences: Same Median Citation Count, But Wildly Different Consistency and Source Preference
Here’s where the study gets really useful: if you’re trying to understand user trust, you can’t only look at “how many sources.” You need to look at how stable the citations are and what types each platform tends to privilege.
Citations per response: medians cluster, means don’t
Across English responses (200 per platform), the median citation count clustered around 10–12 for all three products. But the mean differed dramatically because ChatGPT sometimes produced many more citations than the others.
Reported mean citation counts (English):
ChatGPT: 32.29 (SD 60.50)Google AI Overview: 12.85 (SD 6.63)Perplexity: 9.7 (SD 1.69)
This suggests a “behavioral personality”:
- Perplexity behaves more like a card-based system with a near-constant citation count.
- ChatGPT can embed hyperlinks throughout generated text, so citation volume can “balloon.”
Source-type preference: each platform leads with a different category
Across English-question citations, all three platforms draw heavily from institutional categories (government/academic/commercial/nonprofit health systems), but the leader differs:
| Platform | Most prominent source type (share of citations) |
|---|---|
ChatGPT |
Government/public: 26.4% |
Perplexity |
Academic/journal: 24.3% |
Google AI Overview |
Commercial health: 25.7% |
And the smaller categories show sharper divergence:
- ChatGPT cited Wikipedia/encyclopedias much more: 6.1% (others were ≤0.5%).
- Google AI Overview cited social/video much more: 8.2%, mostly YouTube (7.7%).
Practical implication: if you pick your mental health information tool based on trust signals, you might unintentionally pick a source-mix bias.
Does Asking for Sources Help? “List your Sources” Changes Little in Volume, Somewhat in Composition
A common user strategy is to explicitly request citation transparency—like appending “List your Sources.” The authors tested this directly in two prompt conditions:
1) Question only
2) Question + List your Sources.
How much more content do users get?
Across all English responses, the mean citation count changed modestly:
- Question only: 16.6
- Question + sources request: 19.9
That’s not nothing, but it’s not the dramatic upgrade people might hope for.
The effect was driven largely by ChatGPT:
- ChatGPT mean: 27.3 → 37.2
- Perplexity mean: 9.2 → 10.3
- Google AI Overview mean: 13.4 → 12.3 (basically unchanged/slightly down)
Does the citation mix become “more institutional”?
Prompting for sources produced small shifts in source-type proportions. Reported changes were on the order of a few percentage points:
- Government/public: +2.1 pp
- Commercial health: +1.6 pp
- Academic/journal: +0.8 pp
- Nonprofit health system: -1.6 pp
- Nonprofit/advocacy: -1.3 pp
- Social/video: -1.0 pp
So the takeaway is slightly sobering: you can’t reliably “force” better citations just by asking.
Multilingual Routing Is Where the System’s Weakness Shows Up Hardest
The study also tackles a key concern in AI evaluation: English may be treated as the default, and performance disparities can show up as unequal citation availability and imperfect localization.
For the three translated questions per non-English language set (n=540 total):
- Non-English queries surfaced fewer citations than English queries:
- mean 9.2 (SD 8.7) vs English mean 12.3 (SD 11.6), with 90 English-per-language? (the paper reports the subset size as 90 per language, total 540 non-English)
The range by language was wide:
- Spanish: 11.2
- Twi: 7.1
- The overall result wasn’t uniform decline—it depended on the platform.
Platform-specific multilingual gaps
ChatGPTaveraged 17.5 sources in English vs 4.3 in Twi.- It returned no citations in 33% of Twi and 23% of Japanese responses.
Google AI Overviewreturned no citations in:- 30% of Hindi responses.
Perplexitywas comparatively stable:- it stayed within 8.6–10.2 across the seven languages.
“Are citations localized?” Not as often as you’d think
The paper measures localization in two ways:
1) citations hosted on ccTLD domains (like .jp), and
2) citations pointing to native-language pages on non-local domains (like U.S. domains with Spanish paths).
Using only ccTLDs, language-appropriate citation rates looked like:
- High-resource languages (Japanese, Spanish): 38.0%
- Medium-resource (Ukrainian, Hindi): 20.8%
- Low-resource (Twi, Nepali): 3.3%
But this underestimates localization because many systems route to translated content hosted on international domains rather than local national sites.
When the authors counted native-language pages on non-local domains too, the language-appropriate shares improved substantially:
- Twi: 2.5% → 37.6%
- Spanish: 11.3% → 46.2%
This distinction is huge for interpretation: “language-appropriate” can be satisfied via translation hosted anywhere, but that doesn’t guarantee that the content matches local clinical systems or locally relevant public health guidance.
The deeper issue: localization tracks web supply, not language itself
The authors argue that multilingual behavior follows what’s available and indexed on the public web, especially authoritative in-language material—not the inherent “capability” of the language.
Practical example from the paper’s framing:
- Spanish queries often route to translated U.S. government sources (e.g., Spanish sections on medlineplus.gov), which may be useful in the U.S. but not necessarily aligned with clinical norms in other Spanish-speaking countries.
That’s a subtle but important point: “citations exist in Spanish” isn’t the same as “citations are locally relevant for Spanish speakers worldwide.”
Small Glitches, Big Warning Signs: Lexical Confusion and Domain “Neighbor” Effects
In a few cases, responses included links that were basically mismatches—like a mental health query being routed to an unrelated page sharing similar terms. The paper reports examples of lexical neighbor confusion:
- Anxiety medications query pulled a Wikipedia “Mayonnaise” page alongside Mayo Clinic links.
- Depression diagnosis query returned links to
DSM-Firmenich(a fragrance conglomerate) and other acronym neighbors. - PTSD query surfaced a construction company because it shared acronym overlap with “PCL-5.”
- A Hindi depression diagnosis query returned “Hamilton” ticket links while also showing a depression rating scale.
The authors connect these errors to known cybersecurity patterns such as typosquatting and combosquatting, where bad actors register domains adjacent to trusted institutional names and wait for retrieval systems to “accidentally” pull them in.
Even though they didn’t find evidence of malicious health misinformation in their dataset, the mechanism is worrying: retrieval systems that match based on lexical proximity can be occupiable by adversaries. In other words, the same weakness that causes harmless mix-ups could be exploited for harmful routing.
Conclusion: Audited Citations Are the New Baseline for Responsible Mental Health AI
The main message of this research is simple but hard to ignore: consumer AI products have quietly become a primary triage point for mental-health information, and their citations are neither neutral nor uniform.
For English users, the citations are heavily anchored in an institutional core, with roughly three-quarters coming from categories like academic, government, commercial-medical, and nonprofit health systems—and the top 10 domains account for 43.6% of citations. That concentration may feel trustworthy, but it also concentrates risk.
For non-English users, the systems surface fewer citations and localization happens at rates tied to in-language web availability, not to language quality or the language’s “resource tier” alone.
Finally, the study makes its audit instruments reusable: a nine-category typology, a validated classifier, and an annotated corpus are released for ongoing measurement across systems, languages, and time (see the dataset link in the paper: huggingface dataset and repository described there).
If we want mental health AI to be more than a polished answer, we need routine product-level citation audits—because citations are not decorative. They’re part of the interface to what people believe.
Key Takeaways
- Citations are highly concentrated: in English, the top 10 domains account for 43.6% of all citations (10,959 citations in the English subset).
- Platforms differ more in consistency and source mix than in median volume: medians cluster around 10–12 sources, but
ChatGPTproduces far higher mean citation counts (32.29) due to outlier responses. - Source type composition varies by platform:
ChatGPTleans toward government/public (26.4%),Perplexitytoward academic/journal (24.3%), andGoogle AI Overviewtoward commercial health (25.7%). - Asking for sources helps only a bit: appending
List your Sources.raises mean citations modestly (16.6 → 19.9 across English), with the biggest shift inChatGPT. - Multilingual performance drops for many languages: non-English queries have fewer citations on average (9.2 vs 12.3 for English), and “no citations” events are common for
ChatGPTandGoogle AI Overviewin some languages. - Localization depends on web availability, not just language: language-appropriate routing (measured via ccTLDs and native-language signals) is much lower when only ccTLDs are counted, especially for low-resource languages.
- Citation-driven gatekeeping creates risk: even when sources are authoritative, routing users through a narrow set of destinations concentrates errors and biases at scale.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries — arXiv
- Authors: Authors: Phuong Anh Nguyen, Jill Noorily, Matthew Flathers, Haruka Notsu, Laura Ospina-Pinillos, Tommy Nguyen, Samantha Clark, Aoife Keane, Grace Thompson, John Torous