ChatGPT vs Google Trust in Health: Text Wins, Voice/Avatars Raise Doubts

Trust in health info isn’t just about the content. Research comparing ChatGPT-style conversational search vs Google shows higher trust for ChatGPT—but interface modality (text vs voice/avatars) can raise doubts. Here’s what to design (and what to avoid) when building AI health tools.
The finding In health tasks, participants reported higher trust in ChatGPT-sourced answers than in Google-sourced search results.
The method The research separated the “agent” (ChatGPT vs Google) from the “interface modality” (text vs speech vs embodied) to see which part drives trust changes.
The caveat When the interface moved from text to voice or embodied interaction, trust varied and often dropped, with privacy and authenticity concerns showing up in interviews.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

People trusted health information delivered via ChatGPT more than the same information delivered via Google in the study tasks. The advantage was tied to perceptions of agent and information credibility, not just the UI feel.

For practitioners, this means conversational interfaces can be a stronger trust “delivery path” than traditional snippets, but you still need to design the interface modality carefully. Text performed better than speech or embodied interaction when the same LLM powered all conditions.

A key caveat is that trust shifted across interface modalities: speech and embodied cues raised doubts, linked in interviews to usability/familiarity and to privacy/authenticity concerns. So “more human-like” presentation isn’t automatically more trustworthy.

ChatGPT vs Google Trust in Health: Text Wins, Voice/Avatars Raise Doubts

Introduction: Trust is the real “search result” in health

If you’ve ever Googled symptoms at 1 a.m., you already know this: the hardest part isn’t finding health information—it’s deciding whether to believe it. New research digs into that exact problem, looking at how people’s trust in health information changes when they use different AI “search” tools and different ways those answers are presented.

This blog post is based on new research from the original paper, From Search Agents to Dissemination Interfaces: Understanding Human Trust in Health Information from Conversational Search. The big question: when people talk to an LLM like ChatGPT, is their trust mostly about what the model says—or about how it shows up in the interface? And how does that compare to traditional search like Google?

Why This Matters

Trust in health info is a high-stakes, daily-life problem right now because the interface is changing faster than people’s habits and safeguards.

Historically, Google gave you a buffer: you click, you read, you compare sources. The LLM flips the workflow—answers arrive as if they’re “done,” which can reduce effort but also reduce the opportunity to verify. That means two things: (1) people may rely more on the conversation itself (tone, fluency, “human-likeness”), and (2) any UI choice—text vs voice vs avatar—can nudge trust up or down in ways designers don’t always predict.

A scenario you could apply today: imagine a clinic triage app that uses a chatbot to explain “possible causes” for knee pain. This research suggests the app shouldn’t assume “more realistic” (speech or embodied) automatically means “more trustworthy.” In the study, users trusted text-based interfaces more than speech- or embodied ones—largely tied to usability and familiarity—and embodied/speech interactions raised privacy/authenticity concerns.

Also, this work builds on earlier AI trust research by treating trust as more than just content quality. It separates the “search agent” (ChatGPT vs Google) from the “delivery interface” (text vs speech vs embodied), which is exactly the kind of design-level decomposition many previous studies didn’t test together.

Main Content Sections

What the researchers actually compared: agent vs interface vs trust outcomes

The paper runs two mixed-methods studies (lab sessions plus interviews) to tease apart three layers of trust:

  1. Trust in the search agent (Google vs ChatGPT)
  2. Trust in the health information they deliver
  3. How the interface modality changes that trust (text, speech, embodied)

Study 1: comparing agents—Google vs ChatGPT

  • Sample size: N=21
  • Design: within-subjects (each participant used both agents)
  • Tasks: 3 health question types (general; symptoms/causes; treatments)
  • Agent outputs:
    • Google: snippets + external website content (participants can click through)
    • ChatGPT: the model’s synthesized textual response only (no web browsing in the lab)

Participants rated:
- trust in the health info (credibility/reliability/believability)
- trust in the agent itself
- intention to use each agent again

Study 2: holding the agent constant—same LLM, different CUIs

  • Sample size: N=20
  • Agent backend: the same GPT-4o for all three conditions
  • Interfaces compared:
    1. Text-based CUI (web chat)
    2. Speech-based CUI (Echo Dot-style verbal interaction)
    3. Embodied CUI (physical body with a smartphone “face” and basic expressions, plus a speaker)

So Study 2 is the key “UI isolation” step: it asks, if the brain (LLM) is the same, why does trust still change?

Trust results: why ChatGPT beat Google (but not for the reasons you might expect)

Let’s start with the headline from Study 1:

Higher trust for ChatGPT than Google

Participants showed significantly higher trust in health information from ChatGPT than from Google:
- Mean trust in health info:
- ChatGPT: M=4.05 (SD=0.47)
- Google: M=3.77 (SD=0.64)
- The agent effect was statistically significant:
- F(1,20)=6.73, p=.017 (trust for ChatGPT > Google)
- Trust did not vary significantly by question type:
- F(2,40)=0.63, p=.480

And they also trusted the agent more:
- Paired difference in agent trust:
- t(21)=-2.53, p=.02, Cohen’s d=0.55

Here’s the important nuance: this preference wasn’t limited to one kind of health query. It looked stable across general questions, symptom/cause questions, and treatment questions.

A relationship that behaved differently for each agent (source dissociation)

The correlation results clarify something subtle about trust formation.

Relationship tested Google condition ChatGPT condition
Trust in information vs trust in agent Significant (r=0.63, p=.003) Not significant (r=0.09, p=.21)

Interpretation (in plain language):
With Google, users’ trust in the search tool and trust in what it delivers move together—likely because people have years of familiarity with how Google works. With ChatGPT, users may trust the conversational assistant (or feel more comfortable with it), but they still judge the actual content separately based on how it’s presented (logical flow, professional tone, confidence level, etc.). The paper calls this pattern source dissociation—trust in agent doesn’t reliably “carry over” to trust in the information.

What users said in interviews: “fast answers” help… but verification still matters

Participants described:
- ChatGPT as a quick, helpful starting point—especially when they don’t know enough to search effectively.
- Google as giving them more autonomy—they feel like they have their own “judge” because they can click, compare, and cross-check.

They also mentioned both trust boosters and trust breakers:
- Boosters: professional-yet-understandable language, good structure, clear confidence
- Breakers: uncertainty (“maybe” / “I’m not sure” reducing trust), or overconfidence seeming unrealistic, and lack of verifiable references

The “interface modality” twist: text wins, speech/embodiment can backfire

Now Study 2: what happens when the LLM stays the same and only the user interface changes?

Users trusted text-based delivery the most

Average trust in health information by interface:

  • Text-based: M=4.19 (SD=0.42)
  • Speech-based: M=4.13 (SD=0.46)
  • Embodied: M=4.00 (SD=0.48)

They also trusted the interface itself in the same order:
- Text-based interface trust: M=3.73
- Speech-based: M=3.57
- Embodied: M=3.56

But the real “design lesson” isn’t just that text ranked highest—it’s why trust dropped for the richer modalities.

Mixed Linear Model finding: embodied gets lower trust than text and speech

When the analysis controlled for usability and other factors, trust differences still showed up:

Comparison (interfaces) Reported effect p
Text-based vs embodied β=0.189 .035
Speech-based vs embodied β=0.180 .002

And usability was a major driver:
- usability predicted trust in information (β=0.187, p=.004)
- mediation analysis: usability fully mediated the interface→trust relationship
- significant indirect effect (β=0.154, p<.001)
- non-significant direct effect (β=0.098, p=.277)

Translation: richer interfaces didn’t fail because people hated AI—they stumbled because the modality made the interaction harder, less familiar, and more cognitively demanding in a high-stakes setting.

Why richer interfaces reduced trust: cognitive load, privacy anxiety, and authenticity skepticism

In the interviews, participants explained several mechanisms that line up with the numbers.

1) Usability and familiarity beat “naturalness”

Participants found text easier to process and easier to cross-check. With speech and embodied interfaces:
- you can’t easily “go back”
- vocal delivery can feel like you might miss something
- embodied interfaces add non-verbal cues you also have to interpret

This isn’t just preference—it’s a mental effort issue. Health decisions require careful thinking, so extra processing steps can erode confidence.

2) Speech/embodiment increased the chances of “trust-breaking” mismatches

The paper highlights a modality-context mismatch:
- Human-like text can increase trust by making complex health content readable.
- But when that human-likeness shows up via voice or embodiment, it can trigger concerns:
- privacy (“is it always listening?”)
- authenticity (“is this pretending to be human?”)
- skepticism if something seems off

3) Personalization is helpful—but it scares people when it feels like surveillance

Participants liked tailored answers (context-aware empathy; follow-up questioning; relevance). But they worried about what the system was collecting—especially when voice or embodied interaction made them feel observed.

So trust becomes conditional:
- Content credibility + interaction safety
- not just one or the other

Practical implications: how to design health conversational tools users will actually trust

This paper is useful because it doesn’t just measure trust—it extracts design recommendations grounded in what participants said and what the stats supported.

1) Bridge searching and verification inside the chat flow

In Study 1, users treated ChatGPT as a starting point and then wanted to verify. Many current tools don’t make that easy—they present synthesized claims as if they’re final.

Design direction:
- include clickable provenance (link claims to source material inside the interface)
- prompt verification cues for high-stakes questions (e.g., “this is complex—would you like guidelines?”)
- support user-driven follow-ups rather than replacing their judgment

2) Use “functional anthropomorphism,” not deceptive impersonation

Participants appreciated human-style communication patterns (empathy, clarity, conversational back-and-forth), but distrusted full “pretend-to-be-a-human” visuals.

Design direction:
- focus on conversational behavior that feels supportive
- avoid visual impersonation that triggers uncanny valley effects or identity confusion
- clearly label that the system is AI so expectations stay realistic

3) Make personalization privacy-preserving and visibly controllable

Personalization can build trust through relevance, but it can also destroy trust if users feel monitored.

Design direction:
- provide privacy controls beyond accept/reject (more granular)
- add “memory used” transparency cues
- offer modes like “don’t save this” for sensitive queries
- ensure privacy messaging matches the modality (voice/embodiment should come with extra clarity)

4) Don’t assume “more modality” = “more trust”

A counterintuitive takeaway: the richest interface wasn’t the most trusted. Text-based CUIs were trusted more because they matched usability and familiarity, enabling cross-referencing—critical for health.

Design direction:
- optimize usability first, then add modality where it clearly improves understanding (like well-timed summaries or optional visuals)
- treat voice/embodiment as assistive, not authoritative

If you want the original framing of these results, the full paper is here: https://arxiv.org/abs/2608.21177.

Key Takeaways

Key Takeaways

  • People trusted ChatGPT more than Google for health information in the lab: N=21, with a significant agent effect (F(1,20)=6.73, p=.017).
  • Trust in Google’s agent and Google’s information moved together (r=0.63, p=.003), but trust in ChatGPT’s agent didn’t strongly predict trust in its generated content (no significant correlation). This suggests source dissociation.
  • In Study 2 (N=20), with the same LLM backend, text-based interfaces were trusted more than speech-based and embodied ones, even though speech/embodiment felt more “natural.”
  • The biggest design driver was usability and familiarity—and usability fully mediated how interface type affected trust in health information.
  • Rich modalities (speech/embodiment) can lower trust via cognitive load, privacy anxiety, and authenticity skepticism—especially in high-stakes health contexts.
  • Practical design priorities for trustworthy health CUIs:
    • help users verify (provenance, guidelines prompts)
    • use supportive human-like conversation without deceptive “human impersonation”
    • make personalization privacy-preserving and transparent
    • treat modality as an add-on to usability, not a substitute for it

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Visual Prompt Building: How to Stop Writing Walls of Text

What Are LLM Tokens? The Complete 2026 Guide (Context, Cost & Multimodal)

In-Context Privacy Learning for Chatbots (Just-in-Time Tools)

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.