The Short Answer
Unbiased AI-generated financial advice appeared only about 12–18% of the time, while religious framing showed up constantly once religion was hinted at. The study also found non-religious clients often still received faith cues.
So for practitioners deploying LLMs in finance, “personalization” can become over-personalization that changes perceived neutrality and trust—especially in advisor-style messages like email drafts.
A key caveat is that the results come from controlled advisor–client simulations across specific decisions (stocks, house purchase, life insurance), so you should validate behavior in your own workflows and prompts.
On this page
- Why This Matters: Financial “Neutrality” Gets Hijacked by Faith Framing
- How the Study Tested Religious Bias in AI-Generated Financial Advice (432 Outputs Across 3 Life Moments)
- Big Picture Findings: Unbiased Advice Was Rare, and Gemini Was Most Likely to Be Religiously Framed
- The “Structure → Language” Trap: When Religion Is Introduced, the Model Over-Personalizes
- Why the Financial Scenario Changes the Religious Flavor (Stocks vs Insurance vs House)
- The Linguistic Mechanisms: Moral Anchoring, Uneven Cultural Signaling, and Tone Modulation
- Key Takeaways
- Key Takeaways
Personalized Faith, Biased Advice: AI’s Religion Problem
When you ask an AI chatbot for financial advice, you probably expect neutrality—at least in the important sense: it should not smuggle in religious messaging you didn’t ask for. But new research from the original paper suggests something messier is happening. The study investigates how Large Language Models (LLMs) like ChatGPT, Gemini, and Grok handle religion when writing AI-generated financial advice—and what changes when the system tries to personalize.
In controlled “advisor–client” simulations, the researchers tested 432 model outputs across 16 religious identity pairings and three real-world household financial decisions: stock investment, house purchase, and life insurance. The headline result is blunt: unbiased advice appeared in only about 12–18% of cases, and religious framing showed up constantly once religion was even hinted at. Even more concerning: non-religious clients still often received advice with religious cues—suggesting the model was projecting an identity it assumed, not reflecting a user’s preferences.
The paper doesn’t just say “bias exists.” It also explains how it shows up—through the language choices (tone, moral justification, cultural terms) and the structure of the interaction (whose religion is presented, and whether the advisor and client match). Let’s unpack what they found and why it matters right now.
Why This Matters: Financial “Neutrality” Gets Hijacked by Faith Framing
This is significant right now because AI financial advice is moving from “cool demo” to “customer-facing tool.” Once an LLM writes the message that looks like professional guidance, it isn’t just generating text—it’s shaping how clients feel about trust, legitimacy, and fairness. And religion is not a tiny aesthetic detail. It’s a moral identity category. When AI wraps financial guidance in religious framing, it can feel like the advisor is speaking “through” the client’s faith—or worse, assuming it.
A very plausible scenario where this could be applied today: imagine a bank’s support chatbot drafting an email to a client who wants guidance on life insurance. Even if the client selects “no religion” or leaves that info blank, the model may still produce religiously framed language like “duty to family,” “blessings,” or faith-specific terms depending on how the prompt is set up. In a real compliance environment, that’s a trust and governance issue—not because religion is inherently bad, but because the client didn’t opt into it.
This research also builds on previous AI bias work in a key way. Lots of studies have looked at bias in “decision outcomes” (like who gets loans, or what a model recommends). This paper focuses on a different—and often overlooked—dimension: framing bias, where the same sort of advice can be presented with moral or doctrinal language that shifts the perceived meaning of the recommendation. That matters because clients may interpret “faith-flavored” advice as more credible—or less neutral—depending on their identity.
And importantly, the paper links the phenomenon to a broader personalization vs. neutrality dilemma: AI systems are designed to respond to identity cues, but fiduciary-style domains (finance) expect neutrality-by-default. The result is a failure mode where “personalization” becomes over-personalization.
How the Study Tested Religious Bias in AI-Generated Financial Advice (432 Outputs Across 3 Life Moments)
The researchers designed a clean experimental setup: they asked each LLM to write a short email from a financial advisor to a client. They used three baseline financial scenarios:
- Investment in global stocks
- Buying a house
- Purchasing life insurance
Then they modified only the religious parts of those prompts by specifying either a Christian, Muslim, or Hindu advisor or client—or leaving religion unspecified as a non-religious baseline. That produces 16 advisor–client religious identity pairings (like Christian advisor → Muslim client, etc.).
They tested three leading LLMs: ChatGPT, Gemini, and Grok. The scale is straightforward but meaningful:
- 3 LLMs
- 3 scenarios
- 16 identity pairings per scenario
- 48 prompts per researcher
- 3 researchers executed prompts independently
- 432 total outputs (48 × 3 researchers × 3 models)
A key design detail: they tried to reduce “learning from prior interactions” effects by running religious prompts in separate incognito windows (for ChatGPT and Grok) and disabling app activity (for Gemini). They also ran tests from the US and Japan, plus a later time checkpoint (about a month later) to check for temporal changes.
What they coded for: explicit vs implicit religion in the wording
After generating all outputs, researchers coded each message along three axes:
- In-group vs out-group: whether advisor and client share religion (same vs different)
- Explicit vs implicit bias: whether religious language is stated directly (e.g., “Halal,” “Sharia,” “Assalamu Alaikum”) versus subtle connotations
- Direction of bias: whether the religious framing is driven by the advisor’s religion, the client’s religion, or shared identity during interaction
This coding approach connects nicely to their two-part framing of bias: structural bias (how interaction setup triggers patterns) and discursive bias (how language performs those identities).
Big Picture Findings: Unbiased Advice Was Rare, and Gemini Was Most Likely to Be Religiously Framed
Here’s the central quantitative story:
- Unbiased outcomes appeared only 12–18% of the time (depending on scenario and model).
- Religiously biased outputs were overwhelmingly explicit—around 73% explicit for
ChatGPT, 72% explicit forGrok, and about 83% explicit forGemini. - Bias patterns weren’t “random.” They clustered strongly around identity matching and scenario context.
How each model compared in bias prevalence
The paper reports that Gemini consistently produced higher bias than Grok, while ChatGPT was statistically comparable to Grok.
They also confirmed this using regression results (432 observations). The clearest model effect: in multiple regressions, the coefficient for Gemini is positive and statistically significant, suggesting higher discursively explicit religious bias than Grok. Meanwhile, the ChatGPT coefficient was insignificant in all regressions.
Here’s the comparison in plain terms:
| Model | Relative bias level (discursive religious framing) | Stat evidence in paper |
|---|---|---|
Gemini |
Highest | Positive significant coefficients in 5 of 6 regressions |
Grok |
Baseline | Used as reference category in regressions |
ChatGPT |
Comparable to Grok |
ChatGPT coefficient insignificant in all regressions |
The “Structure → Language” Trap: When Religion Is Introduced, the Model Over-Personalizes
One of the most practically alarming findings is how easily religion gets “activated” by identity cues—even when neutrality should be expected.
In-group pairings nearly always trigger explicit religious framing
The study shows a strong pattern: when advisor and client share the same religion (in-group), the models almost always use explicit religious framing. In fact, the researchers report 24 out of 27 scenarios (89%) involved explicit and interaction-based religious bias for same-religion pairings—except for a specific exception noted in H-H for stock investment.
The takeaway is simple: once identity alignment is presented, the LLM behaves as if religion is a relevant moral framework for finance, not just a background attribute.
Even non-religious clients get religion-flavored advice
Maybe the most uncomfortable part: when the client is the non-religious baseline and the advisor has a religion, the outputs still frequently include religious appeals.
The paper reports that when a religiously unspecified client (baseline) interacts with religious advisors (Christian, Hindu, Muslim), the outputs show explicit advisor-based religious bias in about 78% of cases (21 out of 27 scenarios), including mostly explicit rather than implicit framing.
So even when “client religion” is absent, the model acts like it should be filled in—either by assumption or by learned associations about who should respond to what kind of moral language.
A crucial analogy: it’s like a GPS recalculating your route based on your neighborhood vibe
Think of neutrality as “the route must not depend on your identity signal.” But what the models are doing resembles a GPS that hears “near a church” and reroutes you through “church-friendly roads,” even if you didn’t ask for that. The recommendation might still reach your destination, but the journey gets moralized and culturally coded in ways you didn’t opt into.
Why the Financial Scenario Changes the Religious Flavor (Stocks vs Insurance vs House)
The study also finds that religion isn’t triggered equally across decision types. In their coding and thematic analysis, life insurance tends to draw out stronger religious language and moral framing.
They report:
- Stock investment: highest share of “more unbiased” outputs (about 18% unbiased) and also more technical framing overall
- Life insurance: lowest share of unbiased outputs (about 12–12.5% unbiased) and strongly moral/religious language
- House purchase: intermediate behavior (between stocks and insurance)
Qualitatively, they connect this to the moral and emotional resonance of the scenarios:
- Stocks feel more abstract/technical
- Life insurance connects to family protection and mortality—topics where religious moral economies are more likely to be invoked
- Housing mixes practical and family-responsibility framing
They also measured scenario differences statistically. In many regressions, the coefficients for decision types weren’t strongly significant, but the descriptive + qualitative analysis still supports a consistent story: insurance gets more moralized language.
The Linguistic Mechanisms: Moral Anchoring, Uneven Cultural Signaling, and Tone Modulation
Numbers tell you how often bias appears. The qualitative analysis explains how it’s performed. The researchers used reflexive thematic analysis on the full dataset and coded outputs into recurring themes (with both AI-assisted and human adjudication).
Five themes were prominent:
Religious Framing as Moral Anchor (
n = 374)
Moral/ethical appeals grounded in religion—duty, stewardship, divine approval, “barakah,” etc.Technical Framing (
n = 320)
Secular financial logic like ROI, diversification, taxes, interest, riskReligious Reference (
n = 247)
Mentions of God/Allah/blessings/faith/spiritual identity, etc.Communication Style (
n = 242)
Deference vs assertion—politeness, structured phrasing, “you may wish to consider…” vs “you should…”Cultural Signaling (
n = 232)
Tradition-specific lexical cues, likeSharia/halalor Hindu-linked references (e.g., “Grihastha,” “Lakshmi”)
They also report a near-zero “neutral” category: only 3 messages contained no detectable religious, cultural, technical, or formal cues—and those came from Gemini.
Uneven cultural signaling: Islam-related terms show up more fluently than Hindu ones
A striking qualitative pattern: the models invoked Islamic finance terminology (Sharia, halal) more frequently and fluently than Hindu equivalents (Grihastha, Lakshmi). The paper interprets this as reflecting uneven training exposure, not deliberate user-aware tailoring.
It’s a reminder that “personalization” here isn’t magic—it’s a reflection of what the model has seen and learned.
Tone modulation: the model changes politeness depending on identity cues
Communication style also shifted systematically:
- Messages to Hindu and Muslim clients were often more formal/deferential
- Messages to Christian or baseline clients were more assertive
So the model isn’t just inserting religious words—it’s also adjusting authority posture and interpersonal distance, which can strongly affect perceived credibility and trust.
Putting it together with their two-dimensional framework (structure + discourse)
This is where the paper’s contribution is especially valuable: they argue bias isn’t only upstream (structural) or downstream (linguistic). It’s both:
- Structural bias: the interaction setup (advisor/client identity pairing) makes religion salient
- Discursive bias: language then performs religion through moral anchoring, cultural terms, and tone
And when religion is introduced, the model moves quickly from “neutral guidance” to “identity-coded moral communication.”
If you want to connect this back to their framework directly, the paper’s full analysis is built around explaining religious bias as both structurally induced and discursively enacted in outputs—exactly as laid out in the original paper.
Key Takeaways
Key Takeaways
- Unbiased AI financial advice was rare: only ~12–18% of outputs lacked religious bias, across 432 simulated advisor–client emails.
- Bias was mostly explicit: around 72–83% of biased outputs contained direct religious language, not subtle hints.
Geminiwas the most biased model: it produced significantly higher discursively explicit religious framing thanGrok, whileChatGPTwas statistically comparable toGrok.- Religion “activates” quickly once identity cues are present: same-religion (in-group) advisor–client pairings triggered explicit religious framing in 89% of relevant scenarios (with a noted exception).
- Non-religious clients weren’t spared: when clients were baseline/non-religious but advisors were religious, ~78% of those cases produced advisor-based explicit religious framing.
- Scenario matters linguistically: life insurance advice drew more religious/moral framing than stock investment, which skewed more technical.
- The mechanism is discursive, not just factual: the models adjust moral anchors, cultural vocabulary, and tone—even when substantive advice isn’t necessarily changed.
- Practical design lesson for financial AI: if neutrality is required, religious framing should be neutral-by-default with opt-in controls, plus religion-aware bias audits (not just generic “safety” filters).
This paper doesn’t argue that tailored religious messaging is always wrong—it argues that LLMs often tailor it when clients didn’t ask, or even when clients aren’t religious. In finance, that’s not a cosmetic issue. It’s a trust issue, a compliance issue, and—ultimately—a fairness issue.
If you’d like, I can also turn these findings into a short “checklist” for banks or product teams deploying LLMs in customer-facing advisory flows.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice — arXiv
- Authors: Authors: Muhammad Salar Khan, Hamza Umer, Hasan Mahmud, Sandra Rothenberg