The Short Answer
ChatGPT use in the Global South goes far beyond Q&A: across India, Nigeria, Brazil, and Pakistan, expressive self-reflection grows to roughly a fifth of conversations by the end of the study window. The same assistant is also used for tutoring, translation, health explanations, current affairs, and religious questions—depending on the country.
For practitioners, this means you can’t rely on “typical user” metrics built around work productivity or information-only flows. If your product or policy assumes mainly task completion, you’ll miss common real-world use where people share personal context and seek advice in the same conversation.
The limitation is that this evidence is based on exported histories from a specific set of 1,252 users (not a full population snapshot), so you should validate locally. Still, the core takeaway holds: conversation-level, country-sensitive measurement is essential for accurate safety evaluation.
On this page
- Introduction: ChatGPT use isn’t one-size-fits-all
- Why This Matters: Policy, product design, and safety are all betting on the wrong “typical user”
- What the researchers actually measured: Purpose, topics, and intent—at real conversation scale
- Purpose: Most use is personal—and coursework is nearly as common as work
- Topics: The same ChatGPT gets pulled into different local needs
- Intent (how people engage): Asking declines slowly, Doing stays flat, Expressing grows fast
- Why this changes safety, evaluation, and product design
- Key Takeaways
Conversational AI in the Global South: What ChatGPT Really Gets Used For
Introduction: ChatGPT use isn’t one-size-fits-all
If you think you already know how people use ChatGPT—ask a question, get an answer—this new research will nudge your brain a little. Instead of looking at platform-wide averages, the authors dug into conversation-level evidence: complete ChatGPT exports from real users in India, Nigeria, Brazil, and Pakistan, paired with self-reported age and gender. That means we get to see what people talk about, why they’re using the assistant, and how they engage over time.
This blog is based on new research from the original paper by Shreyasi Roy Chowdhury and Kiran Garimella, who analyze 202,590 conversations from 1,252 users across a multi-year window (Dec 2022–Feb 2026). The study’s big premise is simple but important: the public picture we get from the companies running LLMs is mostly aggregate reporting and fixed taxonomies—useful, but not designed to answer “What does adoption mean in local contexts?” The researchers want the missing ingredient: full conversations you can’t reconstruct from summary stats.
And the twist is this: the same tool shows up as a tutor, a translator, a health explainer, a career coach, and—growing over time—a place where people do something older tools couldn’t do well at scale: express feelings, vulnerability, and self-reflection alongside requests for help.
Why This Matters: Policy, product design, and safety are all betting on the wrong “typical user”
Here’s my take on why this research is especially significant right now: AI companies and governments are moving fast on measurement, regulation, and product defaults—but most of the measurement is still optimized for a world where “AI use” looks like work productivity and English-first interaction. If your metrics don’t capture what people actually do, your safeguards and incentives will be misaligned.
A concrete “apply this today” scenario: imagine a regulator or a safety team tasked with evaluating a chatbot feature meant for “emotional support.” If your tests assume users mainly do either (1) information seeking or (2) task completion, you’re likely missing the common reality described in this paper—conversations where users disclose personal context (health worries, religion, relationships, money stress) and then ask for advice or action in the same flow. The paper finds that expressive conversations grow to roughly a fifth of all conversations by the end of the observation window—and these expressive exchanges are often not in the country’s dominant language. That’s a safety and evaluation problem, not a niche research detail.
This study also builds on previous AI measurement research, but it corrects a blind spot: earlier work has leaned heavily on either (a) platform-level reporting that outsiders can’t re-analyze or (b) public conversation dumps that skew toward unusual user populations. Here, the researchers use data donation with privacy-preserving export pipelines and add self-reported demographics—so we can actually ask: who is doing what, and how does it evolve? For anyone building policy, product, or risk frameworks, that’s the difference between “we guessed” and “we observed.”
What the researchers actually measured: Purpose, topics, and intent—at real conversation scale
This paper’s design is doing something that’s hard to replicate: it’s conversation-level, demographically grounded, multi-country, and spans more than three years.
The dataset: 1,252 users, 202,590 conversations, four countries
Participants were recruited through Clickworker and asked to export their full ChatGPT conversation history from OpenAI’s built-in export feature. A client-side script removed personal identifiers before upload. Users also provided self-reported age and gender.
Key dataset facts from the paper:
- 1,252 users
- 202,590 unique conversations
- Countries: India, Nigeria, Brazil, Pakistan
- Time range: Dec 2022–Feb 2026
- Sample is convenience-based and includes many digitally literate, micro-work-adjacent participants (so it describes the sampled population, not the full national population).
The measurement: three axes, two ways of finding topics
Each conversation is labeled along three dimensions:
- Purpose: Work vs coursework vs personal
- Topics: What people talk about
- Mode of interaction (“intent”): How they use the model—basically, are they asking, doing, or expressing?
To keep results comparable to platform reports, the authors use:
- OpenAI’s 24-category topic taxonomy (so numbers can align with published aggregates)
- Platform intent/purpose classifiers from the broader LLM measurement literature
But to detect local uses that global categories mash together, they also run:
- Unsupervised topic discovery (BERTopic-like clustering) separately per country
This pairing—fixed taxonomy + bottom-up discovery—is one of the paper’s smartest moves, because it makes local patterns visible without losing comparability. If you want to see how the authors operationalize this, the original paper is the place to go: https://arxiv.org/abs/2609.38279.
Methods comparison (the “why should I trust this?” part)
Here’s a high-level comparison of what this study does versus common alternatives:
| Approach | What you see | Biggest limitation |
|---|---|---|
| Platform aggregate reports (company-run) | Large-scale stats with fixed categories | Hard for outsiders to re-analyze; may hide local meaning inside coarse buckets |
| Public datasets (shared logs) | Real conversations | Often self-selected toward unusual users or technical interfaces |
| This study (data donation, conversation exports) | Full conversation history + age/gender + multi-country | Convenience sample; not fully representative of entire countries |
Purpose: Most use is personal—and coursework is nearly as common as work
Let’s start with “why are people even opening ChatGPT?”
Across all four countries, personal use dominates, and coursework shows up about as often as work—which already challenges the workplace-only narrative you hear in many AI economics discussions.
Personal use is the majority everywhere
Using the paper’s purpose classifier, personal use shares are:
- India: 61.5% of conversations
- Brazil: 63.7%
- Nigeria: 55.0%
- Pakistan: 55.4%
And the personal-vs-nonpersonal balance is not a one-time thing. The authors report that the personal share is steadily increasing in all countries over the observation window.
Work is real, but it’s not the headline
Work conversations are roughly a fifth overall in these samples. That’s meaningfully lower than global platform aggregates that often describe higher work shares (the authors note OpenAI-reported global work shares around 27%, and the sample’s demographics likely shift the observed mix because participants skew younger and include many students).
Coursework is right alongside work
Coursework appears about as common as work across the sample. The paper emphasizes two interpretations:
1. ChatGPT may be a low-friction substitute when students face barriers to academic help elsewhere.
2. It may interact with existing education inequality, either compensating for missed resources or reinforcing gaps.
Importantly, this study can tell you what people use it for—not whether the help improves outcomes. That “accuracy and welfare” part remains open.
Gender and age patterns show up differently by country
The study also finds interesting demographic splits (and sometimes only some survive stricter statistical correction). For example:
- Women show higher coursework shares in India, Pakistan, and Nigeria, but not in Brazil.
- Age gradients broadly make sense: younger users tilt more toward coursework and education themes, older users tilt more toward work/career and civic topics.
Even when differences aren’t statistically rock-solid across all tests, the conversation-level lens makes the patterns easier to interpret than name-inferred demographics.
Topics: The same ChatGPT gets pulled into different local needs
If you only looked at fixed taxonomy labels, you might miss what’s actually happening. The paper explicitly shows that global topic buckets compress local meaning.
Under the OpenAI taxonomy: you get a familiar “global” picture
Using OpenAI’s 24-category taxonomy, the broad buckets look roughly similar across countries (and roughly align with published global breakdowns, with some sample-specific shifts). Common categories include:
- Practical Guidance
- Seeking Information
- Writing
- Technical Help
The sample leans somewhat more toward information seeking and writing than technical help.
Unsupervised topic discovery reveals “what the taxonomy hides”
Here’s where the paper becomes really interesting. The authors run unsupervised clustering per country and find clusters that don’t map neatly into the fixed global taxonomy.
Some standout country-specific uses:
- India: health and wellness is the largest cluster (about 8.0% of conversations)
- Brazil: health and wellness is also a top cluster (about 9.2%), and self-reflection/emotional conversations show up as a top theme
- Pakistan: Urdu–English translation is the largest cluster (about 7.2%)
- Nigeria: current affairs is the second-largest theme, and religious questions are prominent
They also report that certain themes are essentially “local-only” in these top lists:
- Religion appears strongly in Nigeria and Pakistan, and is not in India/Brazil’s top-15 lists.
- Online earning clusters show up strongly in India and Pakistan, and also in parts of Brazil.
- Self-reflection/emotion-focused clusters are notably visible in Brazil.
Theme-level comparison: local needs vs global categories
The authors aggregate per-country clusters into ten cross-country themes, and the differences pop.
Here’s a simplified view of the theme emphasis:
- Finance/Earning: heavy in India (17.7%) and Brazil (11.6%), nearly absent in Nigeria (3.3%)
- Religion: concentrated in Nigeria and Pakistan
- Translation/Language: highest in Pakistan
This suggests a core point: the “same product” is being attached to different local value gaps—health access uncertainty, multilingual navigation needs, religious practice explanation, and online economic strategies.
Intent (how people engage): Asking declines slowly, Doing stays flat, Expressing grows fast
Now for the most “future-facing” part: not just what people ask, but how they use the chatbot in conversation.
The paper uses an Asking/Doing/Expressing intent framework:
- Asking = seeking information / decision support
- Doing = asking the model to execute a task
- Expressing = reflecting, emotional communication, or sharing personal context
Asking is still dominant—but it’s slowly declining
Across all countries, Asking remains the largest intent category throughout the study window, but its share declines only modestly.
Doing doesn’t really grow
Task delegation (“Doing”) does not expand the way some people might expect as chatbots get better. It stays roughly stable (with Nigeria showing a mild increase reported by the authors).
Expressing grows to ~one fifth of conversations
The big trend is Expressing:
- By the end of the observation window, Expressing accounts for roughly a fifth of conversations or more in every country.
- This is not just a compositional effect (new users arriving). The authors find within-user increases too.
In their within-user analysis, focusing on users with enough history, Expressing share rises by about:
- +13.7 percentage points overall
- In every country, and for 81% of users individually, Expressing increases from early active months to later months.
Meanwhile Asking and Doing decline within users.
The “hybrid disclosure + request” pattern
A key qualitative finding: expressive conversations usually aren’t pure venting. Instead, users often:
1. Disclose personal context (health symptoms, relationship stakes, religious concerns, money stress)
2. Then ask for advice, explanation, or next steps
The authors give examples like:
- Describing symptoms before asking what a medication does
- Sharing relationship details and asking whether the age gap is acceptable
- Pasting religious text and asking for interpretation while tying it to health constraints
So conversational AI isn’t only giving “answers.” It’s becoming a place where people attach their lives to a request.
Why this changes safety, evaluation, and product design
If you’re building evaluation frameworks, this paper is basically a warning label.
Separate safety categories, miss the overlap
Many safety pipelines treat emotional support and task execution as different domains. But in the expressive mode described here, emotional disclosure and concrete action requests overlap. That makes English-only or “emotion vs completion” separated evaluation likely to miss the typical user reality in these markets.
Language and multilingual behavior matter
The paper finds expressive conversations are less likely to happen in the dominant national language than Doing conversations. For example (share speaking the dominant language):
- India: 68.7% for Expressing vs 88.6% for Doing
- Nigeria: 85.4% vs 96.2%
- Brazil: 76.3% vs 85.2%
- Pakistan: 73.9% vs 86.7%
They also check whether this is just topic selection and find some differences are topic-driven—but language switching still matters in the cases where it changes by mode (like greetings in Nigeria and emotional struggles/self-worth in India).
Tooling and pricing defaults should consider “household production”
Economics research often frames value as workplace productivity. This paper’s conversation data suggests a lot of value is in household production: health understanding, education support, translation, and life advice—things that don’t show up cleanly in labor market metrics.
This changes what matters for adoption:
- If product teams only optimize workflows for office productivity, they miss what users come to ChatGPT for in these contexts.
- If policymakers only regulate around workplace use cases, they may miss safety risks associated with intimate disclosure.
Key Takeaways
- Personal use dominates: across India, Brazil, Nigeria, and Pakistan, personal conversations are the majority (~55% to ~64% depending on country).
- Coursework is nearly as common as work: workplace productivity is real, but it’s not the core story in these conversation logs.
- Fixed global taxonomies hide local meaning: unsupervised topic discovery surfaces country-specific uses like health/wellness (India & Brazil), Urdu–English translation (Pakistan), and religious questions (Nigeria & Pakistan).
- Intent shifts over time toward Expressing: Asking declines modestly, Doing stays roughly flat, and Expressing grows to about a fifth or more of conversations.
- Expressing usually blends disclosure with requests: people share personal context and then ask for advice, explanations, or next steps.
- Safety and evaluation need to be multilingual and overlap-aware: expressive, disclosure-heavy conversations are less likely to be in the dominant language and often combine emotional and task-oriented needs in one flow.
- Adoption means different things locally: the same chatbot interface fills different value gaps—education support, health explainer, translation bridge, career/accounting assistant, and emotional reflection space.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- How People Use ChatGPT: Conversation-Level Evidence from India, Nigeria, Brazil, and Pakistan — arXiv
- Authors: Authors: Shreyasi Roy Chowdhury, Kiran Garimella