The Short Answer
AI can reproduce the surface signals of cooperation—morality language and politeness—without the underlying social mechanisms that create mutual accommodation. The result is that “nice” responses don’t reliably produce turn-to-turn alignment between speakers.
Practically, this means you should evaluate AI dialogue with multi-turn alignment and coordination metrics, not only single-turn “warmth” or “safety-sounding” quality.
Caveat: the paper’s findings emphasize trajectory-level cooperation; focusing only on turn-level polish can lead to over-reliance even when individual responses look socially correct.
On this page
- Introduction: When your AI “gets” manners, but not the point
- Why This Matters: We’re evaluating “niceness,” not cooperation—and that’s costly
- What the Researchers Actually Measured: Morality, politeness, and alignment in the same frame
- Who carries the social load: AI becomes the moral warmth, humans do the rest
- How cooperation unfolds over time: humans stabilize, AI drifts downward
- The biggest shock: cooperative mechanisms reverse direction with AI
- What designers and evaluators should do next: optimize for trajectory, not just turn quality
- Key Takeaways
Talking “Nice” to the AI Doesn’t Mean It’s Cooperative—Here’s Why
Introduction: When your AI “gets” manners, but not the point
Ever had a chat with an AI that sounds socially perfect—warm, polite, morally reassuring—yet somehow the conversation still feels off? Like you’re doing all the real social work while the system just “keeps up appearances”? New research from the paper on arXiv: 2609.21401 digs into exactly that tension: whether human–AI dialogue actually behaves like human–human cooperation, or just mimics its surface.
The study, “Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue,” looks at three layers of cooperative conversation: morality (what values get expressed), politeness (how face and social risk are managed), and alignment (how much speakers converge turn-to-turn). Then it asks a blunt question: when an AI nails these signals, is it really participating in the mutual dance—or is it only broadcasting the right moves?
To test this, the authors analyze 15,881 human–ChatGPT dialogues and 10,784 human–human dialogues, using multi-turn corpora and mixed-effects models. The results are surprising: the AI reproduces the signals of cooperation without the underlying social mechanism that makes cooperation stick between humans. Even more striking, some “cooperative” conversational tactics flip direction when the other participant is an AI.
Why This Matters: We’re evaluating “niceness,” not cooperation—and that’s costly
This research is significant right now because AI systems are no longer confined to “assistive” roles. They’re increasingly used in settings where conversation is the interface to judgment: workplace communication, tutoring, therapy-adjacent support, planning, and decision-making. In these environments, we don’t just want fluent text—we want mutual coordination: the sense that the other party is tracking us, adapting, and staying cooperative over time.
Here’s a concrete scenario you can recognize immediately: imagine using ChatGPT to draft an email after a tense conversation with a teammate. You ask for a reply that’s firm but not hostile. The AI sounds thoughtful, uses hedges (“might,” “consider”), expresses values (“fairness,” “respect”), and stays polite. But when you keep iterating—clarifying the tone, adding context, correcting assumptions—something subtle happens: alignment keeps slipping. The system might remain “warm,” yet the interaction doesn’t settle into the kind of shared rhythm you’d expect with a human collaborator. You end up doing more of the emotional and rhetorical steering yourself.
This builds on earlier AI research in a messy-but-important way. Lots of prior work focuses on single-turn outcomes: correctness, safety, bias, harmlessness, or whether the response matches human preferences. But this paper argues that cooperation is a trajectory-level phenomenon. You can score “good” outputs and still get “bad” interaction dynamics. So if evaluation and training rely too heavily on turn-by-turn signals—or on what “sounds like” alignment—you might accidentally optimize for the wrong thing.
The takeaway from the study is uncomfortable: the AI can look socially competent while the conversation stops being genuinely cooperative. And that means evaluation approaches that focus on surface polish may inadvertently encourage over-reliance.
What the Researchers Actually Measured: Morality, politeness, and alignment in the same frame
To compare human–AI dialogue with human–human dialogue, the authors measure the same constructs in both settings so the results aren’t just artifacts of how the corpora were collected.
Morality: which values show up in a turn?
They use Moral Foundations Theory (MFT) with five foundations—care, fairness, loyalty, authority, sanctity—and score each turn with ME2-BERT, producing a 5-dimensional moral signal per turn.
From those scores, they also derive:
- moral intensity: total moral strength (summing the foundation salience)
- moral entropy: whether moral language is concentrated in one foundation vs spread across many
So morality here isn’t “did the AI sound ethical?” It’s “how much and which moral foundations are expressed in each message.”
Politeness: which social tactics appear in a turn?
For politeness, they extract 32 marker types using a politeness toolkit (grounded in classic politeness theory). They convert each marker into a length-normalized rate (count divided by turn length).
Then they group markers into four functional dimensions:
- modulation (softening and indirectness: hedges, “please,” apologies, gratitude, indirect requests)
- agency (who controls the interaction: asking/giving agency, person framing)
- affect (emotion tone, including positive/negative affect and swearing)
- structure (discourse organization: questions, reasoning, greetings/titles, conjunctions/negation, etc.)
Alignment: does the next turn converge?
Alignment is measured as cross-turn accommodation: how similar the two turns are from each speaker to the next.
They compute similarity across four channels (each scaled to [0,1]):
- lexical: vocabulary overlap
- syntactic: style matching via function-word patterns
- pragmatic: overlap in rapport markers (hedges, please, gratitude, apology, affirmation)
- emotional: sentiment convergence using a sentiment score
They then average the four channels into an overall accommodation score.
The datasets: big enough to see conversation-level patterns
They analyze:
- WildChat for human–ChatGPT: about 1M multi-turn conversations, but filtered down to 15,881 dialogues after cleaning (English-only, ≥7 turns, balanced structure, removed duplicates/malformed cases).
- Topical-Chat for human–human: down to 10,784 dialogues after similar preprocessing.
The filtering leaves:
- human–AI: 473,962 messages
- human–human: 227,704 messages
And crucially: the human–AI corpus spans multiple model generations (e.g., GPT-3.5, GPT-4o, GPT-4.1 Mini), which lets them test whether “newer is better” for cooperation.
Who carries the social load: AI becomes the moral warmth, humans do the rest
One of the clearest differences in the findings is how morality and politeness are distributed between interlocutors.
Moral foundations: AI talks moral—humans share it
In human–AI conversations, the assistant produces significantly more moral content than the paired user across all five foundations. The largest asymmetries include:
- loyalty: Z=60.92, effect size r=.483
- authority: Z=56.79, effect size r=.451
In human–human conversations, the moral load is closer to symmetric—small differences, sometimes not statistically significant. Put plainly: humans avoid imposing binding moral frames on each other, while AI often “takes the moral steering wheel.”
Politeness: AI is warm; humans are socially cautious
Politeness shows a similar asymmetry:
- In human–AI, the assistant carries more affective warmth (emotion/affiliation), with Z=59.55, r=.473.
- But humans carry more of the modulation (face-sensitive softening strategies). The assistant does less of this interpersonal calibration than the user.
They also find the assistant dominates affect, while humans dominate structure (discourse organization), with a contrast emerging where human dyads distribute structure but human–AI shifts the organizational labor largely to the human partner.
Why this matters: “warm and moral” ≠ “mutually responsive”
It’s tempting to interpret warmth and moral language as signs of cooperation. But the study’s bigger argument is that these are surface signals. In human–human dialogue, warmth and modulation are tightly tied to reciprocal social risk management. In human–AI, the warmth arrives without the face sensitivity humans show—and that mismatch shows up later in alignment behavior.
How cooperation unfolds over time: humans stabilize, AI drifts downward
If you’ve ever felt a conversation with an AI “get worse” the longer it goes, this part explains why that intuition might be real—not just annoyance.
The authors track how moral intensity and overall accommodation change across the normalized timeline of a conversation.
Moral intensity: AI stays flat, humans build
- For AI, moral intensity is mostly flat: it starts high (around
~0.40) and stays on a plateau, not really “adjusting” in response to the user. - For humans, moral intensity builds through the middle third of conversations, consistent with moral engagement being negotiated and reciprocally calibrated.
So instead of morality evolving with the relationship, AI morality looks pre-configured.
Accommodation: both decline early, but only humans recover
Both conversation types show a steep drop in accommodation in the first ~10–20%, which the authors interpret as an initial calibration phase. Then:
- human–human conversations stabilize into an equilibrium.
- human–AI conversations keep declining, without recovery.
In other words: humans learn how to meet each other; the AI exchange doesn’t settle into shared ground.
The “who drifts” detail is especially telling
They report turn-level alignment declines per speaker role across channels. Both corpora show declines, but the assistant declines more steeply and consistently.
On emotional alignment, in particular, the AI’s decline is sharp:
- Z=-21.06, r=-.167 (assistant)
- human–human agents are stable on this dimension (Z=2.16, non-significant)
So the channel that anchors human accommodation is the one the AI lets go of fastest.
The biggest shock: cooperative mechanisms reverse direction with AI
Here’s where the paper gets most provocative. The authors don’t just describe differences—they test which features predict next-turn alignment. And they find that some features behave like cooperation signals in humans… but become anti-cooperative signals with AI.
They fit mixed-effects regressions predicting next-turn alignment from moral and politeness properties of the current turn, controlling for prior alignment. Models run separately for:
- assistant turns in human–AI
- agent-2 turns in human–human
Mechanisms that flip sign
A few key reversals:
| Feature (politeness or moral framing) | Human–Human effect on next-turn alignment | Human–AI effect on next-turn alignment | What it suggests |
|---|---|---|---|
| Modulation (hedging/softening) | β=+0.303 |
β=-0.470 |
Softening boosts cooperation between humans, but reduces alignment when produced by AI |
| Sanctity framing (“purity”/degradation) | β=-0.113 |
β=+0.160 |
Humans diverge when sanctity is invoked; users converge toward AI’s sanctity frame |
| Structure (organizational/discourse form) | β=-0.131 |
β=-0.473 |
Structure reduces alignment in both, but much more strongly with AI |
| Agency (giving the other party room) | β=+0.204 |
β=+0.416 |
Sharing control supports cooperation in both settings |
Moral assertiveness vs cooperation: less morality doesn’t fix alignment
They also compare across model generations. Moral intensity decreases across generations (highest in GPT-3.5, lowest in GPT-4.1 Mini), consistent with RLHF iterations reducing assertiveness. But cooperation doesn’t improve in step.
In fact:
- GPT-3.5 shows the strongest alignment trajectory (~0.38)
- while newer models decline more steeply despite being less morally assertive
So “less moral assertiveness” doesn’t automatically yield “better cooperation.”
Path dependence: once it drifts, it stays drifting
Another subtle but crucial result: prior alignment predicts future alignment about 9× more strongly with AI than with humans.
That aligns with the earlier drift observation: human–human exchanges re-equilibrate after low-accommodation turns; human–AI exchanges look more “sticky”—the early level set persists.
What designers and evaluators should do next: optimize for trajectory, not just turn quality
The paper doesn’t just diagnose—it gives a direction for how to think differently about evaluation and design.
1) Stop treating “polite and moral” as proof of cooperation
If hedges and softening are cooperation signals between humans, you might think encouraging those signals in AI would help. But the results suggest the same markers can act in reverse when the AI produces them.
So design work shouldn’t assume that copying human politeness behavior will recreate human social logic.
2) Test “agency” as a cooperative control variable
The one feature that reliably supports alignment is agency—giving users room to shape the exchange. This holds in both interaction types, with stronger effects in human–AI (β=+0.416 vs β=+0.204).
That means a practical design lever could be: increase conversational control for the user, make it easy for them to set tone/constraints, and let the AI actively follow user-led framing rather than imposing its own.
3) Evaluate interaction-level trajectories, not single responses
This is the evaluation punchline. The study argues that judging AI turn-by-turn can be actively misleading: per-output criteria (helpfulness, harmlessness, fluency, etc.) may correlate with “sounds good” while masking “interaction goes bad.”
This directly connects back to the paper’s framing (using the Machine-Integrated Relational Adaptation (MIRA) perspective): AI can simulate linguistic reciprocity, but without true mutual adaptation, the conversation may never mature into stable shared ground.
4) Build experiments that swap roles and isolate grounding
The authors note limitations (like GPT-family specificity and dataset differences, such as knowledge grounding and role asymmetry). They suggest matched-elicitation studies—shared prompts, optional grounding, role-swapped within-subject designs—to tease apart whether these effects generalize.
Key Takeaways
- AI can produce the surface of cooperation without the social machinery underneath. In human–AI dialogue, moral and politeness signals show up, but turn-to-turn alignment behaves differently than in human–human conversation.
- Morality and warmth are one-sided in human–AI. The assistant carries significantly more moral content across all five foundations and dominates affective warmth, while humans handle more modulation/organizational work.
- Human accommodation stabilizes; AI accommodation drifts downward. Humans reach an equilibrium after early calibration; human–AI exchanges keep declining without recovery.
- Some “cooperative” tactics reverse direction with AI. Features like hedging/softening (modulation) and sanctity framing increase or decrease alignment depending on whether the AI or a human produced them.
- Agency is the most consistent predictor. Giving users control supports alignment in both human–human and human–AI interactions.
- Evaluation should shift from turn-level to trajectory-level. Optimizing for single-turn “goodness” can hide worse cross-turn cooperation—and may even train systems toward the wrong social behavior.
If you want, I can also turn this into a practical checklist for designing (or testing) conversational agents so they maintain alignment over time—without relying on superficial politeness signals alone.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue — arXiv
- Authors: Authors: Marina Mitiaeva, Lu Xiao