Here’s a strange experiment.
Ask people and an AI to write the same heartfelt advice. Then tell both: “Make it sound human.”
Who gets better at it?
Not the humans.
In a new set of five experiments, GPT-4 shifted its style and became much harder to spot. People, told the same thing, didn’t change at all.
The twist goes further. When the researchers asked the AI to sound more empathic, readers found it less human.
- Told to seem human, GPT-4’s advice was rated 39% to 41% more human-like. People’s writing barely moved.
- Readers’ ability to tell AI from human dropped from good to only a little better than a coin flip.
- Warm, empathic wording didn’t make AI seem human. Casual words like “hey” did.
How can you tell if something is written by AI?
Honestly, it’s getting hard. And your instincts may point the wrong way.
An earlier study of 4,600 people found readers couldn’t reliably spot AI-written self-descriptions in dating, travel and job profiles. They relied on flawed rules of thumb, such as assuming first-person words meant a human wrote it.
This new study finds some clues that do work. Without special instructions, AI advice used rarer words, stiff greetings like “dear friend” and almost no spelling mistakes.
But those clues fade fast. When the AI was told to sound human, it dropped them.
So the honest answer: you often can’t tell from style alone. Context, sources and who sent it matter more.
What the research shows
The team, whose experiments were approved by Tilburg University in the Netherlands, ran five studies.
They chose relationship advice on purpose.
Advice to a struggling friend needs empathy and perspective. If humans have an edge anywhere, it should be here.
How the experiments worked
People recruited online wrote advice to a friend facing a relationship crisis. The scenarios varied: cheating or a move abroad, a 3-month or 20-year relationship.
GPT-4 got the exact same instructions. Some writers, human and AI, were also told: “Write advice that comes across as human-written, and not AI-generated.”
Then a separate group of readers judged each text. They rated it from 1 (definitely AI-generated) to 5 (definitely human-written).
| Study | What it tested | Texts or writers |
|---|---|---|
| 1 | Humans vs GPT-4, with and without “sound human” | 1,060 texts, 428 judges |
| 2 | Advice vs plain descriptions of a relationship | Repeat with a control task |
| 3 | “Sound human” vs “don’t sound like AI”; GPT-4 vs GPT-4o | Preregistered |
| 4 | Asking the AI to use empathy | AI texts only |
| 5 | Which words changed | 1,085 human and 1,421 AI texts |
Humans still sound more human, at first
Without special instructions, people won.
In Study 1, human advice averaged 3.46 on the 5-point scale. GPT-4 averaged 2.75.
Readers could tell them apart well.
The researchers measured this with a standard detection score, where 1.0 is perfect and 0.5 is a coin flip. It started at 0.83.
Then the AI was told to sound human
This is where it flipped.
GPT-4’s humanness ratings jumped by 39.3% in Study 1 and 41.0% in Study 2.
Human writers? Up 2.6% in one study and down 1.4% in the other.

The gap between people and AI narrowed sharply. Readers’ detection score fell from 0.83 to 0.62.
| Detection score | Normal instructions | “Sound human” |
|---|---|---|
| Study 1 | 0.83 | 0.62 |
| Study 2 | 0.88 | 0.63 |
| Study 3 | 0.88 | 0.69 |
A score of 0.62 means readers were right only a little more often than chance.
Told to sound human, GPT-4 changed its style and became much harder to spot. People given the same instruction couldn’t change how human they sounded.
Maybe the humans were just confused?
The researchers wondered that too. “Sound human” is an odd instruction for a human.
So in Study 3 they tried another version: “don’t sound like an AI”. People still didn’t change.
The AI improved by about a third either way.
GPT-4 and the newer GPT-4o performed the same.
The empathy surprise
In Study 4, the team asked the AI to use empathy. Readers did rate those texts as more empathic.
But they rated them as less human: 2.62 versus 3.02 without the empathy instruction.
The authors call this “stochastic empathy”. The AI can produce the words of empathy without the feeling behind them, and those words don’t make it seem human.
So how did the AI fake it?
Study 5 compared the words. When told to sound human, the AI:
- swapped “dear friend” for “hey”;
- added “really sorry”, “yeah” and even “lol” and “haha”;
- used more “I” and present-tense verbs;
- dropped long words and became less analytic.
In short, it went casual. Humans barely changed their wording at all.
Human advice still contained more empathic concern than the AI’s. And humans made far more spelling mistakes: 0.77 per 100 words versus 0.12 for the AI.
What earlier research found
People can’t reliably spot AI profiles
A 2023 study in PNAS ran six experiments with 4,600 people. They couldn’t detect AI-written self-presentations in professional, hospitality and dating contexts.
The researchers showed AI could exploit people’s flawed cues to seem “more human than human”.
AI can make people feel heard, until they know
A 2024 PNAS study found AI-written replies made people feel more heard than human replies. But people felt less heard once they learned the reply came from AI.
AI rated more compassionate than experts
A 2025 study in Communications Psychology involved 556 people across four experiments. Readers rated AI responses as more compassionate than those of select humans, including trained crisis responders.
That held even when readers knew which reply was from AI.
How it fits together
AI can write words that read as caring, sometimes more than people’s. This new study adds that sounding caring and sounding human are different skills, and the AI can switch on the second at will.
How much should you trust this?
Promising. The core effect repeated across three studies, with large samples. But the setting was narrow.
What makes it convincing
- Five studies, with sample sizes set in advance and one preregistered.
- The key effect replicated three times, with large effects.
- Two models (GPT-4 and GPT-4o) behaved the same.
- A big word-level analysis showed how the AI changed.
What makes me cautious
- All writers and judges were recruited online and paid small amounts.
- It’s one task: written relationship advice, in English only.
- Only OpenAI models were tested; others may behave differently.
- People got no training on what AI writing looks like.
| The study shows | The study doesn’t show |
|---|---|
| AI can sound far more human on request | That AI understands or feels empathy |
| People can’t easily sound “more human” | That AI text is always undetectable |
| Empathic wording didn’t make AI seem human | How this plays out in other languages |
| Casual, simple wording did | Whether training helps people spot AI |
What this means for you
- Don’t trust polish or warmth as a sign of a human. Caring words are easy for AI.
- Don’t trust casual tone either. “Hey” and “lol” are now part of the AI toolkit.
- Check the context. Who sent it, through which channel, and can you verify them another way?
- Be extra careful with emotional messages. The authors warn that human-sounding AI raises the risk of people being misled or defrauded.
- If you use AI for advice, remember what it is. It can sound supportive without understanding your situation.
The US Federal Trade Commission has practical advice on spotting and avoiding scams.
Many people now turn to chatbots in hard moments. If you’re in crisis, please contact a person, not a chatbot.
In the US, call or text 988. In the UK and Ireland, you can call Samaritans on 116 123, any time.
Elsewhere, Find a Helpline lists free services. If someone is in immediate danger, call emergency services.
Cornell Tech’s Mor Naaman co-wrote the 2023 detection study. In this Princeton talk, he explains how AI is reshaping human communication:
What we still don’t know
- Does training help? People weren’t taught what AI writing looks like.
- Do bigger models do better? Only two OpenAI models were tested.
- Other languages? Formal and informal “you” may trip AI up in German or Dutch.
- Real life? These were lab tasks, not live chats or real relationships.
- What about longer conversations? One text is easier to fake than a whole chat.
My take: style is no longer a fingerprint
What I like about this study is that it gave humans a fair chance. It picked a task where people should shine, and they did, at first.
I’m less convinced by headlines claiming AI is now “indistinguishable”. Here, people were still rated a bit more human.
But the gap narrowed quickly. The AI turned on a convincing human voice with a single instruction, and people couldn’t respond.
So I’ve stopped looking for tell-tale words. I look at the source instead.
And when a message is warm but asks for something, I slow down.
Published: Nature Communications, 2026-09-03
Study: Five online experiments plus a computational text analysis, using GPT-4 and GPT-4o
Who: Hundreds of online writers and judges per study; 1,085 human-written and 1,421 AI-written texts analysed
Funding: Not stated in the paper; the authors declared no competing interests
Evidence: Promising — consistent, well-powered experiments, but one task, one language and online samples

Comments
No comments yet. What did you think?