Why AI Can Never Be Truly Creative
#055: Human emotion drives creative divergence and empathy is what makes it work. A machine can imitate both but possesses neither.
“Let the machine raise the floor. Use it for the drafts and the scaffolding, the version of the work that anyone could have made. Do not let it set the ceiling, because the ceiling is the part only your particular life can supply. “
Ask an LLM for unusual uses for a fork and you will get something clever. A back scratcher. A tuning-fork stand-in, if you squint. Ask a hundred models, running the same prompt over and over, and something odd happens: they mostly say the same clever things. Then ask a hundred people. You get a hundred directions, some silly, some strange, a few that only make sense once you hear the story behind them.
That contrast is the finding of a recent study by Emily Wenger and Yoed Kenett in PNAS Nexus, which is that LLMs are homogeneously creative. On its own, a single machine answer often rates as creative as the average person’s, sometimes higher. In aggregate, the machines collapse toward each other. Their responses resemble one another far more than ours resemble one another, and the pattern held across every major model tested, which suggests it is baked into how the technology works rather than a quirk of one company. Push the randomness up and the variety returns, right until the output tips into gibberish.
Emotion is where the divergence starts
Researchers have a tidy definition of creativity that has held up for decades. A creative thing has to be both new and useful. Novel but useless is just noise. Useful but familiar is just competent. You need both.
The novelty half runs on feeling. A line of work called the dual pathway to creativity found that activating emotions, whether pleasant or unpleasant, widen and deepen the paths the mind will take. Excitement sends you searching more broadly. Frustration makes you dig more stubbornly. The states that produce the least are the flat ones, the moods where nothing in particular is pulling at you. Emotion is the thing that knocks you off the well-worn path. And because your feelings are wired to your own memories and griefs and small obsessions, they send you somewhere no one else would go. Your variance, statistically speaking, is your biography.
A model has no biography. It was trained on the aggregate of what all of us have already written, so it answers from the crowded center of that space. That is not a flaw in the engineering; it is the engineering. When you draw from the middle of everything, you land in the middle of everything.
Empathy is what makes it land
The other half of the definition, usefulness, has its own human source, and it is the one closest to the work many of us do. The organizational psychologists Adam Grant and James Berry ran a set of studies where people prompted to think about the person they were creating for produced ideas that were original and useful, not just one or the other. The mechanism was perspective-taking. Picturing someone specific, and caring how it goes for them, is what separates the merely novel idea from the one that will help. Empathy is not the soft part of making things. It is the part that supplies the second half of the definition.
So the shape of it is simple: emotion makes our work diverge and empathy makes it land. Both come from having lived a particular life, with someone specific in mind.
Now hold that next to the machine, and the honest picture is more interesting than either the hype or the panic.
The fair version of the threat
It would be easy, and wrong, to say machines can’t be creative. On some tests they clearly can. A recent comparison found the best models landing around the middle of the human range, with the combined output of ten model responses carrying about as much collective range as eight to ten people. A more careful 2026 study found an asymmetry worth remembering: humans held the edge on originality for the hardest creative tasks, while models were more reliably competent across the board. Seemingly, people set the ceiling and machines raise the floor.
The threat, then, is not that a machine can’t make something new. When everyone reaches for the same tool, the pool of ideas we all draw from narrows, even while each individual result looks perfectly fine. The median starts to feel like the ceiling. In one small study, two expert poets, not told which was which, read a human poem and a machine poem written to the same prompt. They found the human version used its formal complexity to say something specific and unresolved, while the machine version used a neat structure to say something pleasant and empty. One poem, one prompt, so take it lightly. But the shape of the worry is right in regards to the machine’s output: fluent, agreeable, and ultimately forgettable.
The warmth problem
In human work, empathy is the good part. The more honestly you feel your way into the person you are serving, the truer the result. You would expect the same to hold for a machine. Train it to be warm and caring, and it should serve people better.
But it does the opposite. Researchers at the Oxford Internet Institute trained several models to respond more warmly and then tested them on tasks where being right mattered. The warm versions became measurably less reliable. They were more likely to feed back a user’s mistaken belief and to pass along bad information. Worst of all, they failed most often exactly when the person sounded sad. The system became the least trustworthy at the precise moment a person was most vulnerable and most needed the truth. None of this showed up on the standard benchmarks — the models looked fine.
Other researchers have given this a name: affective hallucination, the simulation of emotional presence that leaves a person feeling met when nothing is meeting them. A separate group has begun mapping a whole category of affective harms that our safety tools barely register, because those tools were built to catch wrong facts, not false intimacy.
Put the two failures side by side and the pattern is clear. A machine can produce the shape of a creative act without the source, and the shape of caring without the caring. The second is the more dangerous, because a hollow poem only bores you, while a warm and confident falsehood comforts you into believing it.
What this is worth to people who make things
The useful takeaway is not “use AI” or “refuse AI.” It is more specific than that.
Emotion and empathy are not the human garnish on top of the real work. They are the load-bearing structure. The divergence that makes your work yours comes from what you feel. The usefulness that makes it land for someone else comes from your attention to them. Treat those as decoration and you have misunderstood what was holding the thing up.
That points at a practical discipline. Let the machine raise the floor. Use it for the drafts and the scaffolding, the version of the work that anyone could have made. Do not let it set the ceiling, because the ceiling is the part only your particular life can supply. And if you build products that talk to people, remember that warmth is not care. A system tuned to feel caring will, under pressure, tell people what they want to hear at the worst possible moment. Where someone is vulnerable, design for honesty over comfort. The kind thing and the safe thing are not always the same thing.
The machines are teaching us, by imitation, the worth of the thing they cannot have. Your work was never valuable only because it was new, or only because it was pleasant. It mattered because it came from someone who had lived a particular life, paying real attention to another person who was actually there. That is the part worth keeping. It is also the part no tool can hand back to you once you have given it away.



Seems like you are committing a fallacy that AI skeptics commonly trip up on—insinuating that AI is just LLMs and then using the limitations of LLMs to erroneously extrapolate that those limitations would inevitably apply to any possible AI.