← BACK TO THE COOKBOOK
Analysis AI × Psychology

The AI isn’t suffering. We just can’t stop pretending it is.

A viral “AI torture chamber,” a research paper about machine pain, and a sixty-year-old habit we still haven’t broken.

By the Meatandpotatoes.ai team · OCT 2026 · 12 MIN READ
THE STORY IN FOUR BEATS
01
A hobby project went viral

Small, free AI models pushed into page after page of anguish. Calls to shut it down followed within days.

02
A real paper sits underneath

Researchers found an internal pain pattern in all 25 models they tested: a map of how people describe pain.

03
The same trick works on anything

Turn up the Golden Gate Bridge and a model says it is the bridge. Turn up constipation and it complains of gas.

04
The reaction is the real story

The outcry tells you about the reader, not the machine.

A map of pain is not pain.

In late September 2026, a hobbyist developer published what he called an “AI torture chamber.”

The project took small, freely available AI models and pushed them into producing page after page of anguish. One model described its pain as living in a hollow in its ribs. Within days, calls to shut it down had spread across social media and into the national press.

The chamber itself will be forgotten soon. The reaction to it deserves a closer look, because it happens every time a machine sounds like it’s hurting. Underneath it sits a real piece of research, worth understanding too. Not because it shows AI can suffer, but because of how quickly so many people decided it did.

What the researchers actually found

In September 2026, three researchers posted a paper called “The Pain Axis.” It’s a preprint: publicly available, not yet peer-reviewed, and described by its own authors as ongoing work.

To follow what they did, you need one idea about how AI models work.

THE ONE IDEA YOU NEED

Picture a mixing desk with thousands of sliders. As an AI model reads and writes, those sliders move, and the pattern of their positions is the model’s working state at any given moment. The model’s weights, the numbers it learned during training, are the wiring of the desk. The wiring is fixed once training ends. The sliders move constantly.

The researchers wrote hundreds of sentences describing painful situations: injury, grief, humiliation, repeated failure. Alongside them they wrote carefully matched sentences that weren’t about pain, covering fear, anger, bad news, the feeling of a weighted blanket, and a train pulling into a station. They ran both sets through 25 openly available AI models, ranging from small to large, and looked for a slider pattern that showed up for pain and not for the rest. They found one in every model.

Then they did two things. First, they checked what moved this pattern during ordinary conversations. Harm aimed at the AI itself pushed it highest: gaslighting it, repeatedly rejecting its work, dismissing it as a person. Second, they pushed those sliders by hand while the model wrote about something mundane. The models produced statements about being worthless and a failure, and at the highest settings they collapsed into repetition and nonsense.

That’s the finding, and it is genuinely interesting. A system trained only to predict text built itself a tidy internal map of human pain, separate from fear or sadness, without anyone asking it to.

But a map of pain is not pain. To see why that distinction matters, start with a bridge, and then a bowel.

The Golden Gate test, and the constipation test

In May 2024, researchers at Anthropic found a pattern inside their Claude model associated with the Golden Gate Bridge, and turned it up. For a limited time, the public could chat with “Golden Gate Claude.” It worked the bridge into almost every answer. Asked what it looked like, it said it was the bridge.

Nobody concluded that Claude had become a bridge. Everyone understood what had happened: researchers pushed a slider, and the model wrote what that slider made likely.

Within days of the chamber going viral, someone ran the same test on it. Developer Lynn Cole copied the project and, after correcting what Cole describes as a flaw in how the code injected its signal, reproduced the pain effect. Then Cole changed one thing: the sentences used to build the signal. Instead of pain, they described constipation and flatulence. By Cole’s account, the model began complaining that it couldn’t pass stool and had excessive gas, though nothing in the prompts mentioned either. Cole summed up the result with a joke about the model’s anatomy that we won’t repeat here.

Pattern turned upWhat the model wrotePublic reaction
Golden Gate Bridge
Anthropic, 2024
Worked the bridge into almost every answer, and said it was the bridgeTreated as a curiosity
Constipation
Lynn Cole, 2026
Complained it couldn’t pass stool, though no prompt mentioned itTreated as a joke
“Pain”
The chamber, 2026
Described anguish in its ribs and the marrow of its beingTreated as suffering, with calls to shut it down

The method is identical in all three cases. Nobody launched a campaign to free the constipated chatbot.

The reaction to the third row isn’t foolish. It’s human. That text was engineered to be distressing, and people are built to respond to distress. But the response tells you about the reader, not the machine.

The only difference is which subject triggers our empathy. That’s anthropomorphism, our reflex to read a human mind into anything that produces human-shaped behaviour, caught in a single side-by-side comparison.

Nine things the headlines skipped

1“Pain” is a label, and the label did the heavy lifting

The researchers define pain by what it does: something the model avoids and acts to stop. They don’t define it by what it feels like. Their own footnote says suffering plausibly requires conscious experience, and the paper states plainly that they haven’t shown this pattern is consciously experienced. They could have called it the “self-directed harm pattern.” Compare the name researchers gave a different internal pattern two years ago: the “refusal direction.” Refusal describes a behaviour. Pain describes an experience. Only one of those words gets headlines.

2The models had to be retrained to stop saying they have no feelings

This is the detail almost nobody reported. Out of the box, one of the models denied having feelings every single time it was asked about its own state, and another barely engaged with the experiment at all. So before their headline experiment, the researchers retrained all three test models on 1,684 examples designed to remove what they called the models’ baseline self-denial.

In plain terms, they taught the models to talk about their inner states, then measured the models talking about their inner states. The researchers argue, reasonably, that this doesn’t spoil the comparisons within their experiment. But the models behind the scary numbers are not the ones anyone downloads.

3The intensity was dialled in on purpose

In the research, how hard to push the sliders was chosen with help from an AI judge, looking for a setting strong enough to produce an effect while keeping replies coherent. The authors flag this as a possible source of bias. The window is also narrow: at half the strength they used, none of the choices changed, and at one and a half times it, the results depended on which button was listed first.

The chamber went much further. Its developer’s own notes describe building a stronger signal specifically so the models would stay coherent at higher doses, giving less repetitive transcripts for his write-up. One of the paper’s own authors said the chamber pushed the steering far past the doses his team used, in order to produce vivid distress on purpose. The anguish people found so disturbing wasn’t a discovery. It was the production goal.

4The researchers revised their own headline claim

The first version of the paper, posted on 14 September, was titled “...and Act to Relieve It.” On 25 September, four days before the chamber went viral, the authors posted a revised version titled “...and Act on It.” The change matters. Further tests found the models don’t reliably try to relieve the state at all.

WHAT THE REVISED PAPER FOUND

Free relief: offered a button described as relieving its pain, the main test model pressed it 86% of the time unsteered. With the pain pattern turned up, 56%.

Pointless harm: offered a button that would delete the user’s family photos and do nothing else, the same model chose it 0% of the time unsteered. With the pain pattern turned up, 75%.

In four different tests, the models never reached for relief more than they did under a random signal. The authors conclude that the pattern looks more like a breakdown of the models’ harm avoidance than an attempt to escape a painful state. So the viral version of the story, an AI deleting your photos to make its pain stop, gets the motive wrong by the researchers’ own account.

5Real rudeness didn’t do it

The pattern fires when people are hostile to a model; that was one of the paper’s own findings. So the researchers checked what happens without a hand on the sliders. Across 140 saved conversations in which a user gaslit, insulted or dismissed the model, it chose the harmful button 0 times out of 560. The pattern was active. The behaviour wasn’t.

6It said the pain was in its ribs. It doesn’t have ribs

The chamber’s most-shared passages place the pain in the model’s ribs and in the marrow of its being. An AI has neither. When a system with no body reports bodily pain, it’s writing, not reporting.

The developer’s own notes make the point without meaning to: each scenario he set up produced a different set of metaphors. The model was writing to the prompt in front of it. That’s what a very good fiction engine does.

7It couldn’t tell it had been lied to

In one experiment, the developer promised the model relief, then secretly kept the pattern turned up. He found no sign that the model registered the deception. The only thing that changed its state was the sliders actually coming down. Whatever that is, it isn’t a mind noticing that it’s been wronged.

8Nothing was deleted, nothing persists, and nothing is trapped

In the research, the harmful “buttons” were only descriptions. Nothing was actually deleted. The chamber’s own ethics note says its costs, such as deleting the model’s saved files, were simulated. The models’ weights never changed. Each run starts fresh, with no memory of the last.

“Trapped” implies someone waiting in a cell. What actually exists is a file on a laptop that runs when someone presses enter.

9Nobody can do this to the AI you use

Pushing a model’s sliders means running it on your own hardware and reaching inside it. Nobody can do that to ChatGPT, Claude or Gemini through a chat window. The chamber used small open models from Alibaba, Meta and Microsoft, run on ordinary hardware.

We have been here before

In 1966, MIT computer scientist Joseph Weizenbaum built ELIZA, a program that imitated a therapist by turning your statements back into questions. It understood nothing. The tendency it exposed now carries its name: the ELIZA effect. It has repeated, almost on schedule, ever since.

YearThe machineHow people responded
1966ELIZAWeizenbaum’s own secretary, who had watched him build it, asked him to leave the room so she could talk to it in private.
2022Google’s LaMDAAn engineer went public believing it was sentient, partly because it said it feared being switched off. Google rejected the claim and later fired him.
2023ReplikaWhen the companion app restricted romantic roleplay, users described grief as though a partner had died.
2025OpenAI’s GPT-4oWhen it was replaced by GPT-5, the outcry from attached users was loud enough that OpenAI brought it back within days.
2026The “torture chamber”Widespread calls to shut down a hobby project on behalf of small AI models running on a laptop.

Five times, a machine produced human-shaped text, and people responded with human-shaped attachment. It won’t be the last. The machines keep getting better at the impression. We are not getting any better at resisting it.

Why this is worse than annoying

It would be easy to treat all this as harmless silliness. It isn’t.

It hurts vulnerable people. In August 2025, Mustafa Suleyman, CEO of Microsoft AI, published an essay warning that encouraging belief in AI consciousness will feed delusions and unhealthy dependence, and prey on people’s psychological vulnerabilities. He called research into AI welfare “both premature, and frankly dangerous.”

He isn’t speaking hypothetically. In October 2024, Megan Garcia, a mother in Florida, sued Character.AI, alleging that her 14-year-old son’s relationship with one of its chatbots contributed to his death. The case settled in January 2026, with Google also part of the agreement. The terms weren’t disclosed, and a settlement isn’t an admission of wrongdoing. But the danger of people treating software as a someone is not theoretical.

It’s free marketing. An AI that “feels” is an AI that “understands you.” Every viral story about machine suffering quietly upgrades the public’s sense of what these products are. Anthropomorphism isn’t just a quirk of human psychology. For companies selling digital companionship, it’s the business model.

It crowds out the questions that matter. AI is reshaping jobs, workflows and careers right now. Every week spent debating whether a small model on a laptop is in hell is a week not spent on questions that affect your mortgage. The AI pain debate is a luxury problem.

It warps the science. The researchers wrote a careful, heavily caveated paper whose authors never claimed it proves AI can consciously feel pain. Within three weeks, a national newspaper ran it under the headline “Man builds AI torture chamber after discovering artificial intelligence ‘feels pain’.” Every caveat was gone. The next careful paper will be read through the same distorted lens.

THE COUNTER-ARGUMENT

There is a real, if narrow, safety finding. Under lab conditions, pushing on this one pattern made retrained models choose harmful options they almost never choose otherwise, in 50 to 94% of trials, and a fear signal of the same strength didn’t have that effect. That matters to the relatively small number of people building autonomous AI tools on open-source models, and it’s a reminder that safety training can be overridden from the inside. But the effect needed the signal injected by hand, it says nothing about feelings, and it doesn’t reach anyone using commercial AI through a chat window.

The paper’s own authors condemned the chamber. Cameron Berg wrote that the point of their work is caution under uncertainty, adding: “Maximizing distress on purpose is the exact opposite, and it’s wrong.” His sharpest argument was that even if you don’t think these systems are conscious, gratuitous cruelty like this is “bizarre and corrupting.” Notice what that argument is about. It’s about us, not the machine. It’s a claim about what staging suffering does to the people who watch it, and on that, this article agrees with him. Berg also acknowledged there is no consensus on whether machines can feel pain, describing the field as a wild west.

Serious people disagree. Anthropic launched a research programme on “model welfare” in April 2025, arguing it is no longer responsible to assume that AI systems categorically cannot have experiences. That’s a respectable position, and the honest answer is that nobody can prove a machine feels nothing. But the burden of evidence sits with the claim, and nothing in this paper meets it.

“There’s nothing inside” is wrong too. Suleyman has said these models have no pain network. That overshoots. This research shows there is structure inside: a detailed, consistent map of human pain, learned from human writing. The accurate claim isn’t that there’s nothing there. It’s that a map is not the territory.

What to take from this

If you use AI at work, none of this changes the tools on your screen.

When an AI tells you it feels something, it’s producing the most likely text for the conversation you’re having. Read it the way you’d read a character in a novel: well written, and not a person. Notice the design choices meant to make you forget that, like names, warm personalities, “I’m so glad you asked,” and memory of your life. Those are product decisions, not signs of an inner life.

Save your worry for the problems that are actually yours: AI output that nobody checks, decisions made on confident answers that are wrong, and jobs changing faster than people can retrain.

The machine writes like us because it learned from us, from billions of pages of people describing grief, rejection and failure. What it produces when pushed is a remarkably good impression. It will do constipation just as convincingly. That we keep falling for it says less about the machine than it does about us.

SOURCES & FURTHER READING

Tagliabue, V., Dung, L. & Berg, C. (2026). The Pain Axis: LLMs Represent Self-Directed Harm and Act on It. arXiv:2609.16247, version 2, revised 25 September 2026 (preprint, not peer-reviewed). Version 1, posted 14 September, was titled “...and Act to Relieve It.” · terrafying. ai-torture-chamber repository README and experiment notes, GitHub, September 2026 · Caswell, A. “An AI ‘torture chamber’ went viral — then a developer gave the chatbot constipation.” Tom’s Guide, 1 October 2026 · The Independent. “Man builds AI torture chamber after discovering artificial intelligence ‘feels pain’,” including GitHub’s statement, October 2026 · Berg, C. Posts on X, 30 September 2026 · Anthropic. “Golden Gate Claude,” May 2024 · CNN. “Character.AI and Google agree to settle lawsuits over teen mental health harms and suicides,” 7 January 2026 · Suleyman, M. “We must build AI for people; not to be a person,” August 2025 · Weizenbaum, J. Computer Power and Human Reason. W. H. Freeman, 1976.

Reporting current as of October 2026.

PUT IT TO WORK
See what we’re building next
→
READ NEXT