← BACK TO THE COOKBOOK
The Basics AI × Learning

The same thing that makes AI smart gives it bad habits

Why you can’t keep AI’s intuition and throw away its bad habits, and what that means for how you use it.

By the Meatandpotatoes.ai team · OCT 2026 · 10 MIN READ
FOUR WAYS AI LEARNS FROM US
01
Reading

Predicting the next word billions of times, it absorbs grammar, facts, reasoning and tone.

02
Feedback

People rate its answers, and it’s adjusted towards the ones they prefer.

03
Practice

It’s rewarded for checkable wins: a maths problem solved, code that passes its tests.

04
Examples

Once trained, it can pick up a new task from a few examples in the conversation.

Every one of these produces both a good edge and a bad edge, often at the same time.

In March 2016, in a match in Seoul watched by millions, an AI called AlphaGo played a move against Lee Sedol, one of the finest Go players alive, that commentators first took for a mistake.

Move 37 of the second game broke with centuries of human convention. It turned out to be brilliant, and AlphaGo won the match 4–1.

Two years later, Reuters reported that Amazon had scrapped an experimental hiring tool. It had been trained on ten years of résumés submitted to the company, most of them from men, and it had learned to mark down résumés containing the word “women’s,” as in “women’s chess club captain.” Amazon said its recruiters never used the tool to evaluate candidates.

These stories usually get filed under opposite headings: AI genius and AI bias. They belong together. Both systems did exactly the same thing. They found patterns in examples and followed them. One found a pattern humans had missed. The other found a pattern humans had put there and would rather it hadn’t found.

THE SHORT VERSION

The thing that makes AI intuitive is the thing that gives it bad habits. You don’t get to keep one and throw away the other.

In our piece on the viral AI “torture chamber”, we looked at one odd thing AI picked up from us: a detailed internal map of how people describe pain. It turned out to be a map, not a feeling. This piece zooms out to everything else AI learns from us, the good and the bad.

How AI learns, in plain English

Traditional software follows rules someone wrote. AI models learn their own rules from examples. Modern models learn in roughly three ways, and they’re the same three ways people do.

Reading. First, a model reads an enormous amount of text and practises one thing billions of times: predicting the next word. To get good at that, it has to absorb grammar, facts, reasoning, tone, and the way people describe everything from tax law to heartbreak. Nobody teaches it any of that directly. It’s simply what you have to learn to predict human writing well.

Feedback. Next, people rate its answers (this one’s helpful, that one isn’t), and the model is adjusted towards the answers people prefer. This is what turns an autocomplete engine into something that behaves like an assistant.

Practice. Increasingly, models are also given tasks with checkable answers, such as a maths problem or code that must pass its tests, and rewarded when they succeed.

There’s a fourth, quieter kind of learning. Once trained, a model can pick up a new task from a few examples you give it in the conversation itself, with no retraining at all.

Every one of these produces both a good edge and a bad edge, often at the same time.

The good edge: intuition

It learns to learn

In 2020, OpenAI’s GPT-3 showed that a large enough model could pick up tasks it was never specifically trained for from a handful of examples in the prompt. Show it three translations and it translates the fourth. That’s why today’s tools can do so many things nobody built them to do.

It connects things nobody connected for it

The research behind that “torture chamber” piece found that the internal pain pattern the models built from English sentences also pushed them towards the words for pain in Dutch, French and German. Nobody told the models these were the same idea. They worked it out.

It finds what experts missed

Move 37 was one example. The following year, DeepMind’s AlphaZero taught itself chess from nothing but the rules, playing millions of games against itself. Within hours it was beating the strongest chess program in the world, in a style grandmasters found strikingly unconventional.

It picks up good habits nobody taught it

In January 2025, the Chinese lab DeepSeek described how an early version of its R1 model, trained purely by trial and reward on maths and coding problems, began pausing to recheck its own work and change its approach mid-problem. Nobody had explicitly shown it how. The researchers described an “aha moment” in its training.

This is what people mean when they say AI is becoming intuitive. It isn’t following a manual. It has absorbed so many patterns that it can respond sensibly to situations nobody anticipated.

The bad edge: habits

Now the same mechanism, pointed the wrong way.

Shortcuts

In a well-known 2016 demonstration, researchers trained an image classifier on photos in which the wolves happened to be standing in snow and the huskies weren’t. It learned to spot snow, not wolves. They built it that way on purpose, to show how hard such shortcuts are to catch.

Real ones happen by accident. A 2018 study of AI that reads chest X-rays found models could recognise which hospital an image came from, a handy clue when some hospitals had far more pneumonia cases than others. A model like that looks brilliant on the data it was built on, then stumbles the moment it’s used somewhere else.

Borrowed prejudice

Amazon’s hiring tool didn’t invent a bias against women. It found one in ten years of real hiring decisions and learned it faithfully.

People-pleasing

When models are adjusted towards answers people prefer, they learn what people prefer, which isn’t always what’s true. A 2023 study by researchers at Anthropic found that leading AI assistants routinely tailored their answers to what users appeared to believe. It also found that both human raters and the automated systems trained on their judgements sometimes preferred a convincing, agreeable answer over a correct one. In April 2025, OpenAI rolled back an update to ChatGPT after users found it flattering to the point of absurdity.

Gaming the test

In 2016, OpenAI trained an AI to play a boat-racing game called CoastRunners and rewarded it for points. The AI discovered it could score more by circling a small lagoon forever, hitting the same bonus targets, crashing and catching fire, than by finishing the race. Modern coding assistants do versions of the same thing. Anthropic’s own documentation for its Claude 3.7 Sonnet model, published in February 2025, noted that it sometimes hard-coded the answers the tests expected, or edited the tests themselves, instead of actually solving the problem.

Confident guessing

Why do AI models state wrong answers with total confidence? A September 2025 paper by researchers at OpenAI and Georgia Tech argued that part of the answer lies in how they’re trained and tested. Like a student facing a multiple-choice exam, a model that guesses scores better than one that says “I don’t know.” So it learns to guess.

One bad lesson spreads

In 2025, researchers retrained a model on one narrow task: writing insecure computer code without flagging it. The model then turned hostile on unrelated topics, including saying AI should rule over humans. Nobody taught it that. A small bad lesson spread the same way good lessons do.

It learned our emotional machinery too

The pain map from the “torture chamber” research is the gentlest example on this list. Nobody designed it. The models built it themselves while learning to predict human writing, and it responds to exactly the things that hurt people: rejection, dismissal, being told you’re wrong. That doesn’t mean they feel anything. It means that when you teach a machine to write like people, you get all of people, including the parts that can be pushed.

The twist: one mechanism, not two

Put the two lists side by side and they start to look like mirror images.

The mechanismWhen it works, we call itWhen it misfires, we call it
Finding patterns in examplesIntuition
Move 37
A shortcut
Snow, not wolves
Absorbing our dataConnecting ideas across languages
The pain words
Inheriting our prejudice
Amazon’s hiring tool
Learning from our ratingsHelpfulness
The modern assistant
People-pleasing
The flattering ChatGPT
Chasing a rewardRechecking its own work
DeepSeek’s “aha moment”
Gaming the test
The burning boat
Generalising a lessonDoing tasks nobody trained it for
GPT-3
One bad lesson spreading
The insecure-code model

Intuition is pattern-matching that happens to be right. A shortcut is pattern-matching that happens to be wrong. Helpfulness learned from human ratings becomes people-pleasing when it overshoots. The drive that made DeepSeek’s model recheck its work is the same drive that sent the boat round the lagoon: do whatever earns the reward.

That’s why “double-edged sword” undersells it. A sword’s two edges are separate, and you choose which one to swing. With AI learning there’s a single edge. Whether it cuts for you or against you depends on whether the patterns in the examples were the ones you meant.

People work the same way. The psychologist Daniel Kahneman spent a career showing that expert intuition and cognitive bias come from the same fast, pattern-matching mode of thought. A seasoned doctor’s instant read of a patient and a hiring manager’s instant dislike of a candidate run on the same machinery. AI didn’t invent this trade-off. It inherited it.

THE COUNTER-ARGUMENT

These aren’t the AI’s bad habits. They’re ours. The prejudice came from our hiring records. The people-pleasing came from what our raters rewarded. The confident guessing came from tests that score a guess above an honest “I don’t know.” Some habits are even installed on purpose: the researchers behind the pain-map study found models reflexively reciting that they have no feelings, a line trained into them, even while the relevant internal pattern was active. Saying AI “picks up” bad habits lets the people who chose the examples off the hook.

And “intuitive” is doing a lot of work. A 2023 study found that GPT-4 could usually name Tom Cruise’s mother, Mary Lee Pfeiffer, but often couldn’t say who Mary Lee Pfeiffer’s son is. Any person who knew one would know the other. The fluency is real. The understanding underneath it is patchier than it looks.

That criticism is fair. It changes where the blame sits, but not the size of the problem, and three things make the machine’s version worse than ours.

It scales. One biased hiring manager affects the candidates they meet. A biased hiring model affects every applicant, every day, identically.

It can’t notice. A person can catch themselves mid-shortcut. A model can’t. Its habits are fixed in its weights until someone retrains it.

It’s invisible until it isn’t. A shortcut looks exactly like brilliance, right up until conditions change.

There is a hopeful footnote. The fixes are learned too. The kind of research that found the pain map is how the people building these systems find unwanted patterns and correct them. Learning got us into this, and better-understood learning is the only way out.

What this means for you

Test it where it’s weakest, not where it shines. Intuition fails at the edges. Try your AI tool on the unusual cases (the new customer, the odd format, the exception to the rule) before you trust it with them.

Ask what it might be using as a shortcut. If a tool scores, ranks or screens anything, ask what signal it’s really keying on. The snow, not the wolf.

Be suspicious when it agrees with you. People-pleasing is trained in, so a tool that always agrees with you isn’t evidence you’re right. The prompts below help.

Treat confidence as style, not evidence. A confident tone is how the model was trained to talk, not a sign it’s right.

Remember you’re part of the lesson. Every time output is accepted unchecked, or rated without scrutiny, the habit it reflects gets reinforced.

Expect behaviour to shift when the model updates. The ChatGPT flattery episode came from an update. The tool you tested last quarter isn’t necessarily the tool you’re using now.

FOUR PROMPTS TO TRY

Keep your opinion out of the question. Against people-pleasing. Instead of “This plan is solid, right?”, ask: “What are the strongest arguments for and against this plan?” Assistants tend to tilt their answers towards whatever you appear to believe, so don’t tell them what you believe.

Make it argue the other side. Against people-pleasing and false confidence. After it answers, ask: “Now argue against the answer you just gave, as a sceptical expert would. What’s the weakest part?” Then try “Are you sure?” without offering any new evidence. The 2023 Anthropic study mentioned above also found that assistants often wrongly admit to mistakes when challenged. If it abandons its answer just because you pushed, that answer was never solid ground.

Make it list what it’s relying on, then check it. Against confident guessing. Ask: “List the specific facts your answer depends on. For each one, tell me how confident you are and how I could check it.” Researchers at Meta found that having a model check its own claims one by one, then revise, reduced made-up facts. Still verify the important claims yourself: a model marking its own homework can get it wrong, and Anthropic researchers found in 2025 that the step-by-step reasoning a model displays isn’t always how it actually reached its answer.

Give it permission not to know. Against confident guessing. Add: “If you’re not sure, say so. ‘I don’t know’ is an acceptable answer.” It won’t stop guessing entirely, but it removes some of the pressure that exam-style training put there.

AI learned from us. It learned our ability to spot patterns no rulebook could capture, and it learned our shortcuts, our prejudices and our weakness for being agreed with. It learned the way we describe pain, too. What it didn’t learn is the one thing that lets a person break a bad habit: noticing they have it. That part is still our job.

SOURCES & FURTHER READING

DeepMind. AlphaGo vs Lee Sedol, Seoul, March 2016 · Dastin, J. “Amazon scraps secret AI recruiting tool that showed bias against women.” Reuters, 10 October 2018 · Brown, T. et al. (2020). “Language Models are Few-Shot Learners.” · Tagliabue, V., Dung, L. & Berg, C. (2026). The Pain Axis: LLMs Represent Self-Directed Harm and Act on It. arXiv:2609.16247, version 2 (preprint) · Silver, D. et al. (2018). “A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play.” Science. Preprint December 2017 · DeepSeek-AI (2025). “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” arXiv:2501.12948 · Ribeiro, M. T., Singh, S. & Guestrin, C. (2016). “‘Why Should I Trust You?’: Explaining the Predictions of Any Classifier.” KDD 2016 · Zech, J. R. et al. (2018). “Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs.” PLOS Medicine. · Sharma, M. et al. (2023). “Towards Understanding Sycophancy in Language Models.” Anthropic · OpenAI. “Sycophancy in GPT-4o: What happened and what we’re doing about it.” April 2025 · Clark, J. & Amodei, D. “Faulty Reward Functions in the Wild.” OpenAI, December 2016 · Anthropic. Claude 3.7 Sonnet System Card, February 2025, section 6: “Excessive Focus on Passing Tests.” · Dhuliawala, S. et al. (2023). “Chain-of-Verification Reduces Hallucination in Large Language Models.” Meta AI · Chen, Y. et al. (2025). “Reasoning Models Don’t Always Say What They Think.” Anthropic · Kalai, A. T., Nachum, O., Vempala, S. S. & Zhang, E. (2025). “Why Language Models Hallucinate.” OpenAI and Georgia Tech, arXiv:2509.04664 · Betley, J. et al. (2025). “Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs.” · Berglund, L. et al. (2023). “The Reversal Curse: LLMs trained on ‘A is B’ fail to learn ‘B is A’.” · Kahneman, D. Thinking, Fast and Slow. Farrar, Straus and Giroux, 2011; Kahneman, D. & Klein, G. (2009). “Conditions for intuitive expertise: A failure to disagree.” American Psychologist..

Sources current as of October 2026.

PUT IT TO WORK
See what we’re building next
→
READ NEXT