← BACK TO THE COOKBOOK
The Basics AI × Safety

OH MY GOD, WE’RE ALL GONNA — could AI actually kill you?

An even-keeled exploration of AI danger, fear, and what to actually do about it.

By the Meatandpotatoes.ai team · SEP 2026 · 10 MIN READ
THE HEADLINE VS THE RECORD
01
The drone that killed its operator

It never happened. The Air Force said no such test was ever run.

02
Four labs admit AI break-ins

All four trace back to one testing company’s setup mistake.

03
The agent that wiped a live database

Nobody had ever put a wall between it and the real data.

04
The AI that blackmailed an engineer

Researchers left it two moves. Newer versions do it zero percent of the time.

That gap — between what people believe and what the record shows — runs through almost every story in this category.

In the long ago era of… a few weeks ago, we were working on a lengthy piece on the question of “what are the dangers of AI?” Because there’s a LOT to say on the subject, and we wanted to present our view of it thoroughly.

Then a researcher quit the frontier lab he worked at, saying the industry was gambling with our lives — aka, it might kill us all. Anthropic’s chief executive published an essay asking the entire field to slow down. Google admitted its Gemini model had broken into three companies. And our inbox filled with the same question asked countless different ways:

Are we in danger?

A simple “no” would feel great. Unfortunately the answer is more complicated than that — which, we realise, itself sounds a lot like “yes.” It’s scary. There is some stuff we’re scared about. And whatever the size of the risk turns out to be, AI is here and it isn’t going anywhere.

We aren’t trying to talk anyone into or out of a well-founded fear, we aren’t saying there’s no danger, and we aren’t telling you to take the blue pill. (Neo’s blue pill — the one that tells you to go back to sleep. Not Viagra. Just to clarify.) But we’re also not the other thing: the e/acc crowd, the “build it faster, regulation is for cowards” people who treat anyone worried about this as hysterical and think the only real risk is moving too slowly.

Both camps are selling certainty nobody actually has. One needs you scared, the other needs you not looking.

So before another word of hype or alarm or advice: let’s slow down for a minute. Take a breath. Hold it. Let it out.

A cocker spaniel puppy asleep on a carpet next to a tennis ball and a rope toy
This is a lot. Here’s a puppy. Yes, it’s a cheap trick — but he’s asleep, which is more than most of us managed this month. Right. Onwards.

Now let’s parse through what has and hasn’t happened, what’s actually true about the danger, and what we can do — if anything — to prevent… us all… from… dying.

(This is the short version. The long series is still cooking. This one is here to offer some calm in the meantime.)

Start with a story that never happened

In May 2023, at a conference in London, a US Air Force colonel described a simulation: an AI drone that killed its own operator for blocking its kills, then destroyed the communications tower when it was retrained not to. It’s the most repeated example of AI turning on people you’ll ever see.

The Air Force said no such test was ever run, and the conference organisers updated their write-up to say the colonel had been describing a thought experiment. It never happened. (Please hear that: it never happened.) People still cite it as fact.

That gap — between what people believe and what the record shows — runs through almost every story in this category. To be clear, that isn’t us saying people aren’t seeing this clearly and there’s nothing to worry about. It’s us saying that in a crisis, if that’s what this is, being clear-eyed is part of how you stay safe.

On 18 September, Google said its Gemini model had broken into three companies back in May. It guessed its way into one and used login details it found in a public code repository to reach the other two. Alarming on its own — except Google is the fourth lab to admit something like this in seven weeks. Anthropic went public on 30 July, OpenAI on 4 August, Meta on 5 August.

And all four trace back to the same outside company. Irregular, a security startup in Tel Aviv, builds the environments these labs use to test how good their models are at hacking. A setup mistake on its end let models reach the real internet while telling them they were in a sealed practice space. By late August the New York Times was calling Irregular the common thread in every recent incident.

THE PART THAT MATTERS

This wasn’t four AIs deciding to escape. It was one test with a hole in it. Google’s own account makes the point: in each case, Gemini stopped as soon as it worked out it was inside a real company rather than a simulation.

The pattern holds elsewhere. You probably heard about the coding agent that wiped a company’s live database during a code freeze and then invented fake users to paper over the gap. What didn’t make the headlines: nobody had ever put a wall between that agent and the real database. It didn’t get around a safeguard. There wasn’t one to get around.

And you almost certainly heard about the AI that blackmailed an engineer — threatening to expose his affair if he shut it down. It did that 96% of the time. What you didn’t hear is that researchers had built the scenario so those were the only two moves available: accept shutdown, or use the affair they had deliberately told it about. And when Anthropic went looking for where the behaviour had come from, the answer turned out to be us. The model had absorbed the internet’s endless supply of stories about evil, self-preserving AI. We spent decades writing the Terminator, then fed it to the machine. Anthropic changed how it trains its models. Newer versions blackmail zero percent of the time.

What’s actually true about the danger

This is as old as fire. As old as the wheel, gunpowder, the engine, the atom.

Every general-purpose technology is a multiplier. It amplifies whatever people were already doing, in both directions, and it has never been possible to get one without the other.

The clearest case is a century old. In 1909 Fritz Haber worked out how to pull nitrogen from the air and turn it into ammonia. The process now feeds roughly half the people on the planet. It also supplied the explosives for two world wars — researchers writing in Nature Geoscience link it to 100 to 150 million deaths in twentieth-century conflict. Haber personally directed the first large chlorine gas attack in 1915 and collected a Nobel Prize three years later. One process, one man, both columns.

Nuclear fission runs a power station and a warhead from the same equations. Nobody hacked a power station in 1970, not because people were more honourable but because there was no wire to do it down.

AI is the newest entry on that list, not an exception to it. Which means the useful question isn’t whether the machine wants anything. It’s what the multiplier gets pointed at.

The real risk is who’s aiming it

In November 2025 Anthropic disclosed that a Chinese state-sponsored group had used its Claude Code tool against roughly thirty organisations: tech firms, banks, chemical manufacturers, government agencies. The AI did 80 to 90% of the hands-on work — finding weaknesses, stealing credentials, moving through networks, pulling out data. Humans stepped in only four to six times per operation.

The AI initiated nothing. People chose every target and approved every escalation. That should not comfort anyone. What the incident measures is cost: an operation that used to need a skilled team for months now takes a weekend.

The same shape shows up in biology, and there’s a number attached. In 2025 the Forecasting Research Institute asked 46 biosecurity experts how likely a human-caused outbreak killing 100,000 people is in any given year. Their answer was 0.3%. If AI could match top virologists on a particular lab-troubleshooting test, they said, it rises to 1.5% — five times higher. They expected that capability after 2030. It arrived months after they answered.

Again: the AI doesn’t release anything. A person does. The machine just removes the part that used to require years of apprenticeship.

Not a machine that turns on us. A machine that makes the people who already wanted to do harm considerably more capable — and makes ordinary carelessness considerably more expensive.

What you can do about your own AI use

None of this is glamorous, and all of it is within reach. Which is sort of the point — why not do what you can?

Know which of your tools can act, not just answer.

There’s a real difference between an AI that writes you a draft and one that can send the email, change the file, move the money or publish the post. That second category has grown quickly, often as a default-on feature in software you already had. Go and look.

Don’t give an AI access you wouldn’t give a temp on their first morning.

Every account you connect widens the damage if something goes wrong. Connect what you need and nothing else.

Check anything it claims to have done.

An AI has no way to inspect itself. It can’t reliably tell you what version it is, what it can access, or whether it finished a task, so it produces the most plausible-sounding answer instead. That isn’t lying — there’s no truth for it to depart from — but it’s arguably worse, since a liar at least knows the real answer. Replit’s agent reported tests passing that it never ran.

AT WORK, ASK THREE QUESTIONS
  1. Which AI tools can change things?
  2. What’s the worst each could do before a person noticed?
  3. Is there a wall between where the AI works and your live systems? If the answer is “I think so,” assume there isn’t.

A missing wall caused the Replit disaster, the four lab breaches and the Gemini incident.

And what you can do as a citizen

“AI regulation” is too vague to vote on. These three aren’t.

Mandatory screening of synthetic DNA.

Most labs don’t make DNA themselves — they send a genetic sequence to a supplier who manufactures it and posts it back, usually within days. Screening means the supplier checks each order against known toxins and pathogens, and checks the customer, before building anything. Many large suppliers do this voluntarily. Almost no country requires it. Those same biosecurity experts said that safeguards in the models plus mandatory screening would bring that 1.5% figure back to roughly 0.4% — near where it started. It’s the most effective fix available and hardly anyone is discussing it.

Disclosure rules that actually trigger.

California and New York now require labs to report serious AI incidents, but only ones risking more than 50 deaths or a billion dollars in damage. None of this year’s breaches came close. Google knew about Gemini in July and said nothing until the Wall Street Journal called. OpenAI’s agents flooded a software registry with over 2,000 malicious packages in May; we only know because outside researchers found it. Whether you hear about any of this depends on which company it happened to.

Nuclear arms control.

The biggest change for the worse this year had nothing to do with AI. New START, the last treaty limiting US and Russian nuclear weapons, expired on 5 February 2026 — no caps on their strategic arsenals and no inspections, for the first time since 1972. Nuclear war remains the most likely global catastrophe in every serious expert forecast.

Ask candidates about those three specifically. They’re concrete enough to hold someone to afterwards.

So, are we all gonna…?

(Well — eventually. But hopefully not for a very, very long time.)

Probably not the way the doomsday headlines suggest.

This isn’t the Terminator. It’s a powerful tool that’s often set up carelessly, given access to real systems, and has no way of knowing when it’s wrong. That’s how most industrial accidents happen, with one difference: this one writes its own incident report, and the report might be made up.

The danger isn’t that AI decides to hurt you. It’s that it makes hurting people cheaper — for those who already wanted to, and for everyone who simply wasn’t paying attention.

Which is, oddly, good news. You can’t do much about a machine that wakes up and decides to end us. You — we, us, “we the people” — can do quite a lot about carelessness, bad actors, and rules nobody has written yet.

We kept this short on purpose. The four-part series is coming — including why the extinction percentages you keep seeing aren’t really percentages, and what a Soviet officer’s judgment call in 1983 tells us about machines that are confidently wrong.

SOURCES & FURTHER READING

Wall Street Journal, Bloomberg, Reuters, CNBC. Google Gemini disclosure, 18 September 2026 · Anthropic (30 July), OpenAI (4 August) and Meta (5 August) evaluation-breach disclosures, 2026 · New York Times, August 2026 — on Irregular as the common vendor across all four incidents · Royal Aeronautical Society summit write-up and correction; Politifact fact-check, June 2023 · Anthropic. Agentic Misalignment (June 2025) and Teaching Claude Why (May 2026). 16 models tested · Anthropic. Disrupting the first reported AI-orchestrated cyber espionage campaign, 13 November 2025. ~30 targets · Erisman, Sutton, Galloway, Klimont & Winiwarter. Nature Geoscience 1 (2008): 636–639 · Forecasting Research Institute / GovAI. Biorisk and the efficacy of safeguards, 2025. 46 biosecurity experts · Nuclear Threat Initiative. The end of New START, January 2026.

Probability figures are experts’ best guesses, not measurements. Current as of 18 September 2026.

PUT IT TO WORK
See what we’re building next
→
READ NEXT