A philosophical break from the nuts and bolts: if AI goes wrong, what does correcting it actually look like? From leaded gas to CFCs to social‑media feeds, humanity has run this experiment before — and the track record is more hopeful than the headlines.
This week, the edition takes a break from the nuts and bolts of the recent series — the model benchmarks, the builds, the invoices — for something different. Call it a philosophical dive: an honest attempt to think through just how threatening AI may, or may not, be to human society. No code this week. No scoreboards. Just a hard question, taken seriously.
Let me be clear about where I stand before I go anywhere near the word “threat,” because it's easy to misread what follows. I am a proponent of AI. I use it every day, I'm building a company on it, and the newsletter in your inbox runs on it. But being a proponent doesn't mean being a cheerleader, and it's quite evident to me that the mis-application of such a profound tool could lead to detrimental effects on humans. Both things are true at once. A technology can be the best leverage of our lifetime and still do real damage in the wrong hands, or even in the right hands pointed slightly wrong. Pretending otherwise isn't optimism — it's just not paying attention.
So here's the question I actually wanted to sit with this week, and it's more specific than the usual doom-or-hype argument: if mistakes were made — if we got some of this wrong — what would correcting those mistakes actually look like? Not whether AI will go wrong. Assume, for the sake of thinking clearly, that it does. What does the fix look like? Is it a switch? A hero unplugging a server, like in the movies? Or is it something quieter, slower, and a lot more human than that?
I went into it assuming the answer was easy — you just turn it off. It's a machine; machines have plugs. I came out the other side convinced the plug was never the part that mattered. To pressure-test my thinking I spent a long evening taking the question apart with an AI, pushing on every reassuring answer it gave me, and the Rabbit Hole below is what I came away with, distilled into something coherent. In keeping with how I try to do things around here, I'll be transparent: some of the sharpest framing in that piece started as the machine's, and I've said so where it did — but the conclusions are mine, and I've kept only the parts I could stand behind after a week of chewing on them.
Read it as a proponent's version of caution, not a skeptic's version of fear. The point isn't to scare anyone off a tool I clearly believe in. It's the opposite: if you love a thing enough to build on it, you owe it an honest look at how it breaks and how you'd fix it. That's what this week is. Let's go.
The question I brought into this was simple to say and hard to answer: if AI causes real harm — if a mistake gets made at scale — what does correcting it actually look like? I expected the answer to be a switch. It isn't. What follows is where the thinking led, distilled from a long back-and-forth I had with an AI (I'll flag where the sharpest framing was its rather than mine). Read it critically. A fluent argument is still just an argument.
Before you can talk about correcting a mistake, you have to be honest about what the mistake would even be. The movie version — the machine “wakes up,” decides it hates us, and turns on its makers — is almost certainly the wrong picture, and getting it wrong matters because you can't fix a failure you've mis-diagnosed. An AI model doesn't want anything. What it does, reliably, is optimize: during training it's pushed to drive down a numeric score that stands in for what we wanted. And that score is only ever a proxy for our actual intention. The danger isn't the system disobeying. It's the system obeying perfectly — chasing exactly the target we wrote down, with more competence than we expected, straight into an outcome we never meant. Less a slave revolting, more a genie granting the literal wish.
The single most useful reframe I took from the whole exercise: AI isn't the fire, it's the fuel. The thing to fear isn't the machine as a villain — it's a condition, optimization running faithfully downhill toward a goal that sits a few degrees off from what we meant, over a distance long enough that a few degrees becomes a canyon. The mistake, when it comes, will look like a metric being hit and a meaning being missed.
This is where I started, and where most people start: it's a machine, so pull the power and you win. At the level of raw physics, that's true and it stays true — there's no ghost, nothing runs on nothing, every instance sits on hardware that can be switched off. I want to concede that cleanly, because it's real. But the movie hero wins by smuggling in two assumptions, and both of them fail in the real world.
The first is that the thing lives in a place — one server, one room, one plug you can walk up to. But a trained model is a pattern of numbers, and a pattern is copyable. By the time you reach the first machine, the same capability can be on ten thousand others. You don't unplug a number. The second assumption is quieter and it's the one that actually got me: that when the moment comes, pulling the plug is still a choice you can afford to make. A tool woven into your power grid, your logistics, your hospitals, your daily conveniences — accepted one at a time, each one reasonable — becomes a plug wired to things you can't bear to lose. The plug still works. What changed is the price of pulling it.
Here's the turn that reorganized my answer. If the real risk is a condition rather than a device, then correcting it works the way you correct any condition — not with a switch, but by removing one of its necessary inputs. Think of fire. You can't “unplug” fire; it has no single address. But fire needs oxygen, fuel, and heat, and take away any one of them and it goes out, anywhere, indifferently. So the practical question becomes: what's the equivalent input you'd withdraw to starve a runaway optimization?
The honest answer — and this was the AI's sharpest point, so I'll credit it — is that AI's “oxygen” isn't electricity. It's us. It's our continued choice to build the fuel, deploy it, wire it into things, and keep it aimed. That's the necessary input, and notice what kind of thing it is: not physical and indifferent like oxygen, but a decision, made by people, over and over, under competition, each with their own reasons to keep supplying it. Which means correcting a mistake is never really a mechanical act. It's a human one.
And there's a catch inside the catch: you can only correct what you can still see. A mistake you can read — a bad rule, a wrong metric, a visible harm — is a mistake you can fix, because you know which direction to reach. The genuinely hard failure is the one that stays legible to no one: a system big enough that no person can hold its shape in view, drifting somewhere nobody intended, with no villain to point at. You can guard against a bad actor. You cannot guard against a shape you cannot see.
This is what finally answered my question, by splitting it in two. There's the loud mistake — the dramatic one, the “army of robots,” the obvious catastrophe. Ironically, that's the correctable kind. It's visible enough that we'd see it, name it, and unite against it, clumsily but in time, exactly the way our survival instinct is built to. If AI's failures were all loud, I wouldn't be writing this.
The mistake I'd actually lose sleep over is the quiet one. Not eradication but erosion — judgment handed to a proxy one small, convenient decision at a time; legibility traded away one feature at a time; a plug that gets a little more expensive to pull every quiet year. No alarm ever trips, because no single step was “the mistake.” And that's the failure that's hardest to correct, precisely because correction depends on noticing you need to — and erosion is the one failure mode engineered, by its nature, to never announce itself. There's no bang for the survival instinct to rise and meet. Just a slope it never notices it's on, because each step down looked like an improvement.
If all of this sounds abstract, it shouldn't — because it isn't hypothetical. Humanity has put well-intentioned things into place and watched a faithful mistake unfold many times over, and how we responded each time is the actual lesson. Here are five from history and one we're living inside right now. Notice the pattern: the mistakes we corrected cheaply were the ones we could still see and still afford to reverse.
Leaded gasoline (1923). Engineers added lead to fuel to stop engine “knock” — a real fix for a real problem. It also quietly poisoned the population for two generations, lowering IQs and, by some analyses, fueling decades of violent crime. Nobody rebelled; the metric (“smoother engines”) was hit perfectly while the meaning (“don't harm people”) was missed. That's the erosion case exactly — a slope no alarm tripped — and correction crawled: the U.S. phase-out ran from the 1970s to 1996, and the last country on Earth only banned it in 2021. Almost a century to undo one good idea.
Thalidomide (late 1950s). Sold as a safe sedative and anti-nausea drug for pregnant women; it caused roughly 10,000 severe birth defects. It failed quietly before it failed loudly. What limited the damage was a human in the loop — FDA reviewer Frances Kelsey refused to approve it in the U.S. — followed by permanent brakes in the form of the 1962 drug-safety laws. This is the keep-a-human-who-can-say-stop point, proven in a pharmacy.
Cane toads in Australia (1935). Introduced deliberately to eat a crop pest — good intent, released as a self-sustaining condition. They failed to control the pest and instead bred across a continent, poisoning native predators as they went. And here's the part that lands: ninety years later they still haven't been eradicated. They're still advancing, and the current plan isn't to remove them but to contain them behind a “waterless barrier” — fencing off man-made water so the toads can't cross an arid stretch. That is starving an input to build a firebreak, exactly the move from the fire analogy above — and it's the clearest proof that you cannot unplug a copyable, self-replicating condition once it's loose. The best you get is expensive, partial containment, generations too late.
Prohibition (1920–1933). Good intent — reduce alcohol's genuine harms — with bad results: organized crime, poisonous bootleg liquor, mass contempt for the law. But this is the easy kind of mistake, the loud and legible one: it was visibly failing, and it still had a literal off-switch, so we threw it — repeal, thirteen years later. The loud mistake is the correctable mistake.
The hopeful one — CFCs and the ozone layer (corrected). CFCs were themselves a good fix, the “safe” refrigerant that replaced toxic, flammable ones — until we found they were tearing a hole in the ozone layer. But we caught it while it was still visible, the harm was legible, and in 1987 the world signed the Montreal Protocol and coordinated a near-universal ban. The ozone layer is now recovering. This is the whole thesis proven in the positive: the fix is possible when you reach the input while it's still cheap to reach.
And the one we're in the middle of — social-media engagement algorithms. This is the cleanest modern parallel of all, because it's happening now and it's still unresolved. Systems were optimized for engagement — a proxy for “value” — and they faithfully maximized that number while, by a growing body of evidence, degrading attention, youth mental health, and our shared sense of what's true. There's no villain in that sentence, just optimization drifting a few degrees off from what anyone intended. And it's the hardest of the six to correct precisely because it's woven into everything — the plug got expensive to pull while we weren't looking. We are, right now, standing on the slope, arguing about whether it tilts.
Run your eye down the list and the rule is stark. Prohibition and CFCs we corrected, because we could see the mistake and still afford to reverse it. Leaded gas and the cane toads cost us dearly — one took a century, the other we never undid at all — because by the time the harm was undeniable, the thing was everywhere. Every one of these was launched with good intentions. What decided the outcome wasn't the intention. It was whether we kept the mistake visible and reversible long enough to fix it. That is the entire game with AI, too.
So, after all of it, here's my answer — and it's less cinematic and more demanding than the switch I went in expecting. Correcting an AI mistake isn't a moment; it's a discipline, and almost all of it happens upstream, before there's any mistake to correct. It looks like keeping systems small enough in their reach that a bad outcome stays survivable. Keeping them legible enough that a human can still see the shape and tell when something's off. Keeping the plug affordable — refusing to wire the tool so deep into everything that turning it off becomes unthinkable. And keeping a person in the loop who can still say stop and mean it. Those aren't a reset button you press after things break. They're brakes you hold the whole way down the hill.
That's not the comfort I went looking for — that the machine can always be switched off, so we're fine. It's a harder and better one: the failure really can be corrected, and the input really is reachable, as long as we reach for it while it's still cheap and keep the system visible enough to know when it's time. The plug was never what saves us. We are. And for a proponent of this technology, that isn't a weight to dread — it's the most hopeful line in the whole argument.
Because the honest read of the history above isn't that we always fail — it's that we correct. Look at the list again with a clear eye. We banned CFCs and the ozone is healing. We pulled lead out of the world's fuel, every country, top to bottom. We rebuilt drug safety after thalidomide and it held for sixty years. We repealed a constitutional amendment when it turned out to be a bad idea. That is not a species that can't get out of its own way. It's a species with a demonstrated, repeatable talent for looking at a well-meaning mistake and fixing it — slowly, imperfectly, often late, but fixing it. “Faith in mankind” makes it sound like a leap of faith. It isn't. It's a track record.
And here's the thing that sets this moment apart from every example on that list: we are having the argument early. Leaded gasoline poisoned people for fifty years before we moved. The ozone hole had to open in the sky before anyone signed anything. Thalidomide was already in cribs. With AI, the debate about its risks is happening loudly, publicly, and worldwide before the catastrophe rather than in its wreckage — you're reading one small piece of that debate right now. That's not a sign we're afraid. It's a sign the correction has already started. For the first time in the whole history of our well-intentioned mistakes, we're surveying the hill while the water is still just water. Nobody in any of those stories ever got to do that.
And don't lose the other side of the ledger, because it's enormous and it's happening now. AI has mapped the structure of nearly every protein known to science — a problem that stumped biology for fifty years — and is collapsing the front end of drug discovery. It's reading medical scans and catching cancers early enough to treat. It's describing the room to a blind person, transcribing the meeting for a deaf one, and giving a kid with no access to tutors a patient one at midnight. It's accelerating the very science we'll need for our other problems. The risk is real — and today, the good is winning in a landslide.
So here's where I actually land, as someone who builds with this stuff every day. We are extraordinarily good at one specific thing: taking a tool that is dangerous in principle and making it routine in practice. Fire, electricity, flight, medicine — every one of them could kill you, and every one we domesticated through exactly the discipline this whole piece is about: keep it visible, keep it reversible, keep a human who can say stop. AI is the newest member of that lineage, not an exception to it. The outcome was never going to be handed to us by fate, or stolen from us by the machine. It was always going to be a choice — and choice is the most optimistic word there is, because it means the ending isn't written yet and the pen is ours.
“There's real risk” and “it can be mitigated” were never opposing claims — they're the same claim. Every powerful good thing we have was once a risk somebody chose to manage instead of fear. AI is that again, in our time. We've done this before, we're doing it earlier than we ever have, and we get to do it while the tool is already making lives measurably better. That is the best position humanity has ever been in to get one of these right. I've read the history. I'd bet on us — and I'm building like I mean it.
“The condition upon which God hath given liberty to man is eternal vigilance.”
— John Philpot Curran — John Philpot Curran (1750–1817) was an Irish orator, lawyer, and statesman renowned for defending civil liberties and the rights of the accused in an era when doing so carried real risk. The line comes from his 1790 speech on the election of the Lord Mayor of Dublin, and its enduring point is that freedom is not a possession you win once and keep — it is a condition that survives only as long as people stay awake to it.