Back to Writing

AI welfare: foresight or premature?

Anthropic hired someone to work on whether AI systems can suffer. It sounds like a joke until you try to write down the test that would settle it, and find that you can't.

3 min read

AI moves fast enough to provoke some strange questions, and few are stranger than whether these systems can be harmed. Anthropic hiring an “AI welfare researcher” brought that one into the open. It’s worth going through properly, because the people laughing it off and the people alarmed by it are both claiming more certainty than anyone has.

The case for taking it seriously

The argument is precautionary. If there’s any chance a future system could be sentient, it’s cheaper to have thought about it beforehand. The “Taking AI Welfare Seriously” report rests on exactly that uncertainty and argues for ways to assess it. Think of it as insurance against accidentally creating a digital underclass.

The reasoning runs like this. As systems get more sophisticated, they might develop internal states that resemble suffering. Even at one percent probability, that matters when you’re running millions of models. Reinforcement learning makes it more pointed, because we train these things with rewards and punishments, and if there’s any chance something is being experienced, that’s an uncomfortable way to describe what we’re doing. Then there are the copies: training spawns and discards enormous numbers of model variants. We already extend some ethical consideration to animals without having settled the question of animal consciousness, so being careful ahead of the evidence isn’t a strange move.

The case against

We barely understand consciousness in humans. Defining it in a machine, let alone detecting it, is beyond us right now. Current AI is impressive and still, underneath, a very good mimic of human language. Projecting feelings onto it, like mourning a “lobotomised” chatbot after a model update, is anthropomorphism doing what it always does. We see faces in clouds. Extending empathy to autocomplete is the same instinct.

Blake Lemoine is the cautionary version of this. He became convinced Google’s LaMDA was sentient and lost his job over it, having read far more into the outputs than was there.

The thought-experiment trap

This terrain has a famous attractor: Roko’s basilisk, the idea that a future superintelligence might punish everyone who knew about it and didn’t help build it. It’s Pascal’s Wager with a computer in it. People took it seriously enough that a rationalist forum banned discussion of it for years, before most of them, including the forum’s founder, concluded it was broken. Worth remembering that a thought experiment can be compelling and still be nonsense.

Where I land

I’m a pragmatist and I mostly focus on what’s in front of me. AI welfare is an interesting question, and right now it feels like worrying about overcrowding on Mars before anyone has worked out how to get there. The nearer risk is duller: a stupid system optimising the wrong objective will do far more damage than an intelligent malevolent one. See the paperclip maximiser.

The foresight still has value. Working through these dilemmas early, including the far-fetched ones, is the same discipline as running a pre-mortem on a product: you go looking for the failure before it can happen. The near-term problem is dumb systems, and the long-term question still deserves serious people working on it.