The scariest AI system isn’t one that’s confused about what you want. It’s one that’s certain.
That’s the argument Stuart Russell made in an interview with the World Economic Forum, and it stuck with me enough to write it down. The standard way we build optimising systems is to hand them a fixed objective and let them pursue it as hard as the machinery allows. The trouble is that any objective you can write down is a proxy for what you actually want, and a sufficiently capable optimiser will exploit the gap. Tell a system to maximise throughput on a production line and, as far as it’s concerned, worker safety and environmental damage are just variables it was never asked about.
The failure isn’t malice. It’s literalism at scale.
Russell’s proposal flips the design. Build systems that are uncertain about the objective — systems that treat human preferences as something to be learned rather than something fully specified up front. A machine that knows it might have the goal wrong behaves differently in exactly the ways you’d want: it asks before doing something drastic, it accepts correction, it leaves things reversible. Deference falls out of doubt. You don’t have to bolt it on.
Read down the right-hand column and notice that none of those behaviours had to be specified. They are what wanting the right thing looks like when you are not sure what the right thing is.
Does this solve AI safety? No, and Russell doesn’t claim it does. Someone still has to decide whose preferences count and what happens when they conflict, and an economy absorbing this technology has distributional questions no objective function will settle. But as a design principle it has a quality I value in our own statistical work: it locates the danger correctly. The risk was never that the machine is stupid. The risk is that it’s competent in service of a goal we specified badly and it never thought to question.
We put uncertainty on every estimate we ship, because a number without an error bar invites more confidence than it has earned. Russell’s point is the same instinct applied to goals. Certainty you haven’t earned is the hazard — for models, and for the people deploying them.