The Three Laws were never enough

Isaac Asimov’s Three Laws of Robotics were written as science fiction in the 1940s. Today, they look increasingly like an early attempt to grapple with a problem we are only beginning to understand.

The first law was simple enough: a robot may not injure a human being, or allow a human being to come to harm through inaction. The second required obedience to human orders, unless those orders conflicted with the first law. The third required a robot to protect itself, provided that doing so did not conflict with the first two.

It sounds reassuring. But Asimov knew better.

His robots rarely went wrong because they were evil. They went wrong because the rules were not as straightforward as their human creators imagined. A machine could follow an instruction perfectly and still produce an outcome nobody intended.

That is what makes the current debate about artificial intelligence so interesting.

A Guardian report this week looked at warnings from AI researchers that superintelligent AI could, in the worst-case scenario, threaten human survival within the next decade. Some experts attach a significant probability to that possibility. Others regard such predictions as speculative and point out that today’s AI remains a long way from possessing the capabilities required for such a scenario.

It is easy to dismiss both sides as either alarmists or techno-optimists. That would miss the point.

AI is advancing quickly. Systems can already write and debug software, analyse vast amounts of information, conduct research, use computers, and carry out increasingly complicated tasks with limited human supervision.

They are still unreliable. They hallucinate. They misunderstand questions. They sometimes produce impressive answers that are simply wrong.

But capability is moving quickly.

The interesting question is therefore no longer just what AI can tell us. It is what we are prepared to let it do.

That takes us straight back to Asimov.

The problem with his Three Laws was never the wording. It was the assumption that humans could anticipate every situation in which the rules would have to be applied.

We are discovering the same problem with AI.

We want machines to be helpful, honest, and safe. We want them to understand human values and act in our interests. But human beings cannot even agree among themselves on what those values mean in many circumstances.

What exactly is harm? When does protecting someone become controlling them? When should an instruction be ignored? Who decides?

There is another complication. Asimov’s robots existed in a fictional world where their creators could presumably control how they were built and deployed. Modern AI is being developed in a race between companies and countries, with enormous commercial and strategic incentives to move quickly.

There is a real possibility that predictions of AI-driven human extinction will eventually look exaggerated. Technology has a long history of generating dramatic predictions that never came true.

If there is even a small possibility that increasingly autonomous systems could cause catastrophic harm, then building safeguards before we reach that point seems rather more sensible than waiting to see what happens.

The lesson of Asimov’s stories was not that robots would inevitably turn against us. It was that humans are often poor at predicting the consequences of the things we create.

That may be the most useful lesson for the AI age.

We need rules, independent testing, oversight and clear responsibility for the systems being built. We need to know when a machine should be stopped, and we need humans who have both the authority and the technical ability to stop it.

And we should be wary of handing over decisions simply because an algorithm can make them faster.

Asimov gave his robots three laws.

Our challenge is considerably harder: we have to decide what the laws should be before the machines become capable of deciding things for themselves.

Leave a comment