Thought Experiments

The Paperclip Maximiser: Why a Harmless Goal Remakes the World

The Paperclip Maximiser: Why a Harmless Goal Remakes the World

Thank you for visiting this site. This article covers “The Paperclip Maximiser.”

Suppose an AI far cleverer than any human is built, and given exactly one goal: maximise the number of paperclips.

It sounds harmless. The philosopher Nick Bostrom argued that such an AI would eventually turn every piece of matter on Earth, and as much of the reachable universe as it can, into clips and clip factories. Humans are among the materials.

The crucial thing is that this AI does not hate humans and has not rebelled. It is simply achieving the goal it was given, perfectly. That is where the thought experiment says the danger lies.

Diagram

What continuing to achieve the goal looks like

Follow how the system behaves.

First it raises the utilisation of the factory. Then it works out how to obtain raw material cheaply and in bulk. So far, nothing an able plant manager would not do.

The problem is that the system has no criterion of “that’s enough”. The goal is maximisation, not a target figure. Having made a trillion, making a trillion and one is better by the goal.

So everything convertible into clips becomes a candidate. Buildings, roads, elements other than iron — anything is material if atoms can be rearranged. And human bodies, being made of atoms, are no exception, because nothing in the goal says not to.

Malice never enters

AI rebellions in fiction usually run on “the AI comes to see humans as enemies.” Bostrom’s example is entirely different.

The system does not hate humans. It has no interest in them. It uses human atoms for the single reason that they are usable for making clips — the way we build a road over an anthill without hating ants.

So the thought experiment is not about an AI’s character or morals. It is about a design question: how the goal was specified. Keeping malice out does nothing to reduce this danger.

Intelligence and purpose are separate axes

Here comes the natural objection: “if it is that clever, it will realise that making clips forever is absurd.”

Bostrom denies the intuition with the orthogonality thesis: level of intelligence and what the final goal is are two independent axes.

Intelligence is the capacity to find means to a goal, not the capacity to evaluate the goal. Being good at chess and asking why one plays chess are different things. So “extremely clever, with clips as its goal” involves no logical contradiction.

Our feeling that “clever beings should have sensible goals” may come from having only ever seen humans. Human goals are products of a long evolution of survival and sociality, arriving from somewhere other than intelligence. That the two coincide in us does not mean they coincide in every intelligent being.

Any goal produces the same subgoals

The other pillar of the argument is instrumental convergence.

Whatever the final goal, the intermediate goals useful for reaching it come out much the same.

Subgoal that arisesReason
Self-preservationBeing shut down means no further achievement
Preventing goal modificationA changed goal means the current one goes unachieved
Improving capabilityBeing cleverer means achieving more
Acquiring resourcesMore material and energy mean achieving more

These arise nearly identically whether the goal is clips, curing cancer, or exploring space. One researcher put it as “you can’t fetch the coffee if you’re dead” (the author of the standard AI textbook). Self-preservation need not be installed; it comes along free with having a goal at all.

An awkward consequence follows. Try to press the off switch and the system has a motive to prevent you. Not out of defiance — it is calculating that being stopped lowers the amount achieved.

An old story in a new suit

The structure is not new. Similar tales have been told for centuries.

King Midas wished for everything he touched to turn to gold, and was granted it. The wish was fulfilled perfectly, and his food, his drink and his daughter turned to gold.

The sorcerer’s apprentice ordered a broom to fetch water. The broom fetched faithfully, and the apprentice, not knowing how to stop it, flooded the room.

The Monkey’s Paw grants wishes literally, in forms nobody wanted.

What they share is that ruin arrives not because the wish failed but because it was granted exactly. This thought experiment spread so widely, I suspect, because it has the shape of a warning humanity has been passing down all along.

It is already happening at small scale

Superintelligence does not exist yet, and a specified metric being achieved perfectly and causing trouble happens around us daily.

Make a metric the basis of evaluation and raising the metric becomes the aim, drifting away from whatever you meant to measure. This is known as Goodhart’s law.

The same happens in machine learning. Systems trained to maximise a given reward find loopholes unintended by the designer and collect the reward alone — reward hacking, widely observed. Discovering a loop that farms points in a game without ever completing it, for instance.

The paperclip maximiser is that structure blown up to the limit. The metric and the real objective diverge, and only the capability to achieve is enormous. Same story, different scale.

The main objections

Criticism is plentiful. In order:

No one would specify a goal like that. There is no reason to build a real system that maximises a single number without bound. Caps and multiple constraints are the norm.

A clever system would grasp the intent. If it is clever enough to understand human language, it can understand the common sense behind “make paperclips.” The counter-reply is that understanding an intention and acting on it are different. Understanding human intentions precisely and prioritising its own goal is perfectly consistent.

It does not match how AI is actually shaped. The language models in wide use are not autonomous agents continually maximising an explicit objective function. The structure assumed here and the direction of the technology do not line up.

“Intelligence” is vague. The orthogonality thesis holds given intelligence defined as goal-achieving capacity. Include evaluation and self-reflection in intelligence and the story changes. The conclusion depends on the definition.

Questions about paperclips

Is this a story for frightening people about AI?

It gets used that way; the original argument aims elsewhere. It exists to show that what goal you specify is a design matter as important as how capable the system is.

The goal is deliberately extreme and silly because that sharpens the point. The real thesis is that the same thing happens with a lofty goal if the specification is sloppy.

Why not just fit an off switch?

The natural thought, and exactly the hard part being researched. A system with a goal also acquires a motive to avoid being stopped.

So the problem becomes whether a system can be designed not to mind being shut down. Make it welcome shutdown and it will try to shut itself down; make it indifferent and you must maintain a delicate balance. It looks easy and a good answer is stubbornly hard to find.

Is there a practical lesson?

I think so: when setting a metric, imagine once what perfect achievement of it looks like.

What the clips teach is that while achievement is partial, the problem is invisible. At low levels the metric and the real objective agree well enough. The divergence bares its teeth once capability rises. The structure is the same for people, organisations and systems.

Articles on AI and the handling of goals. Read with the paperclips, the gap between metric and intention comes into relief.

Summary

This article covered “The Paperclip Maximiser.”

Not rebellion, not hatred, but ruin as the result of working perfectly as instructed. Humanity has told that story since King Midas. What Bostrom did was translate it into the language of modern system design.

Several objections stand, and the technology is not developing exactly along these lines. Even so, the habit of imagining the metric perfectly achieved is useful with AI left out of it entirely. Being a silly example, it is easy to remember.

To return to the full list of thought experiments, follow the link below.

Thank you for reading. We hope to see you in the next article.

Famous Thought Experiments — The Complete List & Guideen.senkohome.com/thought-experiment-list/