Thought Experiments

The Turing Test: If a Machine Passes as Human, Is It Thinking?

The Turing Test: If a Machine Passes as Human, Is It Thinking?

Thank you for visiting this site. This article covers the “Turing Test.”

“Can machines think?” Try to answer that head-on and you must first settle what “think” means, and there the conversation stops.

In 1950 the mathematician Alan Turing got around the impasse by swapping the question for a different one. He abandoned the argument about definitions and replaced it with a decidable question: “can it hold a conversation indistinguishable from a human’s?”

The manoeuvre was elegant, and it also dragged along the suspicion that passing as human might not be enough. Now that AI is ordinary, the argument is worth rereading.

Diagram

Swapping the question

Turing’s paper Computing Machinery and Intelligence opens abruptly with the line “I propose to consider the question, can machines think?”

And he immediately stops pursuing it. Starting from definitions of “machine” and “think” turns the exercise into a survey of ordinary usage, with the answer settled by opinion poll. Which would be pointless.

So Turing recasts the question, translating a vague question into a game with a clearly defined procedure. That substitution is, I think, the sharpest thing in the paper.

It began as a game about men and women

The substitution is built on a period parlour game called the “imitation game.” No machine appears in the original setup at all.

  • A man A, a woman B and an interrogator C are in separate rooms
  • C asks questions in writing only and tries to determine which is the man and which the woman
  • A tries to deceive the interrogator; B tries honestly to help

Turing then restates the question: “What will happen when a machine takes the part of A? Will the interrogator decide wrongly as often?”

In that form the answer is settled by actually running it. The interior question of “is it thinking?” has become the externally measurable question of “can it be told apart?”

The Turing test as usually discussed today is a tidied version: a judge converses in text with both a human and a machine and tries to identify the machine. Fail to identify it and the machine wins.

Turing also left a concrete prediction. In roughly fifty years, he thought, machines would deceive judges over five minutes of questioning (on an assumption about storage capacity that was enormous for the time), and the chance of correct identification would fall to around 70%. He offered that as a marker for when “machines think” would become an ordinary thing to say.

Nine objections, pre-empted

Most of the paper’s second half is given over to answering anticipated objections. Turing lined up nine of them himself and dealt with each. The main ones:

ObjectionThe claimTuring’s reply
TheologicalThinking is God’s gift to the soul; machines have noneIt arbitrarily restricts what God is able to do
MathematicalThere are questions a machine necessarily cannot answerHumans get questions wrong all the time too
From consciousnessIt means nothing unless written with feelingAccept that and you cannot verify other people’s consciousness either
Various disabilitiesMachines cannot err, cannot enjoy humourA generalisation from the machines that happen to exist
Lady LovelaceA machine only does what it is told; it originates nothingA learning machine can surprise its author

The reply to the “argument from consciousness” is particularly sharp. Insist on “we cannot grant intelligence without confirming it genuinely feels” and we cannot run that confirmation on other people either. Hold the standard consistently and the only mind you may grant is your own.

Also worth noting is the proposal he placed at the end. Rather than building an adult brain outright, it may be quicker to build something corresponding to a child’s brain and educate it. The idea of acquiring capability through learning is written down at that point.

Attacks on “passing is enough”

Being famous, the test has attracted sustained criticism. The main lines:

You can pass without understanding. The best-known criticism is the “Chinese Room.” A person merely rearranging symbols by rule understands not a word of Chinese, while from outside the responses look fluent. The same behaviour may have nothing inside it.

A vast lookup table can pass. There is a more extreme case still. Prepare an unimaginably large table listing every possible conversation in advance and that machine passes the test. All that is happening is lookup, and nothing there deserves the name intelligence. The case exposes how completely the test ignores “how the answer was produced.”

It becomes a contest in imitating humanness. The shortcut to winning is not to be cleverer but to imitate human weaknesses: be slow and wrong at arithmetic, insert typing errors, change the subject. Many early conversational programs were optimised in exactly that direction. Some used the pretext of a child speaking a second language to excuse the awkwardness.

It is excessively human-centred. An intelligence unlike a human’s is identified as a machine the more honestly it answers. The usual analogy is that aircraft learned to fly once they stopped imitating birds. As long as humanness is the pass condition, the test cannot be a yardstick for intelligence in general.

People are easily fooled. Even a very simple response program from the 1960s led users to confide real troubles to it. The disposition to find a mind in the other party is strongly present on our side. What the test measures may be human gullibility rather than machine intelligence.

Where it is used inverted

Interestingly, the test appears in everyday life with its roles reversed: the distorted characters and image grids that websites present.

Those are a Turing test with judge and candidate swapped, in which a machine sets the question and decides whether you are human. The name itself encodes it: “completely automated public Turing test to tell computers and humans apart.”

The same irony attaches here too. As image recognition improved, these challenges have been broken one after another. The task of continually finding something only humans can do is itself getting steadily harder.

What large language models changed

Recent language models respond in natural prose on any subject, so the argument that “surely the Turing test has now been passed” keeps recurring.

There is no simple answer, because where you draw the pass mark changes the conclusion entirely. Whether the judge is a layperson or a specialist, whether the conversation runs five minutes or several hours, whether the judge knows the tells of AI — the result moves a great deal.

And even granting a pass, there is the question of what follows. Turing did not prove that “passing means it thinks.” His claim was a prediction about usage: that once you get that far, nobody will be able to object to saying machines think.

I think the paper’s value lies less in the criterion than in forcing the question of “on what grounds do we grant intelligence to something whose insides we cannot see?” With humans, we grant minds on behaviour alone. If we demand a different proof from machines, we owe an explanation of why.

Questions people ask about the Turing test

Has a machine actually passed?

Several passes have been declared, all of them conditional. Every time, criticism followed about the brevity of the conversation, the make-up of the judges, or the machine’s framing being convenient.

Turing himself never rigorously fixed a pass mark, so no shared standard for “this counts as passing” exists. That vagueness is the main reason the argument never ends.

Is judging by conversation alone too narrow?

The point is old, and extended versions have been proposed: judging not only text exchange but the ability to see, touch and move objects.

Behind it is the claim that meaning is acquired through engagement with the world, so handling words alone, without embodied experience, cannot amount to genuine understanding. It connects with arguments that meaning is fixed by connections to the external world.

Is this thought experiment useful at work?

What is useful is the instinct that “output that looks correct” and “processing that is correct” are different things. That is precisely the test’s weakness: judging on results alone without asking anything about the process.

Evaluating AI output is the same situation. A plausible answer coming back and a well-grounded answer coming back are not the same. An evaluation that ignores the interior awards high marks to whatever imitates best. A criticism from more than seventy years ago is, unchanged, a practical caution today.

Thought experiments in the philosophy of mind about whether a machine can be granted a mind, each objecting to Turing’s substitution from a different angle.

Summary

This article covered the “Turing Test.”

Replace the unanswerable “can machines think?” with the measurable “can they be told apart?” With that single move Turing pushed the argument forward — and took on the cost of never asking about the interior, which criticism has been probing ever since.

What is striking is how many of those criticisms Turing had already anticipated and addressed himself. What he had no answer for was the later objection that “a vast lookup table would also pass.” Rereading it, the paper feels less like a classic than like a document still standing in the middle of the argument.

To return to the full list of thought experiments, follow the link below.

Thank you for reading. We hope to see you in the next article.

Famous Thought Experiments — The Complete List & Guideen.senkohome.com/thought-experiment-list/