The Machine Question

Artificial intelligence does not simply create new machines. It forces us to reconsider what intelligence means, how we define understanding, and which assumptions about human cognition remain valid in an age of artificial systems.

Why Artificial Intelligence Changes the Meaning of Intelligence

THE MACHINE QUESTION — EDITORIAL PREMISE

This essay does not attempt to determine whether artificial systems are already intelligent, conscious, or capable of understanding. Its purpose is different: to examine how artificial intelligence challenges the concepts through which intelligence, reasoning, creativity and knowledge have historically been defined.

The central question is therefore not only what machines can do, but what human societies mean when they recognise something as intelligent. Artificial intelligence is approached here as an epistemological event: a technology that reveals the assumptions, boundaries and institutional functions hidden inside one of our most fundamental categories.

The question that dominates public discussion of artificial intelligence is whether machines can think. It is an old question. Descartes addressed it in 1637, in the fifth part of the Discourse on the Method, where he proposed two tests that no mechanism would ever pass: the flexible use of language in unforeseen situations, and the general applicability of reason across unrelated domains. Leibniz addressed it in 1714, in section seventeen of the Monadology, with the image of a mind enlarged to the size of a mill: a visitor walking through it would find only parts pushing against other parts, and nothing that could be called perception. La Mettrie addressed it in 1747 by denying the premise, arguing in L’homme machine that the human being was already a mechanism and that the question therefore dissolved. Three centuries of argument have produced no settled answer, and the arrival of systems that generate fluent language has not produced one either.

01 / ABSTRACT

Abstract

Artificial intelligence has transformed an enduring philosophical question into a contemporary problem of definition. This essay examines how artificial systems challenge the conceptual frameworks through which intelligence has historically been understood, shifting the focus from whether machines can think to what intelligence itself represents. From psychometric measurement and symbolic AI to contemporary language models, intelligence emerges not as a fixed natural property but as a changing category shaped by scientific practices, institutional needs and technological developments. The essay explores the distinction between performance and understanding, the limits of behavioural criteria, and the consequences of defining what counts as intelligence. The Machine Question is therefore not only an inquiry into artificial systems, but an examination of the human concepts through which intelligence is recognised, classified and governed.

What has changed is the position of the question. For most of its history it was counterfactual. It asked what would follow if a machine could do certain things, and the things in question stayed safely hypothetical. That condition no longer holds. Systems now perform a substantial share of the operations that the concept of intelligence was built to cover, and they perform them without any of the properties that the concept was assumed to require. The interesting consequence is not what this tells us about the machines. It is what it exposes about the concept.

The category was never stable

Intelligence is often treated as a natural kind, a property that exists in the world and that we discovered and named. Its history suggests something closer to an administrative artifact, assembled at particular moments for particular purposes.

The modern measurable version dates to the first years of the twentieth century. Charles Spearman published his analysis of correlated mental test scores in 1904 and proposed a general factor underlying them. Alfred Binet and Théodore Simon produced their scale in 1905, commissioned by the French ministry of public instruction to identify children who required separate schooling. The category was operationalised in response to an institutional need: sorting populations. Binet himself resisted the interpretation of his scale as a measure of a fixed innate quantity, and lost that argument to the psychometric tradition that followed him.

Fifty years later the category was operationalised again for a different institutional need. The Dartmouth Summer Research Project proposal, written in 1955 by John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon, introduced the term “artificial intelligence” and framed the enterprise around a conjecture: that every aspect of learning and intelligence could be described precisely enough for a machine to simulate it. The proposal was a funding document. Its central move was to make intelligence into an engineering target, which required treating it as decomposable into functions. Twenty years later Allen Newell and Herbert Simon gave that assumption its canonical form in the physical symbol system hypothesis, which held that a physical symbol system has the necessary and sufficient means for general intelligent action.

Both operationalisations were productive and both were partial. Psychometrics produced a scalar suited to ranking people. Symbolic AI produced a functional decomposition suited to building programs. Neither was a theory of what thinking is. The word carried on doing philosophical work that its operational definitions never supported, and the gap between the two was invisible for as long as no machine could occupy it.

Turing’s substitution and what it cost

Alan Turing’s 1950 paper in Mind is usually read as an argument that machines can think. It is closer to an argument that the question should be abandoned. Turing observed that answering it would require definitions of “machine” and “think” derived from ordinary usage, judged that this route led nowhere, and replaced the question with a different one: whether an interrogator communicating through text could reliably distinguish a machine from a human being.

The substitution was disciplined. It converted an unanswerable metaphysical question into a decidable empirical one, and it did so by ruling that the only admissible evidence was behavioural. Everything about the inner constitution of the system was declared irrelevant, on the reasonable ground that we extend the attribution of mind to other human beings on exactly the same terms, having no access to their interiors either.

The cost of the substitution was deferred rather than avoided. Turing’s test works as a criterion only while behavioural competence and the properties we care about remain correlated. He assumed that a system able to sustain open-ended conversation would have to possess whatever it is that makes conversation possible in us. That assumption was empirical, and it has now been tested. Contemporary language models are trained on a statistical objective, predicting continuations of text, using the transformer architecture introduced by Vaswani and colleagues in 2017. They produce extended, contextually appropriate, syntactically well-formed discourse across arbitrary subject matter. Whether they satisfy Turing’s original protocol under rigorous conditions remains contested, and the protocol has been diluted in most popular retellings. What is clear is that the correlation Turing relied on has weakened. Behavioural competence has become cheap relative to the properties it was taken to indicate.

The syntax objection and its limits

John Searle’s 1980 paper in Behavioral and Brain Sciences supplied the standard counter-argument. The Chinese Room describes a person following rules to manipulate Chinese characters without understanding Chinese, producing responses indistinguishable from those of a competent speaker. Searle’s conclusion was that formal symbol manipulation is insufficient for semantic content, and that a program, being purely formal, cannot therefore constitute a mind.

The argument identified a real distinction. Claude Shannon had already established the engineering version of it in 1948, when he set aside the semantic aspects of communication as irrelevant to the transmission problem. Frege had established the philosophical version in 1892, separating the sense of an expression from its reference. Information can be handled correctly without being understood, and every working communication system depends on that fact.

The limits of Searle’s argument matter as much as its force. It establishes that syntactic manipulation alone does not entail understanding. It does not establish what would entail understanding, nor does it supply a test by which the presence of understanding could be detected in any system, including a human one. Searle located the difference in the causal powers of biological brains, which relocates the problem into neurophysiology without resolving it. Thirty-five years of subsequent debate has produced candidate criteria for genuine comprehension, among them intentionality, embodiment, causal grounding in a world, and biological substrate. Each is defensible. None has produced an operational test that a third party could apply.

This is the position we are actually in. We have a strong intuition that simulation and comprehension differ. We have no procedure for detecting the difference from outside. The intuition may be correct and still be useless for any purpose that requires adjudication.

CONCEPTUAL MAP

How Intelligence Has Been Redefined

Each technological breakthrough has not only expanded machine capabilities. It has also forced a revision of what humans considered intelligence.

Period
Dominant conception of intelligence
What changed
Early 20th century
Intelligence as a measurable human capacity. Psychometrics transformed cognition into a score.
Intelligence became an instrument for classification and comparison, shaped by institutional needs.
1950s–1970s
Intelligence as symbolic reasoning. AI was framed as the decomposition of thought into formal operations.
Human reasoning was reinterpreted as a set of functions that could potentially be reproduced by machines.
1990s–2010s
Intelligence as performance under difficult tasks. Chess and Go became symbolic benchmarks.
Once machines succeeded, the definition of intelligence shifted away from the conquered domain.
Present
Intelligence as a contested category. Language models challenge the link between competence and understanding.
The central question is no longer only what machines can do, but what humans mean by intelligence itself.

The reclassification reflex

There is a recurring pattern in the history of the field, sharp enough to function as evidence about the concept rather than about the technology.

Chess was treated for decades as a paradigm case of intelligence, on the reasoning that it demanded foresight, strategy and abstraction. Deep Blue defeated Garry Kasparov in May 1997. Within a short period the standard interpretation was that chess had turned out to be a search problem, and therefore never a good test of intelligence at all. Go was then advanced as the successor case, on the grounds that its branching factor made brute-force search useless and that strong play required something closer to intuition. AlphaGo defeated Lee Sedol four games to one in Seoul in March 2016. The same reinterpretation followed. Language production was then advanced as the frontier, and the reinterpretation is now under way for the third time.

Douglas Hofstadter recorded the pattern in 1979 in Gödel, Escher, Bach, citing a formulation he attributed to Larry Tesler: artificial intelligence is whatever has not been done yet. The formulation is usually quoted as a joke about moving goalposts. It is better read as a report about the structure of the concept. “Intelligence” has functioned as a residual category, defined by whatever human cognitive performance has resisted mechanisation at a given date. A category defined by a shrinking remainder has no fixed content, and its apparent stability across the twentieth century was an artefact of how little had yet been mechanised.

The reflex has a symmetrical form that receives less attention. Hans Moravec observed in 1988 that the tasks we find effortful, such as formal reasoning and calculation, submitted to automation early, while the tasks we find effortless, such as perception and locomotion, proved extraordinarily difficult. Our intuitive ranking of cognitive difficulty tracks introspective effort, and introspective effort tracks evolutionary recency rather than computational complexity. The ranking we use to identify what is most intelligent about us is, on this evidence, close to inverted.

What the question should be

The question “can machines think” presupposes that the extension of “thinking” is fixed and that the only open matter is whether machines fall inside it. Both parts of the presupposition fail. The extension has been revised repeatedly, each revision has been retrospective, and each has been driven by what machines had just achieved. Asking whether machines think is therefore a way of asking what we have most recently decided thinking excludes.

The productive question is what we are naming when we use the word. That question has three components, and they are separable.

The first is descriptive. Cognitive competence is not one capacity. Pattern completion, inference under uncertainty, goal maintenance, error correction, model construction, self-monitoring and the coordination of all of these across time are distinct operations with distinct computational profiles. Current systems achieve some of them at very high levels and others hardly at all. Nothing forces the assumption that they cluster in machines the way they cluster in us, and the assumption that they must has produced most of the confusion in public discussion.

The second is constitutive. Whether comprehension requires something beyond functional organisation remains open, and the arguments run in both directions with roughly comparable strength. Wittgenstein’s position in the Philosophical Investigations, published in 1953, that meaning is constituted by use within a practice, points toward a functional answer. Searle’s position points away from it. The disagreement is not resolvable by inspecting the systems, because the criteria that would settle it are exactly what is in dispute.

The third is institutional, and it is the one that will be settled first, because institutions cannot wait for philosophy. Legal systems are already allocating rights on the basis of implicit answers. Copyright regimes in several jurisdictions have anchored protection to human authorship, which requires a working distinction between generation and creation. Liability regimes require a working distinction between a tool and an agent. The European Union’s AI Act, in force since 2024 and applying in phases, classifies systems by risk and imposes obligations accordingly, which requires a working account of autonomy. Labour markets are repricing occupations according to which of their component tasks have been mechanised, which requires a working account of what the human contribution consists of. In each case a definition is being fixed, and the definitions are being fixed by parties with material interests in where the lines fall.

That last point deserves stating plainly, because it is usually left implicit. Definitions of intelligence are not neutral instruments. A firm that describes its system as reasoning attracts capital and regulatory latitude that a firm describing its system as interpolating does not. A firm that describes the same system as a tool rather than an agent limits its exposure to liability. A profession that defines its expertise in terms of a capability machines lack retains its rents. The question of what counts as thinking is being answered in the middle of a distributive conflict, by participants in that conflict, and the philosophical literature is being cited selectively by all sides.

The premise of this project

Artificial intelligence is treated here as an epistemological event. For the first time, artefacts perform functions that a concept was constructed to describe, without possessing the properties that the concept was assumed to entail. The concept has consequently lost its capacity to sort cases, and the disagreements it generates are no longer resolvable by appeal to it.

The response taken here is to examine the concept rather than to defend a verdict about the machines. That means treating “intelligence”, “understanding”, “creativity”, “reasoning” and “knowledge” as objects of analysis with histories, functions and beneficiaries, and asking in each case what work the term does, which distinctions it was built to draw, which of those distinctions it can still draw, and what follows for the institutions that depend on it.

Two positions are ruled out at the start. The first holds that current systems are already minds and that scepticism reflects human vanity. The second holds that current systems are trivial machinery and that the appearance of competence is a category error committed by naive observers. Both convert an open question into a settled one and both are held with a confidence the evidence does not support.

The work proceeds by cases. What follows examines specific systems, specific claims made about them, specific decisions being taken on the basis of those claims, and the specific interests those decisions serve. Where the evidence supports a conclusion it will be stated. Where it does not, the uncertainty will be stated instead.

Critical Apparatus

References and Intellectual Lineage

The following sources document the philosophical, scientific and historical framework underlying this essay and the development of the concept of intelligence across different intellectual traditions.

01 — Primary Sources
01
Foundational Philosophy
René Descartes. 1637.
Discourse on the Method.
Original philosophical text.
02
Artificial Intelligence
Alan Turing. 1950.
Computing Machinery and Intelligence.
Mind, 59(236), 433–460.
03
Philosophy of Mind
John R. Searle. 1980.
Minds, Brains, and Programs.
Behavioral and Brain Sciences, 3(3), 417–457.
02 — Scholarly References
04
Research Paper
Ashish Vaswani et al. 2017.
Attention Is All You Need.
Neural Information Processing Systems.
05
Cognitive Science
Douglas Hofstadter. 1979.
Gödel, Escher, Bach: An Eternal Golden Braid.
Basic Books.
03 — Institutional Sources
06
Regulation
European Union. 2024.
Artificial Intelligence Act.
European regulatory framework for artificial intelligence.

The Machine Question is an independent research project by Paolo Calvi.