The Language Problem

Language models can master linguistic structure and participate in human linguistic practice. Whether their words refer to anything beyond that structure is a different question.

Research Abstract
Meaning Is More Than One Relation

Contemporary language models make it possible to separate capacities that ordinary language has long allowed the word “meaning” to conceal within a single category. A system may acquire the differential structure of a language, participate competently in linguistic practice and produce contextually appropriate utterances without thereby establishing the causal or environmental relations through which expressions refer beyond other expressions. This essay examines that separation through competing accounts of signification, reference, use and symbol grounding, asking not whether machine language is simply meaningful or meaningless, but which relations have actually been secured by text-based training. The resulting distinction shifts the problem from behavioural resemblance to conceptual architecture: linguistic competence, social use and referential grounding need not coincide, even where human experience has historically made them appear inseparable.

Can Machines Understand What They Say?

A system trained only to predict the next word in a sequence now produces essays, contracts, poems and technical documentation that native speakers cannot reliably distinguish from human writing under casual reading. The training objective contains no representation of meaning. It specifies a probability distribution over tokens, conditioned on preceding tokens, optimised against a corpus. Everything the system does, it does by satisfying that objective. The question this essay addresses is whether satisfying it is sufficient for the system to mean anything by what it produces, or whether meaning requires a relation that a token-prediction objective cannot in principle establish.

The question has to be split before it can be answered, because “meaning” names several distinct relations that happen to share a word.

The sign and its two sides

Ferdinand de Saussure’s Cours de linguistique générale, assembled from his Geneva lectures and published posthumously in 1916, established the vocabulary still used to separate these relations. A linguistic sign joins a signifier, the sound or written form, to a signified, the concept associated with it. The connection between the two is arbitrary: nothing about the sequence of sounds in “tree” makes it suited to the concept it carries, and other languages carry the same concept with unrelated sounds. What holds the system together is not any sign’s connection to the world but the network of differences among signs. A sign has the value it has because it is not any of the other signs in the language.

This is already sufficient to generate a working notion of linguistic competence that requires no contact with the world at all. A system that has correctly modelled the differential structure of a language, which terms exclude which others, which combinations are permitted, which substitutions preserve acceptability, has captured something real about that language. Distributional semantics, the research programme that treats a word’s meaning as recoverable from the contexts in which it occurs, is a direct descendant of this idea. John Firth’s 1957 formulation, that a word is characterised by the company it keeps, states the operating principle that word embeddings and, later, transformer attention mechanisms implement computationally. A system trained this way learns exactly the kind of structure Saussure described: a network of differential relations among signs, derived entirely from their co-occurrence with other signs.

What this structure does not by itself supply is a connection between the sign and anything outside the network of signs. Gottlob Frege drew the relevant distinction in 1892, in “Über Sinn und Bedeutung”, between the sense of an expression, the mode under which it presents its object, and its reference, the object itself. “The morning star” and “the evening star” have different senses and the same reference, both picking out the planet Venus. A system that has mastered sense relations among terms, knowing that “morning star” and “evening star” occupy similar distributional positions and are used in overlapping contexts, has not thereby been given their reference. Nothing internal to the network of co-occurring symbols fixes which object, if any, a symbol picks out in the world.

Conceptual Architecture
Three Relations Hidden Inside Meaning

Linguistic structure can be learned from relations among signs, while use and reference introduce different kinds of connection. The architecture matters because competence in one relation does not, by itself, establish competence in the others.

01 / Differential relation
Linguistic Structure

Signs acquire value through their differences, combinations and distributions relative to other signs. This relation can be learned within language itself.

02 / Practical relation
Linguistic Use

Meaning may also consist in competent participation within a rule-governed practice in which utterances are taken up, corrected and reused by others.

03 / World relation
Reference

Expressions may stand in causal or environmental relations to objects, substances or conditions beyond the network of linguistic signs.

Structure

Relates signs to other signs.

Use

Relates utterances to communal practice.

Reference

Relates expressions to what they are about.

Two ways reference has been thought to work

Two accounts of how words connect to what they are about have dominated analytic philosophy of language since the middle of the twentieth century, and they place the difficulty in different places.

Willard Van Orman Quine’s Word and Object, published in 1960, argued that reference is radically underdetermined by any finite body of evidence about usage. His thought experiment involves a linguist observing a speaker of an unfamiliar language utter “gavagai” while a rabbit runs past. The utterance is compatible with “rabbit”, but equally compatible with “undetached rabbit part”, “rabbit stage”, or “instance of rabbithood”, since every observable circumstance that would confirm the first hypothesis confirms the others as well. No amount of further behavioural evidence closes the gap, because the hypotheses agree on all observable consequences and differ only in how they carve up what counts as an object. Quine’s conclusion was that translation, and by extension reference itself, is indeterminate in a way no accumulation of usage data resolves.

Saul Kripke’s Naming and Necessity, delivered as lectures in 1970 and published in 1980, argued that reference for proper names and natural kind terms is fixed causally rather than descriptively. A name refers to whatever was picked out in an original act of naming, or “baptism”, and the reference is transmitted along a causal chain of usage from speaker to speaker, independent of whatever descriptions later speakers associate with the name. This account makes reference depend on a causal history connecting a token use of a term to an original contact with the referent, a history that exists outside the text and outside any speaker’s internal representations.

Hilary Putnam’s Twin Earth argument, published in “The Meaning of ‘Meaning'” in 1975, reinforced the same conclusion from a different angle. Putnam imagines a planet identical to Earth except that its oceans, lakes and rain are filled with a substance chemically distinct from H₂O but indistinguishable from water by any test available before the development of chemistry. A speaker on Twin Earth and a speaker on Earth, in 1750, would have had qualitatively identical mental states when using the word “water”, yet the word would have referred to different substances on the two planets, because reference tracks the actual chemical composition of the stuff in the environment rather than anything in the speaker’s head. Putnam’s summary of the conclusion, that meaning is not something that can be individuated by what is inside a single speaker’s mind alone, applies with particular force to a system whose entire existence consists of internal states derived from text.

Both accounts, despite their disagreements, converge on the same implication for a system trained exclusively on text. Reference, on either view, depends on something that does not reduce to patterns among symbols: a causal history of contact with objects, on Kripke’s account, or the actual composition of the environment a term’s users are embedded in, on Putnam’s. A system whose only input has ever been other symbols has no access to either. It has extensive, well-organised information about how words relate to other words. It has no demonstrated route to how words relate to anything else, because nothing else has ever been part of its training signal.

Meaning as use

Ludwig Wittgenstein’s later philosophy offers an account that does not depend on this kind of external connection, and it is worth taking seriously as a genuine alternative rather than a consolation.

The Philosophical Investigations, published in 1953 after Wittgenstein’s death, abandoned the picture he had defended in his own earlier work, the 1921 Tractatus Logico-Philosophicus, according to which a proposition means something by picturing a state of affairs in a determinate correspondence between the structure of the sentence and the structure of the fact. The later work replaces the picture theory with the claim that the meaning of a word, for a large class of cases, is its use in the language. Understanding a word is not a matter of associating it with a private mental image or an external object. It is a matter of being able to deploy it correctly within a “language game”, a rule-governed practice embedded in a form of life shared by a community of speakers.

Wittgenstein’s private language argument, sections 243 to 271 of the Investigations, adds a constraint that matters directly here. A language whose rules could be checked only by a single individual, with no external criterion for correct application, could not sustain a distinction between following the rule and merely believing one had followed it. Correctness requires a check external to the individual instance, which for Wittgenstein meant a community of users capable of correcting one another. Applied to a machine, the argument cuts in two directions at once. It denies that meaning could be secured by a private, unshareable relation between a symbol and an inner state, which rules out one route to grounding that might otherwise have seemed available. But it does not require that the community be human, or that the individual using the term be biological. It requires only a practice, with correction, sustained across multiple participants.

This opens a genuine possibility rather than a merely verbal one. A large language model is embedded in an enormous, ongoing linguistic practice: it was trained on text produced within human language games, and its outputs are now returned into that same practice, read, corrected, cited, and in some cases fed back into the training of successor systems. On a use-based account, a system participating in that practice, producing utterances that are taken up, responded to and corrected by competent speakers, is not obviously disqualified from meaning something by its use of a term, in whatever sense a competence in the practice constitutes meaning. What the account withholds is any guarantee that this participation delivers the same thing that Kripke’s causal chains and Putnam’s environmental composition deliver. Wittgenstein’s criterion and the causal-referential criterion are answers to different questions, and a system can satisfy one without the other.

Comparative Research Plate
Where Meaning Is Located

The theories assembled in the essay do not merely offer different definitions of meaning. They locate the relevant relation in different places: within language, between expression and object, in causal history, in environmental embedding, or in socially governed use.

Differential system
Saussure

Linguistic value arises from differences among signs rather than from any intrinsic bond between a sign and the world.

A model can acquire substantial linguistic structure through relations among signs alone.

Sense and reference
Frege

An expression’s mode of presentation and the object it refers to are distinct relations and cannot be collapsed into one another.

Mastery of relations among expressions does not by itself determine what those expressions pick out.

Indeterminacy
Quine

Behavioural evidence may remain compatible with multiple incompatible schemes of reference.

Additional usage data need not uniquely determine what a term refers to.

Causal reference
Kripke

Reference is fixed through a causal history linking present use to an originating act of naming.

The relevant relation depends upon a history outside the internal organisation of textual representations.

Environmental externalism
Putnam

What a term refers to depends partly on the actual environment in which its users are situated.

Internal states alone do not settle reference where environmental composition is constitutive.

Meaning as use
Wittgenstein

Meaning can consist in competent participation within a rule-governed linguistic practice sustained by public criteria of correctness.

Participation in human linguistic practice may support one legitimate sense of meaning without settling referential grounding.

Syntax and semantics
Searle

Formal manipulation of symbols does not, merely by becoming more extensive, establish what those symbols are about.

Linguistic performance cannot by itself decide whether semantic understanding has been achieved.

Symbol grounding
Harnad

A symbolic system requires some relation to non-symbolic input if definitions are not to circulate indefinitely among symbols.

Token-to-token relations alone leave the grounding problem structurally unresolved.

Central contrast

Saussurean and Wittgensteinian accounts allow meaningful linguistic competence to be characterised through relations internal to language or practice; Kripkean, Putnamian and grounding-based accounts require a further relation beyond those structures.

The syntax objection, applied precisely

Searle’s Chinese Room, introduced in the first essay of this series as an argument about symbol manipulation in general, has its sharpest application here. The room’s occupant manipulates Chinese characters according to formal rules, without any of the characters connecting, from the inside, to anything the occupant would recognise as their referents. The argument’s force in the case of language specifically is that syntactic competence, however extensive, is a matter of relations among symbols, and semantic competence is a matter of relations between symbols and what they are about, and no amount of the first constitutes any of the second.

Stevan Harnad’s 1990 formulation of the symbol grounding problem states the same difficulty in computational terms. A symbol system in which every symbol is defined only by its relations to other symbols is, on Harnad’s description, like trying to learn Chinese from a Chinese-Chinese dictionary: each definition sends the reader to further definitions, and the circle never touches anything outside itself. Grounding, on this account, requires some symbols to be connected to non-symbolic input, sensorimotor contact with the world that the symbol-symbol relations can then be built on top of.

This is where the architecture of current systems matters rather than any of their outputs. A model trained exclusively on text sits inside exactly the closed circle Harnad describes: every one of its parameters was shaped by relations among tokens, and none by an independent channel of contact with the referents those tokens are used to discuss. Multimodal systems, trained jointly on text and images, complicate this picture without resolving it. An image is not the world either. It is a two-dimensional array of pixel values, itself a symbol in Harnad’s sense, encoded by a camera and selected for inclusion in a training set by human curators. Pairing text with images grounds text in images. It does not obviously ground either in the world those images depict, though it is a materially different situation from text alone, and the difference is large enough that the two cases should not be assessed together.

Open Problem
What Would Count as Grounding?
What the essay establishes

Text-only training supplies relations among symbols, while image–text training introduces a materially different form of correlation without automatically establishing contact with the world those representations depict.

What remains open
Which kind of non-symbolic relation would be sufficient?

The argument identifies the absence of demonstrated grounding in text-only systems, but it does not determine what would positively constitute grounding in a machine. Multimodal representation may widen the system’s informational contact without settling whether images, sensors or embodied interaction establish reference rather than merely adding further correlated inputs. A stronger criterion would have to specify what kind of causal relation to an environment is sufficient for symbols to be about what they track, rather than only statistically associated with representations of it.

Required distinction

More channels of input are not by themselves equivalent to a demonstrated relation of reference.

Where performance and comprehension diverge

The second essay in this series distinguished linguistic competence from general reasoning competence on neuropsychological grounds, citing patients whose capacity for one survives damage to the other. The same dissociation is visible from the opposite direction in language models, and it is the most direct evidence available for the question this essay is asking.

Systems that produce fluent, well-formed, contextually appropriate text at a level matching or exceeding most human writers have also been documented, across a substantial body of published evaluation work through the mid-2020s, to fail characteristically on tasks that require tracking what a described situation would actually be like: maintaining consistency about the state of an object across a sequence of described actions performed on it, reasoning correctly about negated or counterfactual claims embedded in otherwise fluent prose, and distinguishing between what a passage states and what merely resembles, distributionally, something the passage might state. Emily Bender and Alexander Koller’s 2020 paper “Climbing towards NLU” formalised the diagnostic underlying these observations with a thought experiment: an octopus intercepting a cable between two humans corresponding by exchanging notes could, by learning the statistical regularities of their exchange, produce responses indistinguishable from a competent participant, right up to the point where the humans start discussing objects in their shared physical environment that the octopus has never had access to, at which point the correspondence breaks down in a way that reveals the octopus never had access to what the words were about.

The pattern matches the theoretical diagnosis exactly. Structure without grounding predicts fluent local coherence and predicts failure precisely at the points where correct output requires tracking a referent rather than continuing a pattern. That is what has been observed. It is stronger evidence than any single benchmark score, because it is evidence about where the competence breaks, not merely how often.

The honest answer

The question posed at the outset does not have a single answer, because it is at least two questions wearing one sentence.

If the question is whether a system trained on text alone stands in the causal or environmental relations that Kripke and Putnam take reference to require, the answer available from current architecture is no, and no plausible amount of additional text changes that, because the relation in question is not the kind of thing more text could establish. If the question is whether such a system can be meaningfully described as participating in a linguistic practice, in Wittgenstein’s sense, contributing utterances that are taken up, corrected and built upon by a community of users, the answer is closer to yes, with the caveat that this form of meaning was never guaranteed to coincide with the referential kind, and the divergence between fluent structure and reliable tracking of described situations is exactly what should be expected if it does not.

Both answers can be true simultaneously because they concern different relations that ordinary usage bundles into a single word.

TMQ Proposition
Artificial intelligence makes visible a separation that human language historically concealed: linguistic competence, participation in linguistic practice and referential grounding are distinct relations, and success in one does not by itself establish the others.

The proposition does not claim that machine language is meaningless. It formalises the narrower distinction required once fluent linguistic performance can occur without demonstrated causal or environmental contact with the things language is about.

Domain
Language Use Reference Grounding

The bundling was harmless as long as every system capable of producing well-formed language was also, by virtue of being human, embedded in causal contact with the world its language was about. That correlation no longer holds without exception, and the word has to be unbundled to say anything precise about what has changed.

Critical Apparatus

References and Intellectual Lineage

The argument moves across three connected traditions: structural accounts of linguistic value, philosophical theories of reference and use, and computational arguments about syntax, grounding and natural-language understanding. Read together, the sources trace the widening distance between competence within a symbolic system and a symbol’s relation to what lies beyond it.

10 Verified sources
03 Research strata
01 Reference unresolved
01

Structure, Sense and Reference

01
Structural Linguistics
Ferdinand de Saussure. 1916.
Cours de linguistique générale.

Edited by Charles Bally and Albert Sechehaye, with Albert Riedlinger. Lausanne and Paris: Payot.

Role in the argument Establishes the signifier/signified distinction and the differential structure through which linguistic signs acquire value within a system.

02
Distributional Linguistics
J. R. Firth. 1957.
“A Synopsis of Linguistic Theory, 1930–1955.”

In Studies in Linguistic Analysis, 1–32. Special volume of the Philological Society. Oxford: Blackwell.

Role in the argument Supplies the contextual and collocational conception of linguistic meaning from which the essay connects structural linguistics to later distributional approaches.

03
Sense and Reference
Gottlob Frege. 1892.
“Über Sinn und Bedeutung.”

Zeitschrift für Philosophie und philosophische Kritik 100: 25–50. English translation, “On Sense and Reference”, in Translations from the Philosophical Writings of Gottlob Frege, edited by Peter Geach and Max Black, Oxford: Blackwell, 1952.

Role in the argument Provides the distinction between sense and reference that allows mastery of relations among expressions to be separated from successful reference to objects.

02

Reference, Use and Externalism

04
Indeterminacy of Reference
Willard Van Orman Quine. 1960.
Word and Object.

Cambridge, MA: The MIT Press.

Role in the argument Introduces radical translation and the indeterminacy of reference through the “gavagai” case, showing that behavioural evidence does not uniquely determine what a term refers to.

05
Causal Reference
Saul A. Kripke. 1980.
Naming and Necessity.

Cambridge, MA: Harvard University Press. Based on lectures delivered at Princeton University in 1970; an earlier version appeared in Semantics of Natural Language in 1972.

Role in the argument Supplies the causal-historical account against which reference cannot be reduced to descriptions or relations internal to a symbolic system.

06
Semantic Externalism
Hilary Putnam. 1975.
“The Meaning of ‘Meaning’.”

In Keith Gunderson, ed., Language, Mind, and Knowledge, Minnesota Studies in the Philosophy of Science, vol. 7, 131–193. Minneapolis: University of Minnesota Press.

Role in the argument Uses the Twin Earth case to establish the externalist pressure at the centre of the essay: identical internal states need not determine identical reference.

07
Meaning as Use
Ludwig Wittgenstein. 1953.
Philosophical Investigations.

First published posthumously in 1953. The current fourth edition, edited by P. M. S. Hacker and Joachim Schulte, with the revised English translation by G. E. M. Anscombe, P. M. S. Hacker and Joachim Schulte, was published by Wiley-Blackwell in 2009.

Role in the argument Supplies the use-based alternative to causal reference and the rule-following framework through which linguistic participation can be considered independently of private mental representation.

08
Picture Theory of Language
Ludwig Wittgenstein. 1922.
Tractatus Logico-Philosophicus.

English translation by C. K. Ogden, with assistance from F. P. Ramsey. London: Kegan Paul, Trench, Trubner & Co.; later Routledge & Kegan Paul. The German text first appeared in 1921 as “Logisch-Philosophische Abhandlung”.

Role in the argument Provides the earlier picture-theoretic position against which the later Wittgenstein’s account of meaning as use is defined.

03

Computation, Grounding and NLU

09
Syntax and Intentionality
John R. Searle. 1980.
“Minds, Brains, and Programs.”

Behavioral and Brain Sciences 3 (3): 417–424. Cambridge University Press. DOI: 10.1017/S0140525X00005756.

Role in the argument Supplies the Chinese Room argument used to sharpen the distinction between formal symbol manipulation and semantic or intentional competence.

10
Symbol Grounding
Stevan Harnad. 1990.
“The Symbol Grounding Problem.”

Physica D: Nonlinear Phenomena 42 (1–3): 335–346. DOI: 10.1016/0167-2789(90)90087-6.

Role in the argument Formalises the grounding problem computationally: a symbolic system whose symbols are connected only to further symbols cannot make their semantic interpretation intrinsic merely by extending the symbolic network.

11
Natural-Language Understanding
Emily M. Bender and Alexander Koller. 2020.
“Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data.”

In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185–5198. Association for Computational Linguistics. DOI: 10.18653/v1/2020.acl-main.463.

Role in the argument Transposes the form–meaning distinction directly into contemporary NLP and supplies the octopus thought experiment used to test whether statistical command of linguistic form entails access to what utterances are about.

References Requiring Verification

A
Mid-2020s language-model evaluation literature

The essay refers collectively to published evaluations documenting failures in object-state tracking, negation, counterfactual reasoning and the distinction between stated and distributionally plausible content. No individual studies are named, so no specific bibliographic records are assigned here without further source identification.

The Machine Question is an independent research project by Paolo Calvi.