A Study of Language

Learn fascinating things about language. No linguistics degree required.

What Does AI Reveal About Human Language?

For most of its history, linguistics was a discipline that studied human language by observing humans. Fieldwork, corpora, experiments with native speakers, introspective analysis by trained linguists. The tools were careful but slow, and the data was always limited.

Then AI arrived and accidentally became one of the most interesting instruments ever pointed at the structure of language.

A New Kind of Experiment

When a language model learns from text, it runs something like a vast implicit experiment: given this much data and this learning objective, what can be learned? When a model trained on billions of words handles something easily, it tells us that thing can be extracted from text by statistical means. When a model trained on billions of words consistently fails at something, it tells us that thing may require something text alone can’t provide.

This is not how the experiments were designed. But the signal is real, and linguists have started paying attention.

What AI Learns Easily

Modern language models are remarkably good at syntax. They handle complex grammatical constructions, maintain agreement across long distances, and produce sentences that parse cleanly. They do this without being given any explicit grammatical rules — the syntactic regularities emerge from training.

They also handle a great deal of semantic knowledge: the kinds of things that tend to go together, what words are typically used for, how concepts cluster around each other. They can complete analogies, identify category members, and make reasonable inferences based on what’s stated.

For lexical knowledge — the meanings and uses of individual words — they’re extraordinary. A model trained on enough text learns not just denotation but connotation, register, and the social texture of word choice. It knows that “slender” and “scrawny” mean roughly the same thing and are used very differently.

Where AI Consistently Struggles

The failures are just as informative. Language models struggle, in consistent and revealing ways, with things that require what linguists call grounding — connecting language to the physical world, to embodied experience, or to situational context.

Physical reasoning is one weak point. Ask a model to reason about what happens when you put an ice cube in a hot pan, and it may produce a plausible-sounding description. Ask it something more specific — what shape the ice cube is as it melts, what temperature the pan needs to reach, how long this would take — and it quickly reveals that its “knowledge” of this process is derived from descriptions of it, not from any model of what’s physically happening.

Pragmatics is another sticking point. Understanding that “can you pass the salt” is a request, not a question about ability, requires knowing the social context of a dinner table, the conventions of polite requests, and the expectations of the listener. These aren’t written down in text in a way that’s easily extractable. They’re implicit in the situation.

Discourse coherence over long spans remains difficult. A model can produce a paragraph that holds together beautifully. Maintaining consistent reasoning, consistent character, and consistent structure over a 10,000-word document is substantially harder.

What This Says About Language

These patterns reinforce a view that linguists like Steven Pinker have argued for decades: language is not separable from other aspects of mind. The parts of language that AI learns easily — syntax, lexical semantics, surface pragmatics — are things that are well-represented in text. The parts it struggles with — physical grounding, social context, embodied experience — are things that exist largely outside of text.

This suggests that language is not a self-contained system. It rides on top of a great deal of non-linguistic cognition — sensorimotor knowledge, social intelligence, real-world experience — without which much of what language does cannot function. For a thorough account of this argument from the linguistics side, The Language Instinct by Steven Pinker remains essential reading. Pinker’s argument that language is a biological adaptation — not just learned, but built into our cognitive architecture — gains new resonance when you see what gets left behind when a system learns only from text.

The Debate About Universal Grammar

AI has also become an unexpected data point in one of linguistics’ oldest debates: whether there is innate, species-specific linguistic structure — Universal Grammar — or whether language is learned entirely from input.

Language models learn from input alone and reach impressive proficiency. Critics of the strong nativist view point to this as evidence that explicit innate structure isn’t necessary — that a powerful enough learner trained on enough data can do what children do without a built-in grammar.

Nativists respond that the situations aren’t comparable. Children learn from rich, embodied, socially situated input, in a small fraction of the time, with a much smaller total exposure. Language models are trained on orders of magnitude more text, in a purely statistical way, and still fail at things children handle effortlessly.

The debate isn’t settled. But AI has given linguists new tools for sharpening the question.

What This Means for You

AI has become an unintentional probe of human language — and what it reveals is that language is harder, richer, and more deeply grounded in non-linguistic experience than any purely formal account can capture. The places where AI succeeds tell us what can be learned from text. The places where it fails tell us where language stops being about text and starts being about being human.

Disclosure: This article contains affiliate links. If you purchase through one of our links, we may earn a small commission at no extra cost to you.