Loading...

What Do Large Language Models Do Then?

July 2026

Key Idea:

The machine can produce something that looks remarkably like thought, but its underlying mechanism is different.

It finds patterns across humanity’s accumulated knowledge and recombines them into new ideas, explanations, and conclusions. In that sense, it is best understood as an instrument that amplifies human reasoning—not as a reasoner itself.

I label this concept "associate synthesis" to distinguish it from "reasoning".

The last article ended with a simple conclusion:

the machine doesn’t reason.

But that answer only tells us what it lacks. It doesn’t tell us what the machine actually does.

And something is clearly happening. The technology is useful enough that millions of people rely on it every day to draft documents, summarize information, explain complex topics, analyze problems, and support decisions.

Saying “it doesn’t reason” leaves the most important question unanswered:

if it isn’t reasoning, then what is it doing?

There are two common answers to fill that gap, and neither is very satisfying.

One view is that the machine is essentially a “stochastic parrot”—recombining fragments of text it has seen without any deeper capability. The other treats it as a thinking machine, or something close to one. The previous article addressed the "thinking machine" view, which I refute.

This article takes a different approach. First, we’ll describe what the machine actually does. Then we’ll compare that behavior with what we mean by reasoning and determine how much of that capability it actually reproduces.

Two questions, in order:

What is it?

And how much of reasoning does it actually perform?

Neither Parrot Nor Mind

Start with how the technology actually works, rather than trying to force it into either of two misleading categories. As Tayyar Madabushi and colleagues argue in their paper, it is neither a “stochastic parrot” nor AGI.[1]

“Parrot” undersells what the technology can do. A parrot repeats; an LLM can generalize from the patterns it learned during training.

(Side note: I helped build a 250k parameter LSTM neural network in 2015 that could generate decent english language. Whatever we're doing at 100 billion parameters is different than what we were doing then.)

When you give it a prompt, the model uses that context to identify relevant patterns from its training and generate a response that extends those patterns to a new situation. That is why it can answer questions that were never explicitly written down in its training data, and why the same model can work across domains such as law, software development, or cooking simply by changing the context of the request.

Simple memorization cannot explain that behavior.

“Mind” goes too far, for the reasons the last article laid out. Look at how the system actually works, and none of the four components of reasoning is performing its intended function.

What we have instead is pattern extension over language: the model identifies patterns in its training and extends them to the context in front of it. That can produce remarkably useful results, but it is fundamentally different from reasoning about why something is true.

The simplest description sits between these two extremes: the model predicts the next word, repeatedly, based on patterns learned from a vast amount of human-written text. Viewed this way, both its strengths and its failures become easier to understand. A hallucination occurs when the model generates a plausible continuation without enough reliable information in the prompt or its learned patterns to anchor the answer.[2]

The point is not to criticize the machine. It is to describe how it works.

But “a pile of human text” still doesn’t fully describe what the model has learned, because the text itself has structure. It isn’t random. Much of it is the written record of people explaining, analyzing, arguing, solving problems, and making decisions.

That distinction matters.

The model is not simply learning language;

it is learning patterns from the accumulated record of human reasoning.

That is where the more useful description begins.

The Appearance of Knowledge

This is where the earlier work in the series becomes important.

Writing is not simply a container for reasoning. It is the mechanism that allows reasoning to persist. Writing turns thoughts and conclusions into durable records that can be revisited, examined, challenged, and built upon. Law, science, and scholarship have advanced for centuries through this process: people write down what they know, others examine it, and new reasoning builds on the record.

That is the material on which these models are trained. The training corpus is, to a remarkable extent, a compressed record of human reasoning preserved in written form.

So what the model produces is neither reasoning nor gibberish.

It is reasoning-shaped text:

new writing generated by extending patterns found in the written record of human reasoning.

That is why its outputs so often take the form of premises, conclusions, “therefore” statements, and numbered steps.

These structures are common in the material the model learned from. The resemblance to reasoning is neither a trick nor a coincidence. It is the natural result of extending patterns found in a written record created through human reasoning.

What the model does not do is reproduce the process that created that record.

It can produce the products of reasoning without performing the work of reasoning itself. It does not independently establish that a claim is true, test its conclusions against the evidence, or revise its position when new evidence contradicts it.

Floridi and colleagues describe the distinction in a particularly useful way.[2] The model can produce the output of an explanation—a coherent and plausible account—without performing the act of explanation, which requires testing that account against the world. Quattrociocchi and colleagues describe the resulting problem as well.[3]

Plausibility can take the place of verification, creating the appearance of knowledge without the underlying process of judgment.

That is the distinction to carry forward: reasoning-shaped text—the products of reasoning, generated from its written record, without necessarily performing the work that made those products reasoned in the first place.

An Instrument, Not an Agent

Before we examine what the model actually does, there is one practical question to answer:

if the output looks like reasoning, but the model itself does not reason, where does that reasoning come from?

The answer is: from humans.

Most of it is inherited from the training data. The logical structure in a model’s output reflects patterns created by thousands—or millions—of people whose arguments, explanations, analyses, and solutions became part of that corpus. When an output moves coherently from premise to conclusion, it is often drawing on relationships that humans established in the underlying language.

But recombination is not the same as reasoning.

The model can combine those patterns into a genuinely new argument without independently evaluating the premises, testing the conclusion against evidence, or revising the argument when it fails.

It reproduces the structure of reasoning without necessarily performing the reasoning that created that structure.

The rest comes from the human using the system.

A person is responsible for deciding whether the claims are true, whether the evidence supports the conclusion, and whether the output is good enough to stand behind. The person reads the result, verifies it, corrects it when necessary, and ultimately takes responsibility for what is produced.

That is the missing part of the loop. The machine can generate an argument, but a human must evaluate it and decide whether it should be accepted.

That leads to the more accurate description: the model is an instrument, not an agent. It extends human capability, much like a telescope extends the reach of the eye.

The instrument provides greater capability, but the human still has to interpret what it produces.

Who supplies that missing judgment—and what happens when nobody does—is the subject of future sections of this book.

How Much of Reasoning Is That?

We have now described what the machine does. The next question is how much of that actually constitutes reasoning. To answer it, we can take the definition of reasoning piece by piece and examine which parts the model actually performs.

Start with the core distinction. Reasoning occurs when one thought provides a reason for another—when it justifies the next step, rather than simply coming before it.

The model works differently.

Its next step is determined by what statistically predicts the next step given the context. (note: Transformers learn a latent space of relationships between patterns, expanding the range of things they can generate beyond what they encountered verbatim during training. More on this in a future article.)

That is essentially the same "one-thing-leads-to-another" association that the definition of reasoning was designed to distinguish from genuine justification—the way "seeing your cousin’s BMW might remind you to call Fred".

The important point, then, is not simply that the model falls short of reasoning. Its basic operation belongs to the category that the definition of reasoning was designed to distinguish from reasoning in the first place.[4]

The four components of reasoning do not fare much better. The last article covered the evidence in detail, so the conclusion here can be brief.

A model can generate language that expresses belief, but that signal does not reliably change how the model behaves.[5] It can produce logical connections between statements, but studies have found little evidence that it consistently represents those logical relationships as something it can independently maintain and evaluate. It can state conclusions with confidence, but there is no underlying commitment that forces the conclusion to remain consistent with the evidence.

And it can produce the language of accountability—“I conclude,” “I recommend,” “I believe”—without actually being responsible for the result.

In each case, the form of reasoning is present, but the function is not.

The three major forms of inference also separate in important ways. The model can imitate the structure of deduction, but its ability to carry out reliable deductive reasoning is fragile: changing the symbols or structure of a problem can cause it to fail.[6] When exact, verifiable deduction matters, the more reliable approach is to delegate the calculation or logical check to an external system—the kind of deterministic engine discussed earlier in the series.

Explanation has the same limitation described above: the model can produce a coherent explanation, but it does not independently test that explanation against reality.

That leaves induction—the one form of inference where the model's capabilities are meaningfully different. It deserves its own section.

The One Part It Really Does

There are two parts of reasoning the model genuinely performs. Both come with important limitations.

The first is induction: reasoning from many examples toward a general pattern. That is essentially what happens during model training. The system learns from an enormous body of text and generalizes from those examples to identify patterns in the language. In that sense, statistical learning is a form of induction.[7]

But the fundamental limitation of induction remains the same one David Hume identified centuries ago: no amount of past observations can guarantee what will happen next.[8]

But still --- a pattern learned from the past can be useful without being universally true.

So induction is the clearest case where the model genuinely performs a form of reasoning. More precisely, the model is the result of induction performed at a scale no individual could match.

But three limitations matter.

First, the induction happens primarily over text. The model learns patterns in what people write, rather than learning directly from the world itself. It can therefore learn strong statistical associations between events without necessarily establishing which events actually cause others.[9]

Second, the underlying learned model is largely fixed once training is complete. A person can encounter new evidence, update their beliefs, and carry those changes forward. A deployed model can also consider new information provided in its context—what is commonly called "in-context learning"—and its response can change accordingly. But this is not the same as updating the underlying model. The new information conditions the model's output for that interaction; it does not normally change the learned parameters or permanently incorporate the new evidence into what the model has learned.

That distinction matters for reasoning. A person can revise what they believe based on new evidence and carry that revision into future judgments. A deployed model can reason with new information placed in its context, but the inductive process that established its underlying model occurred before the conversation began.

Third, the model cannot independently evaluate or defend its own inductive conclusions. It can state why an answer might be correct, but that explanation is itself another generated output rather than an independent check on the inference.

What remains is a limited form of induction: a system that has generalized patterns from an enormous body of text, but does not independently test those patterns against the world, update its underlying model as new evidence arrives, or take responsibility for the conclusions it produces. It is Hume’s problem, mechanized—and applied primarily to language rather than direct experience.

The second piece is the social side of reasoning: evaluating arguments made by others. Reasoning is not only about constructing an argument; it is also about challenging, testing, and rejecting arguments when the evidence does not support them. The model is much better at the first task than the second.[10]

That imbalance matters.

A system that can generate convincing arguments without an equivalent built-in process for challenging them can produce explanations that sound authoritative without being grounded in truth. Research has found that fluent AI-generated explanations can actually increase people’s confidence in incorrect answers rather than reduce it.[11]

This means that answers that sound "plausibly correct" yet wrong --- are sometimes more dangerous than ones that are easy to spot as wrong.

The missing skepticism therefore has to come from the human reader. That creates its own problem: people naturally evaluate arguments in a social context, where the speaker is assumed to have knowledge, experience, or something at stake.

A language model has none of those properties, but its fluency can make its output feel as though it does.

What Should We Call It?

The industry has already chosen a name for this capability: “reasoning.”

The latest models are marketed as "reasoning models", and the intermediate steps they generate before an answer are often called reasoning traces.

The terminology is convenient, but it can also be misleading. It suggests that these traces are records of the model’s actual thinking and that the model is performing the same kind of reasoning we described earlier. Theses assumptions don't hold up.

As the last article showed, a reasoning model’s trace is better understood as part of the mechanism used to produce a better answer—not as a transparent record of how the answer was reached. In fact, research has found that the traces most useful for training future models can be the least understandable to human readers.[12]

That distinction matters operationally. A reasoning trace should not automatically be treated as evidence that the model has independently reasoned through a problem, verified its conclusions, or identified its own mistakes.

Dropping the word “reasoning” should not mean replacing it with a dismissive label. Terms like “stochastic parrot,” “pseudo-reasoning,” or “counterfeit” correct one overstatement by creating another. They understate a technology that can perform an enormous amount of useful work.

In seconds, a model can process and connect information across more written material than a person could read in a lifetime. It can identify relevant patterns and recombine them into a useful draft, analysis, explanation, or recommendation. Whatever we call this capability has to account for that.

So instead of defining the technology by what it lacks, name what it actually does.

Earlier, we called the output reasoning-shaped text: the structure of an argument without necessarily performing the reasoning that produced it. Now we can name the process that generates that output.

The underlying operation is association at scale. The model moves from one piece of information to the next based on statistical relationships learned from an enormous body of human writing, then synthesizes those patterns into something new.

Associative synthesis is the more useful name for this capability: searching a vast space of human knowledge through learned associations and synthesizing those patterns into new, usable outputs.

So name what it actually does.

Earlier, we named the output from the perspective of what it lacks: reasoning-shaped text—the form of an argument without necessarily performing the reasoning behind it. Now name the activity that produces that output, and the picture changes from what is missing to what is actually happening.

Underneath the fluent output is association at scale. The model moves from one piece of information to the next based on statistical relationships learned from an enormous body of human writing, rather than by independently establishing that one claim justifies another.

The more accurate name for this process is associative synthesis:

the model searches a vast space of human writing through learned associations and synthesizes those patterns into something new.

Both words matter.

“Associative” keeps the description precise. Association is the fundamental operation: the model moves between patterns based on learned statistical relationships rather than independently establishing that one claim justifies another. The term identifies the distinction without diminishing what the system can do.

“Synthesis” captures the other half of the capability. The model can pull relevant patterns from an enormous body of information and combine them into a coherent draft, analysis, or explanation. That is meaningful cognitive work—work that can save a person hours.

The name also clarifies the machine’s role. Associative synthesis does not reason for you; it amplifies your ability to reason. It is a cognitive amplifier: a tool designed to extend human intellectual capability rather than replace human judgment.[13]

The human supplies the belief, judgment, and accountability. The machine supplies speed, scale, and access to patterns across an extraordinary amount of written knowledge.

Used this way, an associative-synthesis engine is not a substitute for human reasoning. It is a powerful instrument for extending what a person can accomplish with it.

The Work Still Belongs to Us

So, both questions.

What does it do? Associative synthesis: it searches the patterns in humanity’s written reasoning and recombines them into something new. It is an instrument that amplifies a reasoner, not an agent that reasons on its own.

How much of reasoning does that cover? The model reproduces the shape of nearly all of it, but reliably performs the substance of very little. There is one genuine exception: induction, performed at enormous scale but primarily over text rather than direct experience, and without independently reflecting on or updating the conclusions it produces. There is also a half-exception: the ability to construct arguments without reliably evaluating them.

Put plainly, the machine operates through the exact kind of one-thing-leads-to-another association that What Is Reasoning? defined reasoning against—only now it does so with a fluency and scale that history had never prepared us for.[14] The result can look remarkably like careful thought.

That is neither a small capability nor a disappointing one. Fast, fluent pattern work across an enormous body of human knowledge is a genuinely powerful tool. Naming it precisely is not about diminishing what the machine can do. It is about understanding how to use it well.

The missing part of reasoning still has to come from a person: someone must evaluate the output, close the loop, and take responsibility for the result. Treat the machine as an agent, and you risk assigning that responsibility to something that cannot assume it.

Which raises the question I address in future articles:

who supplies what the machine lacks, and can that loop ever be closed without a person in it?


References

  1. Tayyar Madabushi, H., Torgbi, M., & Bonial, C. (2025). Neither stochastic parroting nor AGI: LLMs solve tasks through context-directed extrapolation from training data priors. arXiv:2505.23323.
  2. Floridi, L., Morley, J., Novelli, C., & Watson, D. (2025). What kind of reasoning (if any) is an LLM actually doing? On the stochastic nature and abductive appearance of large language models. arXiv:2512.10080.
  3. Quattrociocchi, W., Capraro, V., & Perc, M. (2025). Epistemological fault lines between human and artificial intelligence. arXiv:2512.19466.
  4. Adler, J. E., & Rips, L. J. (Eds.). (2008). Reasoning: Studies of Human Inference and Its Foundations. Cambridge University Press.
  5. Sanyal, D., Pandey, M., Kumar, D., Deshpande, S., & Mandal, M. (2025). Confidence is not competence: A mechanistic look at the decoupling of belief and action in LLMs. arXiv:2510.24772.
  6. Tang, X., et al. (2023). Large language models are in-context semantic reasoners rather than symbolic reasoners. arXiv:2305.14825.
  7. Vapnik, V. N. (1999). An overview of statistical learning theory. IEEE Transactions on Neural Networks, 10(5), 988–999.
  8. Hume, D. (1748/1977). An Enquiry Concerning Human Understanding. Hackett.
  9. Zecevic, M., Willig, M., Dhami, D. S., & Kersting, K. (2023). Causal parrots: Large language models may talk causality but are not causal. Transactions on Machine Learning Research, 08/2023. arXiv:2308.13067.
  10. Mercier, H., & Sperber, D. (2011). Why do humans reason? Arguments for an argumentative theory. Behavioral and Brain Sciences, 34(2), 57–74. (See also The Enigma of Reason, Harvard University Press, 2017.)
  11. Palod, V., Biswas, U., & Kambhampati, S. (2026). Evaluating the false trust engendered by LLM explanations. arXiv:2605.10930.
  12. Bhambri, S., Biswas, U., & Kambhampati, S. (2025). Do cognitively interpretable reasoning traces improve LLM performance? arXiv:2508.16695.
  13. Engelbart, D. C. (1962). Augmenting Human Intellect: A Conceptual Framework. Stanford Research Institute.
  14. Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.