Embeddings and the Latent Space of Language
Key Idea:
An LLM navigates a rich, implicit map of how humans have written about concepts and their relationships. That map is real and useful. But it is a map of language about the world, not a map of the world itself. There is no internal validity check — every output requires a reader who knows the territory.
The previous articles in this series, Do Large Language Models Reason? and What Do Large Language Models Do Then?, named "associative synthesis" as what large language models do:
search and recombine patterns from training data rather than reason.
That claim depends on a mechanism — what does the search actually run on? The answer is embeddings.
The Encoding Problem
A list of words has no structure.
"Doctor" and "physician" are different strings. A simple index assigns them different numbers and stops there. That index has no way to represent that the two words are nearly interchangeable, or that both are far from "carburetor" in some meaningful sense.
Early approaches — bag-of-words, simple frequency counts — captured word presence but not relationship. Shannon's 1948 information theory[1] showed how to measure information and how reliably symbols can be sent over a noisy channel. Shannon set meaning aside on purpose. He wrote that the "semantic aspects of communication are irrelevant to the engineering problem."
Meaning lives in the relationships between symbols, not inside any single one.
Stevan Harnad named the deeper problem in 1990: the symbol grounding problem.[2] A formal system can manipulate symbols according to rules without having any connection to what those symbols represent in the real world. Harnad asked how the meaning of such symbols could be intrinsic to the system, instead of "parasitic on the meanings in our heads." Embeddings leave that problem open. What they add is structure: each word gets a position defined by its relationships to other words.
Meaning from Context
The underlying idea comes from the distributional hypothesis, proposed by linguist Zellig Harris in 1954. Harris framed it as a correlation: the more two words differ in meaning, the more they differ in the contexts where they appear.[3] The usual short form is that words appearing in similar contexts tend to have similar meanings.
"Oncologist" and "cardiologist" both appear near "diagnosis," "patient," "hospital," and "treatment." That overlap in context is a signal about conceptual overlap. Harris claimed only a correlation. Context is a measurable proxy for meaning. The signal turns out to be rich enough to be useful at scale.
What an Embedding Is
An embedding maps each word to a point in high-dimensional space — typically hundreds to thousands of dimensions. The position encodes the contexts that word tends to appear in.
The positions are learned during training. The basic idea is simple: adjust the position of each word so that words appearing in similar contexts move closer together. Given enough text, the resulting geometry begins to reflect real relationships between concepts.
The landmark demonstration came from Mikolov and colleagues in 2013:[4] vector arithmetic on word embeddings produces meaningful analogies.
king − man + woman ≈ queen.
The arithmetic works because the geometry captures relationships between concepts. The same method gives:[5]
Paris − France + Italy ≈ Rome
This works because the embedding space contains a direction that corresponds roughly to the relationship between a country and its capital.
No one explicitly labeled that direction. It emerged from patterns of co-occurrence across billions of words.
Later models made embeddings contextual: the same word gets a different position depending on the words around it, so "bank" in a financial sentence lands far from "bank" beside a river. ELMo did this in 2018 using recurrent networks.[6] BERT did it with the transformer,[7] the architecture Vaswani and colleagues introduced in 2017 for machine translation.[8] Today's large language models are built on the transformer.
Latent Space: The Geometry Between Concepts
The high-dimensional space contains structure that was never explicitly labeled. It emerged from the statistical patterns in the training data and was encoded as a coordinate system. This is what we mean by latent space: relationships that are implicit in the data become explicit in the geometry of the model.
Three types of structure are particularly important.
Neighborhoods. Semantically related concepts tend to cluster together. Legal terms occupy one region; medical terms another. When a model processes a question about employment law, the relevant concepts are located in the same general neighborhood, influencing the patterns of text the model produces.
Directions. Some relationships appear as directions through the space. Gender, capital-to-country relationships, and differences in formal versus informal language can each correspond to directions in the geometry. These dimensions were not defined in advance. They emerged from patterns in the training data.
Paths. The geometry can also capture relationships between concepts that go beyond simple proximity. Human writing repeatedly moves through certain patterns of reasoning: evidence leads to a hypothesis; a claim is followed by a rebuttal and then a conclusion. These recurring patterns can become reflected in the structure of the space.
The important point is that the space is not random. It contains structure derived from the relationships present in the training data, and models use that structure when producing new text.
Geometry Is Not Knowledge
The latent space encodes statistical patterns in text about the world.
The world itself is outside the space, and there is no built-in representation of truth.
Three limits follow from this.
Truth has no reliable address. In a static word embedding such as word2vec, “The moon orbits Earth” and “Earth orbits the moon” are built from the same word vectors. Nothing in those vectors marks one sentence as true and the other as false.
Contextual models do place the two sentences differently, and researchers have found some structure related to truth inside them. Burns and colleagues recovered a direction in a model's internal activations that separates true statements from false ones, without using any labels.[9] On average that direction was about 4% more accurate than the model's own answers. Marks and Tegmark found that true and false statements separate along a roughly linear direction, and that the separation gets cleaner in larger models.[10]
These signals are partial. Both studies tested simple yes/no or true/false items with clear answers. Levinstein and Herrmann tested two probing methods, Burns's among them, and found that both failed to generalize in basic ways.[11] A supervised probe trained on plain factual statements scored worse than chance on negated versions of similar statements. Marks and Tegmark later reported probes that held up better across datasets, so the evidence on how far these signals generalize is mixed. The model also has no step where it checks a claim against this signal before writing it. The Burns result shows the gap directly: on average, the model's internal state held a slightly better answer than the one it gave.
This helps explain why hallucinations are a structural problem rather than simply a bug to patch. When the model generates from a region where the training data provides no strong path to a correct answer, it still has to produce something. It follows the statistical patterns available to it.
The result can be fluent and internally coherent while being factually wrong.
Whatever truth-related structure the geometry holds is weak, and nothing in generation requires the output to agree with it.
Grounding is still absent. Harnad's symbol grounding problem[2] has been deferred, not solved. The things that words refer to—the actual objects, events, and experiences in the world—are not present in the embedding space. “Invoice” occupies a position because of its relationships to terms such as “payment,” “due date,” and “net 30.” The model has never issued an invoice, waited for payment, or dealt with an overdue account.
Proximity is not entailment. “Promise” and “obligation” may be close neighbors in the latent space. But “A promised” does not logically entail “A is obligated.” That conclusion depends on conditions and exceptions that formal logic can represent but geometric proximity cannot. The embedding tells us that two concepts are related; it does not tell us what kind of relationship they have. Association and logical implication can both appear as distance in the same space.
Here Be Dragons
The latent space is a real and useful map of how humans have written about concepts and their relationships, compressed from a corpus far larger than any individual could read. That is where much of an LLM’s capability comes from: it can move across domains because the same geometry contains patterns from all of them.
But it is a map of language about the world, not a map of the world itself. A map of the London Underground can accurately show how stations connect while saying nothing about their elevation. In the same way, latent space can accurately capture patterns in human language while saying nothing about whether those statements are true.
Systems built on this substrate cannot verify their output by inspecting the space itself. There is no reliable internal truth register or formal validity check. Verification comes from outside the model:
someone who knows the subject must compare the output with what is actually known about the world.
Every generation is therefore a new traversal through a space with no ground truth. Whether the destination is correct depends on whether someone—or something outside the model—can check it against the territory.
References
- Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379–423.
- Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1–3), 335–346.
- Harris, Z. S. (1954). Distributional structure. Word, 10(2–3), 146–162.
- Mikolov, T., Yih, W., & Zweig, G. (2013). Linguistic regularities in continuous space word representations. Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), 746–751.
- Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv:1301.3781.
- Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., & Zettlemoyer, L. (2018). Deep contextualized word representations. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), 2227–2237.
- Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), 4171–4186.
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
- Burns, C., Ye, H., Klein, D., & Steinhardt, J. (2023). Discovering latent knowledge in language models without supervision. International Conference on Learning Representations (ICLR). arXiv:2212.03827.
- Marks, S., & Tegmark, M. (2024). The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. Conference on Language Modeling (COLM 2024). arXiv:2310.06824.
- Levinstein, B. A., & Herrmann, D. A. (2025). Still no lie detector for language models: Probing empirical and conceptual roadblocks. Philosophical Studies, 182(7), 1539–1565.