How models work

Latent Space

The internal, many-dimensional space in which a model represents what it has learned; positions and directions in it correspond to concepts and relationships.

manwomankingqueengender directionroyalty directionDIRECTIONS CARRY MEANINGking − man + woman≈ queenThe model encodes relationshipsit saw in the data as directions.Illustrative sketch: real latent spaces have thousands of dimensions.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/latent-space

In plain terms

A model never stores the sentence “Paris is the capital of France”. During training it arranges its internal numbers so that related ideas sit near each other and consistent relationships point in consistent directions. That hidden arrangement is the latent space. Everything the model reads is translated into it, and everything it writes is translated back out.

Why it matters

Latent space is the reason models generalise: they can handle a sentence they have never seen because it lands near ones they have. It is also why their behaviour is hard to audit. The knowledge is real but it is held as geometry, with no entry you can look up or delete. Interpretability research, and much of AI safety, is an effort to read that geometry.

Example

In an image model, moving a point slightly along one direction in latent space makes a face look older; along another, adds a smile. Nobody programmed an “age” control. The model organised its space so that age became a direction, and researchers found it afterwards.

Most often confused with

Latent Space vs. Embedding

Latent SpaceThe whole internal space of representations
EmbeddingOne point in a space, used as an output

An embedding is a specific vector: the coordinates of one text. Latent space is the space in which such vectors live, including the model's intermediate, hidden representations at every layer. An embedding model exposes one layer of its latent space for you to use.

Under the hood

“Latent” means not directly observed: the dimensions are learned, not designed, and most have no simple human label. Evidence of structure includes linear relationships between concept vectors (the well-known king − man + woman ≈ queen example), features found with sparse autoencoders, and steering experiments in which adding a vector to the activations changes behaviour. In diffusion and other generative models, generation consists of moving through latent space and decoding the result. Limits: single concepts are spread across many dimensions and single dimensions mix many concepts (superposition), which makes clean interpretation hard.

Written by Mehmet Erkek · Last updated: