How models work

Embedding

A list of numbers that represents the meaning of a piece of text, so that texts with similar meaning have similar numbers.

catdogbirdapplepearbananainvoicepaymenttransferFROM TEXT TO VECTOR“cat”[0.21, −0.63, 0.08, …]hundreds or thousands of numbersClose points =close meaningSemantic search and RAG work by finding the points closest to the question.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/embedding

In plain terms

Imagine placing every sentence on a huge map where distance stands for difference in meaning. “How do I reset my password?” and “I forgot my login” land close together though they share no words; “quarterly tax filing” lands far away. An embedding is the map coordinates of a text. Computers cannot compare meanings, but they can compare coordinates.

Why it matters

Embeddings are what make search by meaning possible, and they sit behind most practical AI systems: finding the relevant passages for RAG, matching support tickets to known issues, recommending similar items, detecting duplicates. If a company's AI assistant “knows” its documents, embeddings are almost always how it finds them.

Example

A support team embeds 20,000 past tickets. When a new ticket arrives saying “the app logs me out every few minutes”, the system finds the five nearest tickets in under a second, including one titled “session expires too quickly”, and shows the fix that resolved it.

Most often confused with

Embedding vs. Token

EmbeddingA vector that captures meaning
TokenA piece of text with an ID number

A token is a unit of text identified by a number that carries no meaning. An embedding is a vector of hundreds or thousands of numbers whose position encodes meaning. Tokens are counted for billing; embeddings are compared for similarity.

Under the hood

Embedding models map text to a fixed-length vector, commonly 384 to 3,072 dimensions. Similarity is measured with cosine similarity or dot product, and vector databases use approximate nearest-neighbour indexes to search millions of vectors quickly. Practical points: documents are split into chunks before embedding; queries and documents must be embedded with the same model; changing the model means re-embedding everything; quality varies by language and domain, so test on your own data. Embeddings also exist for images and audio, and multimodal models place different media in one shared space. Inside an LLM, the first layer is itself an embedding table that maps token IDs to vectors.

Written by Mehmet Erkek · Last updated: