Knowledge & retrieval

Vector Database

A database built to store embeddings and to find, very quickly, the stored items whose meaning is closest to a query.

Query: “how do I get my money back?” → [0.11, −0.52, …]VECTORTEXTSIMILARITY[0.12, −0.48, …]“Returns policy: within 30 days…”0.93[0.09, −0.55, …]“Refunds are paid within 5 working days.”0.89[0.71, 0.20, …]“Shipping fees and delivery times”0.31[−0.33, 0.64, …]“Corporate membership terms”0.12No shared keywords needed; closeness comes from meaning.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/vector-database

In plain terms

A normal database answers “find the rows where the city is Izmir”. A vector database answers “find the passages that mean something similar to this question”. It stores each piece of text as a point in a space of meaning and, given a new point, returns its nearest neighbours.

Why it matters

It is the storage layer behind most RAG and semantic-search systems. For decision-makers the practical point is that it is rarely a reason to buy a new product: the databases many companies already run, such as PostgreSQL, Elasticsearch and the major cloud databases, now offer vector search. A specialised product earns its place at very large scale or with demanding latency.

Example

A retailer embeds descriptions of 2 million products. A shopper types “something warm for a toddler to wear at the beach in the evening”. No product contains those words, yet the vector database returns hooded towelling ponchos and fleece-lined swim robes in forty milliseconds.

Most often confused with

Vector Database vs. Relational database

Vector DatabaseFinds what is similar in meaning
Relational databaseFinds what matches exactly

A relational database is exact: a row either meets the condition or does not. A vector database is approximate: it ranks everything by closeness and returns the nearest. Real applications need both, exact filters (this customer, this date range) and similarity, which is why vector search is increasingly a feature inside ordinary databases.

Under the hood

Each record holds a vector, the original text or a reference to it, and metadata for filtering. Search uses approximate nearest-neighbour indexes such as HNSW or IVF, which trade a little recall for large gains in speed; distance is cosine, dot product or Euclidean. Engineering concerns: filtered search (combining metadata conditions with similarity), index build time and memory, updates and deletes, multi-tenancy and access control, and re-embedding when the embedding model changes. Options include pgvector for PostgreSQL, search engines with vector support, and dedicated systems such as Pinecone, Qdrant, Weaviate and Milvus.

Written by Mehmet Erkek · Last updated: