Knowledge & retrieval

Knowledge Graph

A way of storing knowledge as a network of entities, such as people, companies and products, and the named relationships between them, so that connections can be followed and queried.

Ayşe KayaAcme Inc.Beta Ltd.X-200Part P7works atcustomer ofmakessuppliescontainsNODE = ENTITYEDGE = RELATIONSHIPQuestion: “Which supplier does the X-200 depend on?”Two hops in the graph: X-200 → Part P7 → Beta Ltd.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/knowledge-graph

In plain terms

A map of what is connected to what. Where a document holds knowledge as paragraphs, a knowledge graph holds it as small facts: “Ayşe works at Acme”, “Acme makes the X-200”, “the X-200 contains part P7”. A question is answered by walking from one point to the next along those links.

Why it matters

Some business questions are about connections: which customers are exposed to this supplier, who has worked with this client, which contracts refer to this clause. Text search handles these poorly; a graph handles them exactly and can show the path it took. The cost is real: someone has to define the entities, keep them clean and keep the graph current. Language models have made building one cheaper, though the quality control remains.

Example

A supplier, Beta Ltd., announces a six-week delay. A manufacturer asks which products and orders are affected. The graph follows three links: Beta Ltd. supplies part P7, part P7 is in the X-200 and two other products, and those products appear on 38 open orders. The answer arrives in seconds, with the chain of links that produced it.

Most often confused with

Knowledge Graph vs. Vector database

Knowledge GraphStores explicit, named relationships
Vector databaseStores closeness of meaning

A vector database knows that two passages are about similar things; it does not know how the things in them are related. A knowledge graph knows that one company supplies another, and nothing about how similar two texts sound. The first is cheap to build and fuzzy; the second is costly to build and exact. Many systems use both.

Origin: The term was popularised by Google, which launched its Knowledge Graph in 2012.

Under the hood

The basic unit is a triple: subject, relationship, object. Two main families: property graphs (nodes and edges with attributes, queried with languages such as Cypher) and RDF graphs (standardised triples, ontologies, queried with SPARQL). The hard work is modelling and data quality: a schema or ontology that says which entity and relationship types exist, and entity resolution, deciding that “Acme Inc.”, “ACME” and customer 40117 are the same thing. Graphs are built from structured systems and, increasingly, by language models extracting entities and relationships from text. With language models they are used in three ways: translating a question into a graph query, feeding retrieved subgraphs into the context, and GraphRAG.

Written by Mehmet Erkek · Last updated: