Knowledge & retrieval

Knowledge Base

An organised, maintained collection of an organisation's information, such as policies, procedures, product details and answers, that people and AI systems consult as the authoritative source.

Scattered knowledgeeveryone keeps a copyThree versions of the same policyDocuments with no clear ownerWhich one should the assistant read?Knowledge baseone owned, current sourceOne valid document per topicA named owner and review datePeople and assistant read one sourceAn assistant cannot be better than the knowledge base it draws on.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/knowledge-base

In plain terms

The place where the company's official answers live. When an AI assistant is “connected to your knowledge”, this is what it reads: help articles, policy documents, product sheets, internal guides. The assistant does not know your business; it knows how to read, and the knowledge base is what you give it to read.

Why it matters

An assistant cannot be better than what it draws on. Most disappointing company assistants are content problems before they are model problems: three versions of the same policy, documents nobody owns, guidance that expired two years ago. Cleaning and owning the knowledge base usually raises answer quality more than switching to a stronger model.

Example

A telecom operator connects an assistant to 4,200 help articles. In the first test it gives three different answers about roaming fees, because three versions of the tariff article exist. The team retires 1,100 outdated or duplicate articles and names an owner for each topic. Accuracy on the test questions rises from 71% to 93% with the same model.

Most often confused with

Knowledge Base vs. Training data

Knowledge BaseRead at question time; can be corrected today
Training dataLearned during training; fixed once the model ships

Training data shaped the model before you met it and cannot be edited afterwards. A knowledge base sits outside the model and is read at the moment of the question, so a corrected article changes the next answer. Putting company knowledge into a knowledge base, and not into training, is what keeps it current and removable.

Under the hood

For AI use, a knowledge base is a content set plus a pipeline: connectors pull documents from their systems, a parser extracts text (tables, scans and slide decks are the usual trouble), the text is chunked, embedded and indexed, and changes are synchronised on a schedule. Useful metadata per document: owner, last review date, validity period, audience and access rights. Writing that retrieves well has one topic per article, descriptive headings and sections that make sense on their own. Operations: deduplication, archiving of expired content, permission-aware retrieval, and a feedback loop in which unanswered or badly answered questions become new or corrected articles.

Written by Mehmet Erkek · Last updated: