Foundations

Foundation Model

A large AI model trained on broad data that can be adapted to many different tasks, and so serves as the common base on which applications are built.

Foundation modeltrained once, on very broad dataWith promptsgive instructions and examplesWith RAGconnect your own documentsWith fine-tuningtrain on your own dataOne model, thousands of applications: the costly training is done once and adapting is cheap.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/foundation-model

In plain terms

Before foundation models, each AI task needed a model of its own: one for translation, one for sentiment, one for summarising. A foundation model is trained once, at great expense, on a very wide range of material, and then serves as the starting point for thousands of uses. You adapt it with instructions, with your own documents or with a little extra training.

Why it matters

It changed the economics of AI. A company no longer needs a data-science team and years of collected data to begin; it rents a model that already works and puts its effort into adapting it. The other side is dependency: a small number of providers build these models, and any flaw, bias or change in a foundation model is inherited by everything built on it.

Example

One foundation model serves three teams in a hospital group. Customer service uses it with a system prompt to answer questions about appointments. The quality team connects it to clinical guidelines through RAG. The billing department fine-tunes a version on 20,000 examples to assign billing codes. One base model, three adaptations, and no model trained from scratch.

Most often confused with

Foundation Model vs. Large language model (LLM)

Foundation ModelThe wider category: any adaptable base model
Large language model (LLM)The most common kind: a base model for language

Every major LLM is a foundation model, but the category is wider: it includes models for images, audio, video, protein structures, weather and robotics. “Foundation model” describes the role a model plays, as a base for many uses. “LLM” describes the material it works with.

Origin: The term was introduced in 2021 by researchers at Stanford's Institute for Human-Centered AI.

Under the hood

Defining features: self-supervised pre-training at scale (the model learns by predicting hidden or upcoming parts of unlabelled data), generality across tasks, and adaptation after training. Methods of adaptation, from lightest to heaviest: prompting, few-shot examples, retrieval (RAG), parameter-efficient fine-tuning such as LoRA, and full fine-tuning. Access: closed (through an API only) or open-weight. In regulation, the EU AI Act uses the related term general-purpose AI model and places obligations on providers concerning documentation, copyright and, for the most capable models, safety. Risks that follow from the structure: concentration among a few providers, shared weaknesses across all downstream applications, and limited visibility into the training data.

Written by Mehmet Erkek · Last updated: