Foundations

Large Language Model

LLM

An AI model trained on very large amounts of text to predict the next token, which enables it to write, summarise, translate, analyse, answer questions and produce code.

LLMone model, one job: predict the next tokenWrite and summarisedrafts, reports, translationAnalyseclassify, extract, compareCode and use toolsthe engine of agents“Large”: billions of parameters and training data at the scale of the internet.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/large-language-model

In plain terms

A model that has read a large part of what people have written and learned one skill from it: given some text, continue it sensibly. That single skill turns out to be enough to answer questions, draft documents, translate, write software and reason through problems, because doing each of those well comes down to producing the right next words.

Why it matters

It is the engine inside almost every AI product now on sale: assistants, copilots, agents, AI search. Understanding three of its properties prevents most bad decisions. It predicts plausible text and does not look facts up, so it can be confidently wrong. It knows nothing about your organisation unless you supply that information. And it remembers nothing between conversations unless a system built around it provides a memory.

Example

A lawyer pastes a 40-page lease into an assistant and asks for the break clauses, the rent review dates and anything unusual. The model returns a structured summary in twenty seconds and flags an uncapped service charge. It was never trained to review leases as a task; it reads the document in its context and applies general language ability to it.

Most often confused with

LLM vs. Chatbot

LLMThe model: an engine that predicts text
ChatbotA product: an interface that people talk to

ChatGPT, Claude and Gemini are products. Inside each runs a large language model, together with a system prompt, tools, memory and safety layers. The same model can power a chatbot, a coding agent or an invisible step in a back-office process. When comparing offers, separate the model from the product built around it.

Under the hood

An LLM is a transformer neural network with billions of parameters. Text is split into tokens; the model outputs a probability for every possible next token; one is sampled and appended; the process repeats. Life cycle: pre-training on a very large body of text by next-token prediction, then further training for following instructions and for safety, often using human feedback. Key characteristics: context window (how much it can read at once), knowledge cutoff, the kinds of input it accepts, speed and price per token. Abilities beyond plain text generation come from the surrounding system: tool use, retrieval, structured output, extended thinking. Limits: hallucination, sensitivity to phrasing, uneven reliability on long multi-step tasks, and no learning from use unless it is retrained.

Written by Mehmet Erkek · Last updated: