In plain terms
A doctor spends years in general education before specialising. Pre-training is that general education for a model: it reads a large slice of what people have written and learns how language works and how the world is described. The result is a capable generalist that later stages turn into a helpful assistant.
Why it matters
Pre-training is why one model can be used for thousands of tasks nobody trained it for, and why a handful of well-funded labs supply models to everyone else. For a buyer, it explains the economics: the expensive part has been paid for once and is shared across all customers, and your own investment goes into adapting and applying.
Example
A hospital group wants a model that summarises clinical notes. It does not gather billions of documents to train one. It takes a pre-trained model, which already reads and writes fluently, and supplies a few hundred of its own examples, through prompting or fine-tuning, to teach the house format.
Most often confused with
Pre-training vs. Fine-tuning
Pre-training builds general ability from scratch over months at enormous cost. Fine-tuning starts from that result and adjusts it for a domain, a style or a task in hours or days. The P in GPT stands for pre-trained.
Under the hood
The objective is self-supervised: predict the next token (in GPT-style models) or a masked token (in BERT-style models), so the data needs no human labels. Corpora mix web text, books, code and licensed sources, filtered and deduplicated. The output is a base model that continues text and does not yet follow instructions; chat behaviour comes from later instruction tuning and alignment. Because cost grows with data, parameters and compute together, only a few organisations train frontier base models, while distillation and open-weight releases spread the results more widely.