In plain terms
Nobody programs a language model with rules of grammar or facts about the world. It is shown an enormous amount of text with the next word hidden, asked to guess, told how wrong it was, and adjusted slightly. Repeat that trillions of times and the ability to write, translate and reason emerges as a by-product of getting better at the guess.
Why it matters
Training is where a model's capabilities, blind spots and biases are fixed, and it is the most capital-intensive step in AI: frontier training runs are estimated to cost hundreds of millions of dollars. Almost no company needs to train a model from scratch. The useful question is which trained model to build on, and whether to adapt it.
Example
A lab assembles a few trillion tokens of text, runs them through tens of thousands of GPUs for several months, and ends up with a single file of weights. From then on, that file can answer questions about tax law, write Python and draft emails, without having been taught any of those tasks explicitly.
Most often confused with
Training vs. Inference
Training happens once, before release, and produces the model. Inference happens every time someone uses it. A model does not learn from your conversations at inference time; anything it appears to remember comes from the context window or from a later training run.
Under the hood
The standard recipe is self-supervised learning with gradient descent: the model predicts the next token, a loss function measures the error, backpropagation computes how each parameter contributed, and an optimiser updates them. Stages usually follow one another: pre-training on broad data, then fine-tuning for instruction-following and safety, often using human feedback (RLHF) or model-written feedback. Key quantities are data volume and quality, parameter count and compute; scaling all three has been the main driver of progress. Training data ends at a cutoff date, and the model's knowledge with it.