In plain terms
A model is, physically, a very large table of numbers. Each number says how strongly one artificial neuron should influence another. Training nudges these numbers, billions of times, until the network's outputs are useful. Nothing else is stored: no database of facts, no rule book. What the model knows is spread across the numbers.
Why it matters
Parameter count is the usual shorthand for model size, and it drives cost: more parameters need more memory and more computation for every token. It is a poor guide to quality on its own. Newer, smaller models regularly outperform older, larger ones, so compare results on your own tasks before you compare sizes.
Example
A model described as “7B” has seven billion parameters. Stored at 16-bit precision that is about 14 GB, which fits on a single high-end graphics card. A model fifty times larger needs a cluster, and its price per token reflects that.
Most often confused with
Parameters vs. Hyperparameters
Parameters are the numbers training discovers. Hyperparameters are the settings engineers choose for the training process itself: learning rate, number of layers, batch size. At inference time, settings such as temperature are also chosen by the user and are sometimes loosely called parameters, which adds to the confusion.
Under the hood
Parameters comprise weights (connection strengths) and biases, arranged in matrices; in a transformer they sit in the embedding table, the attention layers and the feed-forward layers. They are stored as floating-point numbers, and quantization reduces their precision (for example from 16 to 4 bits) to shrink memory use at a small cost in quality. “Open-weight” models publish the parameter files so anyone can run them; closed models expose only an API. Fine-tuning changes parameters; prompting never does. Interpretability research tries to work out what groups of parameters compute, with partial success so far.