In plain terms
Think of a printed textbook and a transparent overlay sheet. The book stays exactly as it was printed; your corrections and additions are written on the overlay. Lay it on top and the book reads your way; lift it off and the original is back. A LoRA adapter is that overlay for a model: a small add-on file holding the adjustments, while the model's billions of original numbers are never touched. Different overlays can sit on the same book.
Why it matters
LoRA turned fine-tuning from a project that needs a cluster into one that fits on a single graphics card, and it is why adapting open-weight models became routine. Training is cheaper, the result is a file of megabytes, and one copy of the base model can serve many teams, each with its own adapter. The limits: an adapter only works with the exact base model it was trained on, it cannot be applied to a closed model unless the provider offers that, and for large changes in behaviour full fine-tuning can still do better.
Example
A bank runs one open-weight model of 8 billion parameters, 16 GB in size, on its own servers. Legal, customer service and finance each want it to write in their own format. Three full fine-tunes would mean three 16 GB models and three sets of hardware. With LoRA each team trains an adapter of about 50 MB in an afternoon, and all three run on the single shared copy of the base model.
Most often confused with
LoRA vs. Full fine-tuning
Both teach a model from examples. Full fine-tuning rewrites the model and leaves a second multi-gigabyte copy to store, serve and maintain. LoRA leaves the model as it is and produces a small file that can be attached, removed or swapped. For most business adaptations the quality is close, which is why LoRA is the usual starting point; full fine-tuning is kept for deep changes with plenty of data.
Origin: Introduced in 2021 by Edward Hu and colleagues at Microsoft in the paper “LoRA: Low-Rank Adaptation of Large Language Models”.
Under the hood
A weight matrix W is kept frozen and the update is expressed as the product of two much smaller matrices, B and A, whose shared inner dimension is the rank r. For a 4,096 × 4,096 matrix, about 16.8 million numbers, rank 8 means training 65,536 numbers, under half a percent. Main settings: the rank (often 4 to 64), a scaling factor called alpha, and which layers receive adapters; the original paper used the attention matrices. After training, an adapter can be merged into the weights, which adds no delay at inference, or kept separate so that several can be swapped on one base model. QLoRA combines the method with a base model quantised to 4 bits, cutting memory further. The Hugging Face PEFT library is the common implementation. The same technique is widely used to teach image-generation models a style or a subject. LoRA belongs to a family known as parameter-efficient fine-tuning (PEFT).