In plain terms
Someone who plays the violin learns the viola in weeks, because reading music, pitch and bowing carry over. Only what is specific to the new instrument is new. Models work the same way. One that has already learned to see edges, shapes and textures from millions of ordinary photos can learn to spot a cracked weld from a couple of thousand examples, because most of what it needs was learned before your project began.
Why it matters
Transfer learning is the reason AI projects no longer start with collecting a million examples. It moved the entry cost from a research budget to a pilot budget, and the whole foundation-model economy is built on it: one expensive general training, reused everywhere. The limits are real. Transfer works when the new task resembles what the model has already seen; on very unfamiliar data, such as unusual sensor signals, the gain shrinks. And whatever the original model learned badly, including its biases, travels with it.
Example
A steel fabricator wants to detect faulty welds from photographs and has 2,000 labelled images. A model trained from scratch on those reaches 71% accuracy. The team then takes a vision model already trained on millions of general images, keeps its early layers and retrains the last one on the same 2,000 photos. Accuracy reaches 96% after an afternoon on one graphics card.
Most often confused with
Transfer Learning vs. Fine-tuning
Transfer learning is the idea; fine-tuning is the most common way of carrying it out. Other ways exist: keeping the model frozen and training only a small new part on top, or using its embeddings as input to a simple classifier. Even a prompt relies on transferred knowledge. In everyday use the two words are often swapped; the distinction matters when deciding how much of a model to retrain.
Under the hood
Two classic recipes. Feature extraction: freeze the pre-trained network, replace its final layer (the head) and train only that. Fine-tuning: also update some or all of the earlier layers, usually at a low learning rate, sometimes unfreezing them gradually. Early layers hold general features and later layers task-specific ones, which is why the head is replaced first. In vision, networks pre-trained on ImageNet were the standard starting point through the 2010s. In language the same pattern took over in 2018 with models such as BERT and the first GPT, and it is the basis of today's foundation models. Related terms: domain adaptation (same task, different kind of data) and few-shot or zero-shot use, in which the transfer happens through the prompt alone. Risks: negative transfer, where an unsuitable source model hurts results; inherited bias; and licence terms of the source model that carry over to what you build.