In plain terms
Supervised learning is studying with an answer key: here are ten thousand past loan applications and whether each was repaid, now predict the next one. Unsupervised learning is being handed a box of unsorted photographs and asked to put similar ones together; nobody says what the piles should be. The first learns to predict something you have named. The second shows you patterns you had not named.
Why it matters
The choice is made by your data and your question. If you can say exactly what should be predicted and you have past cases with known outcomes, supervised learning gives measurable accuracy, and it carries the cost of labelling. If you have plenty of data and no labels, unsupervised methods will find segments, outliers and themes, but their output is a suggestion that a person must interpret: there is no answer key to score it against. Many projects stall because a supervised result was promised on data that has no reliable labels.
Example
A telecom operator holds records on 2 million customers. To predict who will cancel, it trains a model on 18 months of history in which every customer is marked “stayed” or “left”: supervised learning. To design new tariffs it asks an algorithm to group customers by usage, with no categories given. Six segments come back, one of which nobody in marketing had noticed: unsupervised learning.
Most often confused with
Supervised and unsupervised vs. self-supervised learning
Self-supervised learning is how language models are pre-trained, and it sits between the two. No person labels anything, as in unsupervised learning. Yet the model is trained to predict a known answer, as in supervised learning: the next word of a real sentence is hidden and then used as the label. Text labels itself in this way, which is what made training on a large share of the internet possible.
Under the hood
Supervised tasks: classification (a category) and regression (a number). Typical algorithms: logistic regression, gradient-boosted trees and neural networks, judged on held-out labelled data with measures such as accuracy, precision and recall. Unsupervised tasks: clustering (k-means, hierarchical methods), dimensionality reduction (principal component analysis), anomaly detection and topic discovery; evaluation is indirect and relies on expert review or on usefulness downstream. In between: semi-supervised learning, which combines a few labels with much unlabelled data, and self-supervised learning, where the training signal is built from the data itself, as in next-token prediction, masked-word prediction and contrastive methods for images. A modern language model passes through several of these: self-supervised pre-training, supervised fine-tuning on demonstrations, and reinforcement learning from feedback. The third classic paradigm, reinforcement learning, learns from reward and has no fixed dataset of answers.