In plain terms
Earlier machine learning needed a specialist to decide what to measure: for a photo, perhaps edges, colours and textures. Deep learning skips that step. Given enough examples, the network works out for itself what to look for, building up from simple patterns in its first layers to whole concepts in its last. “Deep” describes the number of layers; it makes no claim about depth of understanding.
Why it matters
Almost everything called AI today is deep learning: language models, image generators, speech recognition, translation, the perception systems of self-driving cars. Its strength is unstructured data, which is most of what organisations hold. Its costs are an appetite for data and computing power, and opacity: a deep network cannot readily explain why it reached a result, which matters wherever decisions have to be justified.
Example
A factory inspects circuit boards for soldering defects. The older system used hand-tuned rules about brightness and shape and missed 9% of faults. A deep-learning model trained on 40,000 labelled photos misses 1.5%, and it picks up a kind of hairline crack for which no engineer had written a rule.
Most often confused with
Deep Learning vs. Classical machine learning
Both learn from examples. In classical machine learning a person first turns the raw material into meaningful measurements, and the algorithm learns from those. In deep learning the network learns the measurements too. That is why deep learning took over images, audio and language, where good features are hard to define by hand, and why it needs far more data and computing power.
Origin: The turning point came in 2012, when the deep network AlexNet won the ImageNet image-recognition competition by a wide margin.
Under the hood
A deep network stacks layers of simple units; training adjusts millions to billions of weights with backpropagation and gradient descent, on GPUs. Main architectures: convolutional networks (images), recurrent networks (sequences, now largely replaced), transformers (language and, increasingly, everything else) and diffusion models (generating images, audio and video). What made the breakthrough of the 2010s possible: large labelled datasets such as ImageNet, GPU computing and better training techniques. Practical themes: transfer learning (start from a pre-trained model and adapt it), self-supervised learning (learn from unlabelled data by predicting hidden parts of it), regularisation against overfitting, and interpretability research, which tries to open the black box.