In plain terms
Imagine a photograph slowly dissolving into television static. A diffusion model is trained to do the reverse: given a noisy picture, make it slightly cleaner. Having learned that from millions of images, it can start from pure static and clean its way, over a few dozen steps, to a picture that never existed. Your text prompt steers every step towards what you described.
Why it matters
It is the technology behind most AI image and video generation, and so behind a large shift in design, advertising, media and product visualisation, where a draft visual now takes seconds. It is also behind convincing fake images and video, which makes provenance and verification a concern for every communications and security team. The legal questions about training data and about ownership of the output are not settled.
Example
A furniture retailer needs each of its 400 sofas shown in five room settings. Photographing them would take months. A diffusion model, given the product photo and a description of each room, produces the 2,000 scenes in two days. A designer rejects about a fifth of them: wrong proportions, a fabric pattern that changed, a table with five legs.
Most often confused with
Diffusion Model vs. Large language model (LLM)
Both are generative AI, with different mechanics. A language model builds its output in sequence, piece after piece. A diffusion model starts with the whole canvas as noise and refines all of it together. The line is blurring: some current image generators are built into language models and produce images token by token, and many systems combine the two.
Origin: The modern form dates from a 2020 paper by Jonathan Ho and colleagues; public image generators built on it arrived in 2022.
Under the hood
Training: add Gaussian noise to real images in many small steps and train a network to predict the noise that was added. Generation: start from random noise and apply the network repeatedly to remove it. Most systems work in a compressed latent space and not on pixels (latent diffusion), which cuts cost sharply. Text conditioning comes from a text encoder, and the guidance strength sets how closely the result follows the prompt. The denoising network was at first a U-Net and is increasingly a transformer. Controls: seed, number of steps, guidance scale, negative prompts, image-to-image, inpainting and structural conditioning. Related approaches such as flow matching, and distilled models that need only a few steps, reduce generation time. Typical weaknesses: text inside images, hands, counts of objects and keeping a character consistent across images, all much improved in recent models.