In plain terms
Earlier language models read a sentence one word at a time and tended to lose the beginning by the time they reached the end. A transformer looks at the whole text at once and works out, for every word, which other words matter for understanding it. In “the bank raised its rates because it expected inflation”, it connects “it” to “the bank”. The T in GPT stands for transformer.
Why it matters
It is the invention on which the present wave of AI is built. Because it processes all positions in parallel, it could be trained on far more data than anything before it, using GPUs, and it kept improving as it was made larger. That is why progress since 2017 has come largely from bigger models, more data and more computing power. It also sets today's constraints: the context window and the cost of long inputs both follow from how attention works.
Example
Take the sentence “The trophy did not fit in the suitcase because it was too big.” What was too big? A reader knows it is the trophy. Change “big” to “small” and “it” becomes the suitcase. Older models often got such sentences wrong. A transformer's attention links “it” to the right noun in each version, because it weighs every word against all the others.
Most often confused with
Transformer vs. Recurrent neural network (RNN)
Recurrent networks, including LSTMs, were the standard for language before 2017. They processed text in order, which made training slow and connections across long distances weak. The transformer removed that step-by-step bottleneck, and that one change made training on text at the scale of the internet practical.
Origin: Introduced in 2017 by Ashish Vaswani and colleagues at Google in the paper “Attention Is All You Need”.
Under the hood
Input text becomes tokens, tokens become embedding vectors, and positional information is added so that word order is known. The vectors pass through a stack of identical layers, each with multi-head self-attention followed by a feed-forward network, with residual connections and normalisation. The final layer yields a probability for each possible next token. Variants: encoder-only models (BERT) for understanding tasks, decoder-only models (the GPT line and most current LLMs) for generation, and encoder-decoder models (the original design, and T5) for tasks such as translation. The cost of attention grows with the square of the sequence length, which is what limits context windows and motivates many efficiency techniques. The architecture now also underlies models for images, audio, video and protein structure.