In plain terms
When you read “the bank raised its rates”, you settle what “bank” and “its” mean by glancing at the other words. Attention gives a model the same ability. For every token it works out how much to draw on every other token, so the meaning of each word is built from its surroundings, however far apart the relevant words are.
Why it matters
Attention is the idea that made modern AI possible. It is what allows a model to connect a clause on page forty of a contract with a definition on page two. It also explains a practical limit: attention is spread across everything in the context, so as the context grows, the share given to any one detail shrinks.
Example
In “Ayşe handed the report to Elif because she had asked for it”, a model has to decide who “she” is. Attention links “she” strongly to “Elif” through the cue “asked for it”. Change the ending to “because she had finished it” and the strongest link moves to “Ayşe”.
Most often confused with
Attention vs. Context Window
The context window is the capacity: the number of tokens that can be present. Attention is the process applied to them: the relationships the model computes among those tokens. A large window guarantees that information can be present, and does not guarantee that the model attends to it.
Origin: Attention became the basis of language models with the 2017 paper “Attention Is All You Need”, which introduced the transformer.
Under the hood
In self-attention every token produces a query, a key and a value vector. A token's query is compared with all keys to give attention weights, and its new representation is the weighted sum of the values. Many attention heads run in parallel, each picking up different relations (syntax, coreference, position), and layers are stacked dozens of times. Cost grows roughly with the square of the sequence length, which is one reason long contexts are expensive. Attention weights can be inspected, though they are an unreliable explanation of why a model produced an output.