How models work

Attention

The mechanism that lets a model weigh, for each token, which other tokens in the text matter most for understanding it.

ATTENTION · illustrativeWho does “he” refer to?AnnagaveTomthebookbecauseheaskedstrong linkweak linksAttention builds the meaning of each token by looking at the other tokens in the sentence.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/attention

In plain terms

When you read “the bank raised its rates”, you settle what “bank” and “its” mean by glancing at the other words. Attention gives a model the same ability. For every token it works out how much to draw on every other token, so the meaning of each word is built from its surroundings, however far apart the relevant words are.

Why it matters

Attention is the idea that made modern AI possible. It is what allows a model to connect a clause on page forty of a contract with a definition on page two. It also explains a practical limit: attention is spread across everything in the context, so as the context grows, the share given to any one detail shrinks.

Example

In “Ayşe handed the report to Elif because she had asked for it”, a model has to decide who “she” is. Attention links “she” strongly to “Elif” through the cue “asked for it”. Change the ending to “because she had finished it” and the strongest link moves to “Ayşe”.

Most often confused with

Attention vs. Context Window

AttentionHow the model weighs what it reads
Context WindowHow much the model can read at once

The context window is the capacity: the number of tokens that can be present. Attention is the process applied to them: the relationships the model computes among those tokens. A large window guarantees that information can be present, and does not guarantee that the model attends to it.

Origin: Attention became the basis of language models with the 2017 paper “Attention Is All You Need”, which introduced the transformer.

Under the hood

In self-attention every token produces a query, a key and a value vector. A token's query is compared with all keys to give attention weights, and its new representation is the weighted sum of the values. Many attention heads run in parallel, each picking up different relations (syntax, coreference, position), and layers are stacked dozens of times. Cost grows roughly with the square of the sequence length, which is one reason long contexts are expensive. Attention weights can be inspected, though they are an unreliable explanation of why a model produced an output.

Written by Mehmet Erkek · Last updated: