How models work

Tokenization

The step that converts text into the sequence of tokens a model works on, and converts the model's output tokens back into text.

Text“easier”Tokenizerlooks up, splitsToken IDs6606 · 2845Modelworks on numberson output the path runs in reverse: IDs become text againToken IDs are illustrative, not from any specific model.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/tokenization

In plain terms

Before a model can read your sentence, a tokenizer chops it into pieces from a fixed vocabulary and replaces each piece with its number. The model only ever sees the numbers. When it answers, the process runs backwards: numbers out, text reassembled.

Why it matters

Tokenization decides how many tokens your text becomes, and so what it costs and how much fits in the context window. It also explains some odd model behaviour: models stumble over spelling, counting letters and long numbers partly because they never see individual characters.

Example

The same paragraph is sent to two models from different providers. One counts 412 tokens, the other 455, because each uses its own tokenizer. A budget calculated with the first provider's counter is about ten percent short for the second.

Most often confused with

Tokenization vs. Embedding

TokenizationSplits text into pieces and numbers them
EmbeddingTurns each piece into a vector that carries meaning

Tokenization is a lookup with no understanding in it: text in, ID numbers out. Embedding is the next step, in which each ID is mapped to a vector the model has learned. The tokenizer decides where the pieces are; the embedding decides what they mean.

Under the hood

Modern tokenizers are subword schemes, mainly byte-pair encoding and its relatives. They are trained on a corpus by repeatedly merging the most frequent adjacent pairs until the vocabulary reaches a target size. Frequent words become single tokens; rare words, names and non-English text break into more pieces, which is why languages with rich morphology such as Turkish use more tokens per word. Spaces, capitalisation and punctuation affect the split. A tokenizer is tied to its model: text tokenized with one cannot be fed to another. Always count with the provider's own token counter.

Written by Mehmet Erkek · Last updated: