How models work

Token

The unit of text a language model reads and writes: a word, a piece of a word, or a punctuation mark.

TEXTUnderstanding AI got easier.TOKENSillustrative splitUnder8179standing352AI11203got4467easier29871.915Token IDs are illustrative, not the real numbers of any specific model.4 words → 6 tokenscommon words stay whole≈ 4 charactersper token, rough English averageinput + outputyou are billed for both

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/token

In plain terms

Models don't see letters or words; they see tokens. Text is cut into pieces from a fixed vocabulary, and each piece becomes a number. Think Lego: a common word is one brick, a rare or long word is built from several.

Why it matters

Tokens are the currency of AI. You pay per token in and out, the amount of text a model can consider at once (the context window) is measured in tokens, and response time grows with the tokens generated. Budgeting an AI product is largely token arithmetic.

Example

In English one token is roughly four characters, about three-quarters of a word; 1,000 tokens come to around 750 words. Turkish, with its long suffixed words, usually needs more tokens for the same meaning, so the same document costs more to process.

Most often confused with

Token vs. Embedding

TokenA piece of text with an ID number
EmbeddingA vector of numbers that carries meaning

A token is a piece of text, represented by an ID number. An embedding is the list of numbers that carries the meaning of that piece, or of a whole text. Tokens answer “which piece”, embeddings answer “what does it mean”. Tokens drive the bill; embeddings power semantic search.

Under the hood

Most LLMs use subword tokenizers (BPE, WordPiece, SentencePiece) with vocabularies of roughly 50k–200k entries. Each token ID maps to an embedding vector; the model outputs a probability distribution over the next token. Counts differ between models, so the same text is not the same number of tokens everywhere. Output tokens usually cost several times more than input tokens, and cached input is cheaper. Images and audio are converted to tokens too.

Written by Mehmet Erkek · Last updated: