In plain terms
You do not hand someone a 200-page manual to answer one question; you hand them the right page. Chunking is cutting the manual into pages in advance. Cut too coarsely and each piece covers five topics and matches nothing well. Cut too finely and a piece says “the limit is 30 days” without saying which limit.
Why it matters
Chunking is unglamorous and decides a large share of RAG quality. A poorly cut document cannot be retrieved well however good the model or the database. It is also where document reality bites: tables split in half, headings separated from their sections, footnotes detached from the claims they qualify.
Example
A bank's RAG pilot gives wrong answers about fee schedules. The cause: the PDF's fee tables were cut every 500 characters, separating each row of numbers from its column headings. Re-chunking by section and keeping each table whole, with its title, fixes most of the errors without touching the model.
Most often confused with
Chunking vs. Tokenization
Both cut text, at very different scales and for different purposes. Tokenization breaks text into word pieces so a model can process it at all. Chunking breaks documents into passages of a few hundred tokens so that search can find the relevant part. A chunk is made of many tokens.
Under the hood
Strategies: fixed size with overlap (simple, ignores structure); recursive splitting on paragraphs and sentences; structure-aware splitting by headings, sections, tables and code blocks; semantic splitting where topic shifts; and late or contextual approaches that attach a summary of the surrounding document to each chunk before embedding. Typical sizes run from 200 to 1,000 tokens with 10 to 20 percent overlap. Keep metadata with every chunk: source, section title, date, access rights. A parent-child setup retrieves on small chunks and hands the model the larger section around them. The best setting depends on the documents; test alternatives against real questions.