In plain terms
Most applications send the same long opening on every request: the instructions, the tool list, perhaps a large document. Without caching, the model reads all of it from scratch each time and you pay full price each time. With caching, the provider keeps its work on that opening for a short while; the next request that starts the same way picks up from there and pays a fraction for the reused part.
Why it matters
For any product with long system prompts, large documents or multi-turn conversations, caching is among the largest cost and latency savings available, and it requires no change in output quality. On repeated content the discount is often between fifty and ninety percent, depending on the provider. It is frequently overlooked because it has to be designed in: what repeats must come first, and what changes must come last.
Example
A contract-analysis tool sends a 60,000-token contract with every question a lawyer asks about it. Over twenty questions, that is 1.2 million input tokens at full price. With the contract cached after the first question, the following nineteen read it at a fraction of the price (a tenth with some providers), and answers start arriving noticeably sooner.
Most often confused with
Prompt Caching vs. Memory
Caching does not make the model remember anything. It is a billing and speed optimisation on text you are sending anyway, and it expires within minutes or hours. Memory is information deliberately saved and brought back into later sessions. A cached prompt and an uncached one produce the same answer.
Under the hood
Caching matches on an exact prefix: one changed token invalidates everything after it. Order content from most to least stable: tool definitions, system prompt, documents, conversation history, new message. Keep timestamps, user names and other volatile values out of the prefix. Some providers cache automatically above a minimum length; others require explicit cache breakpoints. Pricing typically has three rates: a surcharge for writing to the cache, a deep discount for reading, and the normal rate for uncached tokens. Lifetimes run from about five minutes to an hour and are refreshed on use. Monitor the cache hit rate; a silent drop usually means something volatile crept into the prefix.