In plain terms
The difference between a courier and the ordinary post. When a person is waiting for the answer, each request goes by courier: immediately and at full price. When nobody is waiting, thousands of requests can go in one sack. The provider works through them whenever its machines have spare capacity and hands back all the answers together, usually within a day, for about half the price.
Why it matters
A large share of AI work has no reader waiting: classifying archives, extracting fields from documents, enriching product records overnight, running evaluation suites. Paying real-time prices for it is waste. Batch discounts are commonly around fifty percent, and batch jobs usually draw on a separate quota, so they do not crowd out live traffic. The cost is time and certainty: results can take up to a day, a job that overruns its window can come back incomplete, and nothing interactive can be built on it. Decide per workload by asking who is waiting.
Example
A retailer wants better descriptions for 300,000 products. Sent as ordinary requests, the job would cost about 9,000 dollars and would compete with the shop's live assistant for the same rate limit. The team uploads the requests as batch jobs on Friday evening. By Saturday morning the result files are ready, the bill is about 4,500 dollars, and the assistant's customers noticed nothing.
Most often confused with
Batch Processing vs. Prompt caching
Both lower the bill without changing the answer, so they are often filed together as “discounts”. They reward different things. Batch processing rewards patience: any request qualifies if it can wait for hours. Prompt caching rewards repetition: it works in real time, and only on the part of a prompt that was sent before. With some providers the two can be combined on the same job.
Under the hood
Mechanics: requests are written to a file or a list, each with its own identifier, and submitted as one job; the job runs asynchronously and the results are collected by polling or through a callback. Results come back in any order and are matched by identifier. Typical terms at the large providers: a completion window of 24 hours, with many jobs finishing much sooner; limits on the number of requests and the file size per job; results kept for a limited period. Each request succeeds or fails on its own, so the pipeline needs retry logic for failed and expired items and must tolerate the same item being processed twice. Good fits: classification, extraction, embedding large collections, evaluation runs, synthetic data generation, back-office enrichment. Poor fits: chat, agents that need each result to choose the next step, and anything with a deadline tighter than the window. The term has a second meaning: an inference server grouping simultaneous requests on one GPU to raise throughput, which the customer never sees.