In plain terms
The difference between eating at a restaurant and running your own kitchen. With a model provider's API you order, pay per dish and never see the kitchen. Self-hosting means the kitchen is yours: you buy or rent the equipment, hire the cooks and decide exactly what happens to every ingredient. On-premise is the strictest form, with the kitchen inside your own building.
Why it matters
Four reasons lead here: rules that keep data inside a country or a network, low latency next to factory or branch systems, independence from a supplier's changes, and unit cost. The economics turn on utilisation: a GPU server costs the same busy or idle, so self-hosting pays only at high, steady volume. Add the engineers who run it and slower access to the strongest models, most of which are offered only as a service. Before buying hardware, check the middle options: many residency requirements are met by a regional, dedicated or private cloud deployment.
Example
A state-owned bank may not send customer data outside its own network. It buys eight GPU servers for 2.4 million dollars, hires three engineers and runs an open-weight model for its 9,000 employees. Utilisation is 60% by day and 5% at night. In the same bank the marketing team, which works only with public content, keeps using a provider's API: for 200 users, paying per token is cheaper.
Most often confused with
On-premise / Self-hosted vs. Private cloud deployment
Suppliers call several arrangements “private” that differ from self-hosting. Dedicated capacity is throughput reserved for one customer on the provider's machines. A private or regional deployment keeps processing inside a chosen region or network, with contractual limits on access. In both, the provider still operates the model. Both often satisfy a residency or isolation rule at far lower cost; ask exactly where data is processed and who can reach it.
Under the hood
Building blocks: GPU servers sized by the model's memory needs and the required throughput; an inference server such as vLLM, or a packaged option such as NVIDIA NIM, with Ollama and llama.cpp for small setups; an OpenAI-compatible endpoint, so that applications do not care where the model runs; and the surrounding pieces a provider normally supplies: autoscaling, monitoring, guardrails, access control and model updates. Sizing questions: concurrent users, tokens per second, latency target, context length, and whether quantisation is acceptable. Compare on total cost: hardware depreciation or reservation, power and cooling, staff, and realistic utilisation; idle capacity is the usual reason self-hosting loses the comparison. Middle options: provisioned throughput billed per reserved unit, private network endpoints on a cloud platform, sovereign-cloud regions, and vendor-managed systems that place a provider's models in the customer's data centre, including air-gapped variants such as Gemini on Google Distributed Cloud. Many organisations settle on a hybrid: sensitive or steady workloads self-hosted, everything else through an API, with one gateway in front of both.