Infrastructure & economics

On-premise / Self-hosted

On-prem · self-hosting

Running AI models on infrastructure the organisation controls, in its own data centre or its own cloud account, so that prompts and data never pass through a model provider's shared service.

Shared APIpay per tokenDedicated capacityat the providerSelf-hosted in cloudyour cloud accountOn-premiseyour own data centreRuns the modelproviderprovideryouyouGPUs and serversprovider, sharedreserved for youyou rent themyou buy themData centreproviderprovidercloud provideryoursModel choiceincl. frontierincl. frontieropen-weightopen-weighthighlighted = your responsibilitymore control, more work →Many data-residency requirements are met by the two middle options; check those first.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/on-premise-self-hosted

In plain terms

The difference between eating at a restaurant and running your own kitchen. With a model provider's API you order, pay per dish and never see the kitchen. Self-hosting means the kitchen is yours: you buy or rent the equipment, hire the cooks and decide exactly what happens to every ingredient. On-premise is the strictest form, with the kitchen inside your own building.

Why it matters

Four reasons lead here: rules that keep data inside a country or a network, low latency next to factory or branch systems, independence from a supplier's changes, and unit cost. The economics turn on utilisation: a GPU server costs the same busy or idle, so self-hosting pays only at high, steady volume. Add the engineers who run it and slower access to the strongest models, most of which are offered only as a service. Before buying hardware, check the middle options: many residency requirements are met by a regional, dedicated or private cloud deployment.

Example

A state-owned bank may not send customer data outside its own network. It buys eight GPU servers for 2.4 million dollars, hires three engineers and runs an open-weight model for its 9,000 employees. Utilisation is 60% by day and 5% at night. In the same bank the marketing team, which works only with public content, keeps using a provider's API: for 200 users, paying per token is cheaper.

Most often confused with

On-premise / Self-hosted vs. Private cloud deployment

On-premise / Self-hostedYou operate the model on infrastructure you control
Private cloud deploymentA provider operates it for you, in an isolated or regional setup

Suppliers call several arrangements “private” that differ from self-hosting. Dedicated capacity is throughput reserved for one customer on the provider's machines. A private or regional deployment keeps processing inside a chosen region or network, with contractual limits on access. In both, the provider still operates the model. Both often satisfy a residency or isolation rule at far lower cost; ask exactly where data is processed and who can reach it.

Under the hood

Building blocks: GPU servers sized by the model's memory needs and the required throughput; an inference server such as vLLM, or a packaged option such as NVIDIA NIM, with Ollama and llama.cpp for small setups; an OpenAI-compatible endpoint, so that applications do not care where the model runs; and the surrounding pieces a provider normally supplies: autoscaling, monitoring, guardrails, access control and model updates. Sizing questions: concurrent users, tokens per second, latency target, context length, and whether quantisation is acceptable. Compare on total cost: hardware depreciation or reservation, power and cooling, staff, and realistic utilisation; idle capacity is the usual reason self-hosting loses the comparison. Middle options: provisioned throughput billed per reserved unit, private network endpoints on a cloud platform, sovereign-cloud regions, and vendor-managed systems that place a provider's models in the customer's data centre, including air-gapped variants such as Gemini on Google Distributed Cloud. Many organisations settle on a hybrid: sensitive or steady workloads self-hosted, everything else through an API, with one gateway in front of both.

Written by Mehmet Erkek · Last updated: