Infrastructure & economics

LLMOps

LLM operations

The practices and tools for running applications built on large language models in production: versioning prompts and models, evaluating changes, tracing behaviour, controlling cost and releasing safely.

1 · Versionprompt, model, settings2 · Evaluatetest set before release3 · Releasestaged, ready to roll back4 · Observetraces, cost, qualityTHE LLMOPS LOOPfor every changeCases that fail in production are added to the test set; that is how the loop improves.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/llmops

In plain terms

Getting a demo to work takes an afternoon. Keeping it working for thousands of users, month after month, while prompts are edited, models are replaced and costs creep up, is a different job. LLMOps is the discipline for that job: the routines that let a team change an AI application on Tuesday and know by Wednesday whether it got better or worse.

Why it matters

Most AI applications that fail in production fail for operational reasons: a prompt edit that quietly broke one case, a provider's model update, a bill that tripled, a complaint nobody can reproduce. LLMOps turns these from surprises into routine. It sets how fast a team can change things safely, and it produces the evidence an auditor or a regulator will ask for. It also costs effort: a test set to maintain, tooling to run and someone who owns it. A pilot can live without it; a product cannot.

Example

An insurer's claims assistant starts giving wrong deductible amounts. Because every request is traced, the team finds within an hour that the errors began with a prompt edit two days earlier. They roll back to prompt version 41, add the failing cases to the 400-case test set and repair the edit. The corrected version now has to pass that set before it reaches the first 5% of users.

Most often confused with

LLMOps vs. MLOps

LLMOpsOperates applications built on models someone else trained
MLOpsOperates the training and serving of your own models

MLOps grew up around teams that train their own models: data pipelines, feature stores, retraining, model registries. In most LLM applications nobody trains anything. The model is rented, and what changes is the prompt, the retrieved context and the tools. So the versioned artefact is the prompt and its configuration, quality is judged on open-ended text, and cost is counted per token. Teams that fine-tune or self-host need both.

Under the hood

Core practices. Versioning: prompts, model identifiers and settings, tool definitions and retrieval configuration are stored and released together, with model versions pinned where the provider allows it. Evaluation: a golden dataset run offline before each release, using code checks, LLM judges and human review, plus online evaluation on sampled production traffic. Tracing: every request recorded as a tree of model calls, retrievals and tool calls with inputs, outputs, tokens, latency and cost; OpenTelemetry's generative-AI conventions are emerging as the common format. Cost and performance: budgets, caching, routing and handling of rate limits. Release: staged rollout, A/B comparison and fast rollback. Feedback: user signals and failed cases flow back into the test set. Tools include LangSmith, Langfuse, MLflow and Arize Phoenix. For agents, where the label AgentOps is sometimes used, the unit of evaluation becomes the whole trajectory. Typical gaps: prompts edited directly in production, no test set, and model updates that arrive unannounced behind an alias.

Written by Mehmet Erkek · Last updated: