In plain terms
Getting a demo to work takes an afternoon. Keeping it working for thousands of users, month after month, while prompts are edited, models are replaced and costs creep up, is a different job. LLMOps is the discipline for that job: the routines that let a team change an AI application on Tuesday and know by Wednesday whether it got better or worse.
Why it matters
Most AI applications that fail in production fail for operational reasons: a prompt edit that quietly broke one case, a provider's model update, a bill that tripled, a complaint nobody can reproduce. LLMOps turns these from surprises into routine. It sets how fast a team can change things safely, and it produces the evidence an auditor or a regulator will ask for. It also costs effort: a test set to maintain, tooling to run and someone who owns it. A pilot can live without it; a product cannot.
Example
An insurer's claims assistant starts giving wrong deductible amounts. Because every request is traced, the team finds within an hour that the errors began with a prompt edit two days earlier. They roll back to prompt version 41, add the failing cases to the 400-case test set and repair the edit. The corrected version now has to pass that set before it reaches the first 5% of users.
Most often confused with
LLMOps vs. MLOps
MLOps grew up around teams that train their own models: data pipelines, feature stores, retraining, model registries. In most LLM applications nobody trains anything. The model is rented, and what changes is the prompt, the retrieved context and the tools. So the versioned artefact is the prompt and its configuration, quality is judged on open-ended text, and cost is counted per token. Teams that fine-tune or self-host need both.
Under the hood
Core practices. Versioning: prompts, model identifiers and settings, tool definitions and retrieval configuration are stored and released together, with model versions pinned where the provider allows it. Evaluation: a golden dataset run offline before each release, using code checks, LLM judges and human review, plus online evaluation on sampled production traffic. Tracing: every request recorded as a tree of model calls, retrievals and tool calls with inputs, outputs, tokens, latency and cost; OpenTelemetry's generative-AI conventions are emerging as the common format. Cost and performance: budgets, caching, routing and handling of rate limits. Release: staged rollout, A/B comparison and fast rollback. Feedback: user signals and failed cases flow back into the test set. Tools include LangSmith, Langfuse, MLflow and Arize Phoenix. For agents, where the label AgentOps is sometimes used, the unit of evaluation becomes the whole trajectory. Typical gaps: prompts edited directly in production, no test set, and model updates that arrive unannounced behind an alias.