In plain terms
A conventional dashboard tells you the lights are on: the service responds, quickly and without errors. An AI system can pass all of that and still give customers wrong answers, politely and at speed. Observability combines a flight recorder with a quality inspector: it keeps the record of every conversation, what the model was given, what it did, what that cost and how good the result was.
Why it matters
A language model fails quietly. No error message appears when an answer is wrong, when a prompt change doubles the bill or when a new document version misleads the assistant. Without observability a company hears about these from a customer complaint; with it, from its own data, usually days earlier. It also makes improvement possible, because real conversations become test cases. The price: tooling, storage, someone who looks at the data, and an archive of user conversations that must be protected like any other personal data.
Example
An insurer's claims assistant shows green on every classic dashboard: 99.9% uptime, answers in 1.8 seconds. The observability data tells another story. Since a prompt change on Monday, the share of answers judged to be grounded in the policy documents has fallen from 92% to 81%, and the cost per conversation has risen from 0.04 to 0.09 dollars. The team rolls the change back the same day.
Most often confused with
Observability vs. Monitoring
Monitoring watches known signals against thresholds and raises an alarm: the server is down, responses are slow. Observability keeps enough detail to answer questions nobody planned for, such as why this customer received that answer. Software needed it once systems became distributed. AI systems need it more, because the same input can produce different outputs and a wrong answer looks, technically, like a successful request.
Origin: The term comes from control theory, where Rudolf Kálmán introduced it around 1960; software engineering adopted it in the 2010s.
Under the hood
Signals collected per request: the trace (model calls, retrieval, tool calls and their timings), token counts and cost, latency including time to first token, model and prompt version, errors and guardrail blocks, and user feedback. Quality is scored on top: automated checks, LLM judges run on a sample of live traffic (online evals) and human review queues. Useful practices: tag every trace with the prompt and model version, so that a change can be compared before and after; alert on quality and cost as well as on errors; turn failed production cases into eval datasets. For a common data format, OpenTelemetry is defining semantic conventions for generative AI (attributes under gen_ai.*), which were still marked as in development in 2026. Tools include Langfuse, LangSmith, Arize Phoenix, Braintrust, W&B Weave and the LLM modules of general platforms such as Datadog. Decide retention, masking of personal data and access rights before collection starts.