Evaluation & quality

Tracing

Trace · LLM tracing

The step-by-step record of how an AI system handled one request or one agent run: every prompt, model call, retrieved document, tool call and output, with timings and cost.

TRACE · one conversation · 5 steps · 6.4 seconds · illustrativeThe customer's question6.4 s1 · Input check0.2 sis the request in scope?2 · Document search0.5 sreturned a policy dated 20233 · Model call1.4 sdecides to look up the order4 · Tool: order lookup2.8 sslowest step5 · Model call1.5 swrites the answer0 s6.4 sEach row is a span with its input, output and duration. A trace holds user data and must be protected.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/tracing

In plain terms

A trace is the itemised receipt for one answer. Behind a single reply a lot has happened: a search, perhaps three model calls, a lookup in the order system. The trace lists each of those steps in order, with what went in, what came out, how long it took and what it cost. When an answer is wrong or slow, the trace shows at which step the trouble began.

Why it matters

With agents, the final answer hides most of the work, and “the AI got it wrong” is no diagnosis. A trace turns it into a finding a team can act on: retrieval brought the wrong document, the model skipped an instruction, a tool returned an error. Traces are also the raw material for audits and test sets. The other side is what they contain. A trace holds the user's words, retrieved company documents and tool results in full, so it needs access control, masking and a retention period, like any system that stores personal data.

Example

A customer complains that a retailer's assistant quoted the wrong return period. The team opens the trace of that conversation: five steps, 6.4 seconds. Step two, the document search, returned a policy dated 2023, and the model quoted it faithfully. The trace also shows that the order lookup took 2.8 of the 6.4 seconds. One record yields two fixes: retire the old document and speed up the lookup.

Most often confused with

Tracing vs. Logging

TracingAll steps of one request in one linked, nested record
LoggingSeparate lines of events, each written on its own

A log lists events in the order they were written, mixed with those of every other request. A trace ties together all the steps of one request and keeps their hierarchy: this tool call happened inside that agent step, which belonged to this conversation. For a single model call a log line is enough. For an agent that takes twenty steps, only a trace shows how they connect.

Under the hood

A trace is a tree of spans. Each span is one unit of work (a model call, a retrieval, a tool execution, a guardrail check, a subagent) with a start time, duration, status, inputs, outputs and attributes such as model name, token counts and cost. Spans share a trace ID and point to their parent, which is how a multi-agent run across several services is joined into one record; sessions group the traces of one conversation. Instrumentation comes from SDK wrappers, framework integrations or an AI gateway, usually built on OpenTelemetry. Design choices: whether to store full prompts and outputs or only metadata, masking of personal data before storage, sampling at high volume, and retention. Uses beyond debugging: cost attribution per feature or customer, latency analysis, evaluating the path an agent took as well as its answer, and replaying a failed run against a new prompt or model.

Written by Mehmet Erkek · Last updated: