Evaluation, Benchmarking & Observability · 7 min
Observability: Reading the Trace
An agent run is an execution graph — model calls, retrievals, and tool calls, each with its own numbers. Instrument it with a shared convention and debugging stops being archaeology.
A support ticket says the agent "gave a wrong refund amount." You open the logs. There are 400 lines of JSON, three of them are the actual model call, and none of them tell you whether the model hallucinated the number or a tool handed it back garbage. So you re-run the request four times trying to reproduce it. That is archaeology, and it is the default state of most agent systems in production.
The fix is to stop treating an agent run as a black box that emits a final string, and start treating it as what it actually is: an execution graph. Once you can see the graph, "why did it do that" becomes a question you answer by reading, not by guessing.
An agent run is a tree of spans
Distributed tracing already solved this shape for microservices, and it maps onto agents cleanly. A trace is one end-to-end request. A span is one unit of work inside it, with a start time, a duration, a status, and a bag of attributes. Spans nest: a parent span has children, and the children have children. Draw one agent turn and it looks like this:
trace: user_request "refund my last order" [2.9s]
└─ span: agent_run [2.9s]
├─ span: model_call (planning) [640ms]
│ gen_ai.request.model = claude-sonnet-4
│ gen_ai.usage.input_tokens = 1830
│ gen_ai.usage.output_tokens = 42
├─ span: retrieval "refund policy" [120ms]
│ results = 4, top_score = 0.81
│ doc_ids = [pol_17, pol_04, faq_09, pol_22]
├─ span: tool_call get_order [310ms]
│ arguments = {order_id: "A-5592"}
│ authorized = true
│ result.total = 48.00
├─ span: model_call (answer) [1.6s]
│ gen_ai.usage.output_tokens = 96
└─ output = "Refunded $48.00 to your card."Now the debugging session is a scan, not an excavation. The order total in the tool result is 48.00. The final answer says $48.00. The model did not invent the number, the retrieval pulled the right policy, and the tool was authorized. If the answer had instead read $84.00, you would know in five seconds that the model transposed a digit, because the tool span shows the truth it was handed. The trace turns "somewhere in this run something went wrong" into "this span, at this number."
Each node type carries its own numbers, and they are different numbers:
- Model calls care about tokens, latency, the model id, and the finish reason. Input tokens drive cost and tell you when your context is bloating. Output tokens plus latency tell you where the wall-clock time went.
- Retrievals care about how many results came back, their similarity scores, and the document IDs. A retrieval that returns four docs with a top score of
0.32found nothing useful, and that is often the real root cause of a bad answer three spans downstream. - Tool calls care about the arguments the model chose, whether the call was authorized, and what came back. This is where you catch a model calling
issue_refundwith the wrongorder_id, or reaching a tool it should never have been allowed to touch.
The nesting hands you one more signal for free: the critical path. In the trace above, the answer model_call is 1.6s of the 2.9s total, so if this turn feels slow, that span is where the budget went — not the 120ms retrieval you might have blamed on reflex. Read top to bottom, follow the longest child at each level, and the trace tells you which span to optimize next instead of leaving you to guess.
OpenTelemetry gives you a vendor-neutral vocabulary
You could invent your own attribute names for all of this. Don't. OpenTelemetry publishes semantic conventions for generative AI — an agreed set of span names, attribute keys, and metric names for exactly these operations. Adopt them and your traces are legible to any backend that speaks OpenTelemetry (Jaeger, Grafana Tempo, Honeycomb, Langfuse, and others) instead of being welded to one vendor's SDK.
The core conventions are worth knowing by name. A model span uses gen_ai.operation.name to say what kind of call it is (chat, embeddings, execute_tool), gen_ai.request.model and gen_ai.response.model for the model ids, and gen_ai.usage.input_tokens / gen_ai.usage.output_tokens for token counts. Span names follow a {operation} {model} pattern, so chat gpt-4 is a valid, self-describing span name. A tool execution gets the execute_tool operation and carries attributes like gen_ai.tool.name and gen_ai.tool.call.id, which let you tie the model's decision to call a tool to the span where that tool actually ran.
In practice you rarely hand-instrument all of this — an SDK does it for you. But the shape underneath is small:
from opentelemetry import trace
tracer = trace.get_tracer("agent")
with tracer.start_as_current_span("chat gpt-4") as span:
span.set_attribute("gen_ai.operation.name", "chat")
span.set_attribute("gen_ai.request.model", "gpt-4")
resp = client.chat.completions.create(model="gpt-4", messages=msgs)
span.set_attribute("gen_ai.usage.input_tokens", resp.usage.prompt_tokens)
span.set_attribute("gen_ai.usage.output_tokens", resp.usage.completion_tokens)Most teams never write that block by hand. Auto-instrumentation libraries — the opentelemetry-instrumentation wrappers around the OpenAI and Anthropic SDKs, or the tracing built into frameworks like LangChain and LlamaIndex — emit these spans for you. An OpenTelemetry Collector then sits between your app and the backend as the one place to batch, filter, and redact before anything leaves your network, which matters more than it sounds like once you get to the sensitive part.
The same convention also defines metrics, not just spans — gen_ai.client.token.usage and gen_ai.client.operation.duration, both histograms. Spans let you debug one bad run, metrics let you watch token usage per model creep up across a million runs and alarm before the bill does. The token count you set on a span is the same quantity you aggregate in the metric. That is the whole point of a shared convention: instrument once, read it two ways.
Fair warning: the GenAI conventions are still in-development, not yet stable, and the attribute names have already shifted (the token keys were prompt_tokens / completion_tokens before the input_tokens / output_tokens rename). Pin the version you build against and expect to update it. It still beats a bespoke schema you have to migrate alone.
The trace is a copy of everything sensitive
Here is the part teams discover too late. A trace that is genuinely useful for debugging contains the full prompt, the retrieved document chunks, the tool arguments, and the tool results. That means your telemetry backend now holds the user's actual message, the private policy docs your RAG system pulled, the order_id and email a tool was called with, and whatever that tool returned. The observability store becomes a second, shadow copy of your most sensitive data — one that engineers browse casually, that ships to a third-party SaaS, and that nobody wrote a retention policy for. One leaked trace can expose more than a stray log line ever did: a single span can hold an entire system prompt alongside a customer's name and card number sitting in the tool arguments.
The conventions take this seriously enough to bake it into the schema. Message content — the actual text of prompts and completions, carried on gen_ai.input.messages and gen_ai.output.messages — is marked opt-in: off by default, enabled only when you flip an explicit setting in the instrumentation. The structural attributes (token counts, model, latency, tool names) are safe to collect everywhere; the message content is the loaded gun. Treat that split as your policy boundary.
Concretely, telemetry needs its own privacy and access-control regime, separate from your application's:
- Redact before export, not just in the app. Run span content through PII scrubbing in the collector so a leaked API key or a customer's card number never lands in your trace store in the first place.
- Set a retention window on the telemetry store. Debug traces do not need to outlive your business records, and a shorter window shrinks the blast radius of a breach.
- Gate access. "Anyone with a Grafana login can read every prompt any user ever sent" is a data-exposure incident waiting to be named. Scope who can see content-bearing spans.
- Sample deliberately. Keep full content only for errored traces plus a small percentage of the rest, which cuts both cost and exposure.
The reframe that makes this stick: your traces are not logs, they are a dataset of real user interactions. If you would not drop that dataset into a public bucket, do not let it pile up untended in your observability tool either.
Instrument the graph, name it with a shared convention, and guard it like the sensitive data it is. Do those three things and the next 2 a.m. ticket is a five-minute read of one span, not a night of re-running requests hoping the bug shows itself.
Sources
- OpenTelemetry, "Semantic conventions for generative AI" (entry point), 2024–2025. https://opentelemetry.io/docs/specs/semconv/gen-ai/
- OpenTelemetry, GenAI semantic conventions repository, 2024–2025. https://github.com/open-telemetry/semantic-conventions-genai
- OpenTelemetry, "GenAI spans" (operation names, model/token attributes, opt-in message content), 2024–2025. https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md
- OpenTelemetry, "GenAI metrics" (token-usage and operation-duration histograms), 2024–2025. https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-metrics.md
- OpenTelemetry, "Traces" (signals and concepts), 2024. https://opentelemetry.io/docs/concepts/signals/traces/
- Sigelman et al., "Dapper, a Large-Scale Distributed Systems Tracing Infrastructure," Google, 2010. https://research.google/pubs/pub36356/