Observability for AI Agents in Production

· 7 min read

Traditional application observability answers “was this request slow, and where.” An AI agent in production raises a different question that standard APM was never built to answer: “why did the agent do that.” A span duration tells you a tool call took 400ms; it tells you nothing about whether the agent chose the right tool, whether it was one retry away from a completely different action, or whether the same input would produce the same output tomorrow. Agent observability has to be designed around decisions, not just requests.

What standard APM misses

A conventional trace captures a call graph: service A called service B called the database. Applied to an agent, this captures that the agent called a tool, and the tool took some time — which is true and almost useless on its own. What it does not capture is the reasoning step in between: which tools were available, which one the model selected and why (to the extent “why” is even legible from the model’s output), what the previous turn’s context was, and whether a retry silently changed the outcome. Debugging a bad agent action from a standard trace is like debugging a bug report that only says “the function returned the wrong value” with no arguments.

The four things worth logging on every agent turn

  1. The decision context. What tools were in scope for this call (see the Substrate Pattern’s tool substrate), what the relevant memory/context window contained, and what the model was actually asked to do — not just the final action taken.
  2. The action and its pre/post-conditions. Not just “tool X was called” but what the pre-condition check evaluated to, what the action changed, and what the post-condition check confirmed afterward. This is what makes an incident reconstructable instead of just visible.
  3. The identity and scope in force. Who this action was attributed to, and what scope was active at the time — critical once delegation is involved, since scope can narrow or (if something is misconfigured) widen partway through a call chain.
  4. The retry and fallback path. If the first model call failed, timed out, or produced an unusable response, what happened next: same provider retried, fallback provider used, or the turn abandoned. This is exactly the kind of state that multi-provider fallback logic needs to be observable, not just functional.

Traces vs. logs vs. metrics for agents

SignalGood forWeak for
Metrics (latency, error rate, token cost)Alerting on drift and cost; dashboardsExplaining any single bad decision
Logs (structured, per-turn)Reconstructing exactly what happened in one interactionSpotting trends across thousands of interactions
Traces (decision + action + pre/post-condition)Debugging a specific incident end to end, including delegationCheap to collect at very high volume without sampling

None of the three replaces the others. Metrics tell you something changed; logs and traces tell you what and why for a specific case worth investigating.

What to alert on

Alerting on agent systems fails in a specific way when copied wholesale from web-service alerting: p99 latency and 5xx rate are necessary but not sufficient. Add alerts for the agent-specific failure modes that a latency dashboard is structurally blind to:

  • A spike in post-condition check failures — the agent’s actions increasingly don’t match what the pre-condition predicted, which usually means the environment changed underneath it or a tool’s behavior drifted.
  • A spike in fallback-provider usage — the primary model provider is degrading in a way that isn’t yet a hard outage.
  • A change in the tool-call distribution — the same class of request is now routing to a different, unexpected tool, which often surfaces a prompt or model-version regression before users report anything wrong.
  • Any action taken outside its expected scope, even if the action itself succeeded — this should page immediately, not wait for a retrospective, because it means a guardrail failed to fire rather than that something was merely slow.

A worked example

An agent that files support tickets starts, after a routine model version bump, occasionally closing tickets instead of filing them — both are valid tool calls the agent has access to, so no error is thrown and no latency spike appears. Standard APM shows a fleet of fast, successful requests. What catches it: a post-condition check that expects “ticket status: open” after a file-ticket action starts failing at a low but nonzero rate, and a tool-call-distribution alert flags that the close-ticket tool is being invoked from a code path that historically never called it. Neither of those exists in a conventional latency-and-error-rate dashboard; both come directly from decision-level logging.

A starter checklist

  • Log the tool scope available at decision time, not just the tool ultimately called.
  • Record pre-condition and post-condition evaluations as structured fields, not free text, so they’re queryable after the fact.
  • Thread one request ID through every delegated call so a trace spans the whole call tree, not just the top-level agent.
  • Alert on post-condition failure rate and tool-call-distribution shift as first-class signals, not as an afterthought bolted onto latency dashboards.
  • Sample traces at 100% for any action with write access to a production system; sample everything else if volume demands it, but never sample away the write path.

Limitations

Full decision-level logging is not free: it increases storage cost, and logging the full decision context on every turn can itself leak sensitive data into observability systems that don’t have the same access controls as the production data store, which needs its own scoping and retention policy. There is also a ceiling on how much “why” is recoverable from a model call in the first place — logging the prompt and the output is not the same as logging the model’s actual reasoning, and treating a plausible-sounding explanation as ground truth is its own failure mode.

FAQ

Do I need a specialized AI observability vendor for this, or can I build it on existing infrastructure? The four signal types above can be built on existing structured logging and tracing infrastructure (OpenTelemetry spans with custom attributes, for instance) — the important part is deciding what to log, not which vendor’s SDK captures it.

How does this relate to the kill switch in Defence in Depth? Defence in Depth’s kill switch is the response; agent observability is what tells an operator it’s time to use it. A kill switch nobody knows to reach for is only theoretically a safety mechanism.

What’s the single highest-value thing to add first if I have none of this today? Pre/post-condition logging on any action with write access. It is the smallest change that turns “the agent did something wrong” into “the agent did something wrong, and here specifically is where the expectation and the reality diverged.”

Bottom line

Standard APM tells you an agent’s request was slow or failed; it does not tell you why an agent made the decision it made, which is the question that actually matters in production. Log decision context, actions with their pre/post-conditions, identity and scope, and the retry/fallback path on every turn, and alert on post-condition failures and tool-call-distribution shifts rather than latency alone.

Related Articles