Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q23IntermediateConcept

What does observability look like for agents, and why is it harder than for normal apps?

30-second answerSay your answer out loud first, then reveal.

What to capture per run

  • Trace ID, user/session ID, agent version (prompt + model + tool versions).
  • Spans for: each LLM call, each tool call, retrieval, guardrail checks, sub-agent runs.
  • Inputs and outputs, token counts, latency, cost, errors and retries.
  • Final outcome, plus user feedback (thumbs, edits, escalations).

Standards and tools: OpenTelemetry with the GenAI semantic conventions. Platforms include LangSmith, Langfuse, Arize Phoenix, Braintrust, and vendor dashboards.

What you do with traces

  1. Debugging: replay a failed run step by step.
  2. Error analysis: sample 50–100 traces, label failure modes (wrong tool, hallucinated argument, bad retrieval, gave up early), and fix the biggest bucket first. This is the most valuable habit in agent engineering.
  3. Monitoring: dashboards for success rate, steps per task, cost per task, tool error rate, latency at p50/p95, escalation rate.
  4. Turning traces into evals: a failed production trace becomes a test case.

Privacy: traces contain user data. Redact PII, apply retention limits, and restrict access.

Common mistakes

  • Logging only the final input and output.
  • Not versioning prompts, so you can't tell which change caused a regression.

Slow is fine. Stopping is the only problem.