1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What does observability look like for agents, and why is it harder than for normal apps?
30-second answerSay your answer out loud first, then reveal.
What to capture per run
- Trace ID, user/session ID, agent version (prompt + model + tool versions).
- Spans for: each LLM call, each tool call, retrieval, guardrail checks, sub-agent runs.
- Inputs and outputs, token counts, latency, cost, errors and retries.
- Final outcome, plus user feedback (thumbs, edits, escalations).
Standards and tools: OpenTelemetry with the GenAI semantic conventions. Platforms include LangSmith, Langfuse, Arize Phoenix, Braintrust, and vendor dashboards.
What you do with traces
- Debugging: replay a failed run step by step.
- Error analysis: sample 50–100 traces, label failure modes (wrong tool, hallucinated argument, bad retrieval, gave up early), and fix the biggest bucket first. This is the most valuable habit in agent engineering.
- Monitoring: dashboards for success rate, steps per task, cost per task, tool error rate, latency at p50/p95, escalation rate.
- Turning traces into evals: a failed production trace becomes a test case.
Privacy: traces contain user data. Redact PII, apply retention limits, and restrict access.
Common mistakes
- Logging only the final input and output.
- Not versioning prompts, so you can't tell which change caused a regression.
Related
Slow is fine. Stopping is the only problem.