1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Design an observability and tracing platform for LLM applications handling billions of spans per day.
30-second answerSay your answer out loud first, then reveal.

Design considerations
- Data volume: LLM payloads are large (prompts can be 10K+ tokens). Store metadata for 100% of traces, full payloads for a sample (plus all errors and flagged traces), and compress them.
- Sampling strategy: head sampling (decide at the start) vs tail sampling (decide after seeing the outcome: keep errors, slow traces, low feedback scores).
- Privacy: redaction before storage, tenant isolation, retention policies, role-based access to payloads.
- Cost attribution: compute cost from tokens × model price per span; aggregate by feature, team and user.
- Query patterns: "p95 latency by model last hour" (columnar aggregate), "show this user's failing traces" (index lookup), "find traces where the tool X errored" (filters).
- Online evaluation: asynchronous judges score sampled traces for faithfulness and relevance; results become metrics and alerts.
- Integrations: export to existing APM tools; link traces to deploys and versions for regression analysis.
Related
This is what real progress feels like.