Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q49HardSystem design

Design an observability and tracing platform for LLM applications handling billions of spans per day.

30-second answerSay your answer out loud first, then reveal.
Instrumented apps send spans to collectors and a streaming bus, which fans out to a search index, online evaluators, a columnar store and object storage; the stores feed dashboards, trace views and alerts, from which datasets are curated into eval sets.

Design considerations

  1. Data volume: LLM payloads are large (prompts can be 10K+ tokens). Store metadata for 100% of traces, full payloads for a sample (plus all errors and flagged traces), and compress them.
  2. Sampling strategy: head sampling (decide at the start) vs tail sampling (decide after seeing the outcome: keep errors, slow traces, low feedback scores).
  3. Privacy: redaction before storage, tenant isolation, retention policies, role-based access to payloads.
  4. Cost attribution: compute cost from tokens × model price per span; aggregate by feature, team and user.
  5. Query patterns: "p95 latency by model last hour" (columnar aggregate), "show this user's failing traces" (index lookup), "find traces where the tool X errored" (filters).
  6. Online evaluation: asynchronous judges score sampled traces for faithfulness and relevance; results become metrics and alerts.
  7. Integrations: export to existing APM tools; link traces to deploys and versions for regression analysis.

This is what real progress feels like.