1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What are the key production metrics for an LLM service?
30-second answerSay your answer out loud first, then reveal.
| Category | Metrics |
|---|---|
| Latency | TTFT, TPOT/ITL, end-to-end, per-stage (retrieval, tools, LLM), p50/p95/p99 |
| Throughput | RPS, input/output tokens per second, concurrent requests |
| Reliability | Error rate by type (429, 5xx, timeout), fallback rate, retry rate |
| Cost | Tokens per request, cost per request, cost per active user, cost per resolved task |
| Quality | Judge pass rates (groundedness, policy), thumbs up/down, edit and regenerate rates |
| Safety | Guardrail block rates by category, injection detections, PII leak alerts |
| Agents | Steps per task, tool error rates, task success rate, loops hit |
| Self-hosted infra | GPU utilisation, GPU memory, KV-cache utilisation, queue depth, batch size, preemptions |
Interview tip: emphasise percentiles over averages, per-stage breakdowns, and cost per outcome rather than per call.
Related
Little by little, you're building something great.