1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What metrics would you track for an AI product: offline vs online?
30-second answerSay your answer out loud first, then reveal.
Offline and online metrics by layer
| Layer | Offline (pre-release) | Online (production) |
|---|---|---|
| Quality | Golden-set accuracy, rubric scores, faithfulness | Thumbs up rate, edit rate, sampled human / LLM-judge review |
| Task / business | Simulated task completion | Resolution rate, conversion, time saved, deflection |
| Safety | Red-team pass rate, PII leak tests | Incidents, blocked-content rate, reports |
| Performance | Benchmark latency | p50/p95 TTFT and total latency, error and timeout rates |
| Cost | Tokens per test case | Cost per request, user, outcome |
| Engagement | — | DAU using feature, retention, repeat usage |
North-star vs guardrail metrics: pick one primary outcome (e.g. "tickets resolved without human") and guardrails that must not degrade (CSAT, escalation correctness, cost per ticket).
Pitfalls
- Optimising thumbs-up rate can reward sycophancy or verbosity.
- Engagement can rise because the AI is confusing (more turns needed). Interpret it in context.
Related
PreviousHow do you design systems that tolerate the non-determinism and occasional failures of LLMs?
Every expert started right here.