Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q4EasyConcept

How do you do back-of-the-envelope estimation for an LLM system?

30-second answerSay your answer out loud first, then reveal.

Worked example: chat assistant, 1M daily active users

text
Assumptions: 5 conversations/user/day × 4 turns = 20 turns/user/day
Requests/day   = 1M × 20 = 20M
Average QPS    = 20M / 86,400 ≈ 230   → peak (3x) ≈ 700 QPS
Input tokens   = 2,000 per turn (system prompt + history + retrieved docs)
Output tokens  = 300 per turn
Daily input    = 20M × 2,000 = 40B tokens
Daily output   = 20M × 300   = 6B tokens

Cost (illustrative prices: $1 per 1M input tokens, $5 per 1M output tokens):

text
Input:  40,000M tokens × $1/M = $40,000/day
Output:  6,000M tokens × $5/M = $30,000/day
Total ≈ $70K/day ≈ $2.1M/month  → $0.0035 per turn

Immediate design insights from the numbers

  • Input tokens dominate the volume, so prompt caching, shorter history and fewer retrieved chunks matter.
  • Routing 60% of turns to a model 10x cheaper could cut cost by more than half.
  • 700 peak QPS × 300 output tokens = 210K output tokens/s, so check provider rate limits (tokens per minute) or GPU capacity.

Storage: logging 20M turns × ~10 KB = 200 GB/day, so plan retention and sampling.

Common mistakes

  • Forgetting that conversation history makes input tokens grow with every turn.
  • Ignoring peak vs average, since rate limits apply at peak.

This is what real progress feels like.