1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you do back-of-the-envelope estimation for an LLM system?
30-second answerSay your answer out loud first, then reveal.
Worked example: chat assistant, 1M daily active users
Assumptions: 5 conversations/user/day × 4 turns = 20 turns/user/day
Requests/day = 1M × 20 = 20M
Average QPS = 20M / 86,400 ≈ 230 → peak (3x) ≈ 700 QPS
Input tokens = 2,000 per turn (system prompt + history + retrieved docs)
Output tokens = 300 per turn
Daily input = 20M × 2,000 = 40B tokens
Daily output = 20M × 300 = 6B tokensCost (illustrative prices: $1 per 1M input tokens, $5 per 1M output tokens):
Input: 40,000M tokens × $1/M = $40,000/day
Output: 6,000M tokens × $5/M = $30,000/day
Total ≈ $70K/day ≈ $2.1M/month → $0.0035 per turnImmediate design insights from the numbers
- Input tokens dominate the volume, so prompt caching, shorter history and fewer retrieved chunks matter.
- Routing 60% of turns to a model 10x cheaper could cut cost by more than half.
- 700 peak QPS × 300 output tokens = 210K output tokens/s, so check provider rate limits (tokens per minute) or GPU capacity.
Storage: logging 20M turns × ~10 KB = 200 GB/day, so plan retention and sampling.
Common mistakes
- Forgetting that conversation history makes input tokens grow with every turn.
- Ignoring peak vs average, since rate limits apply at peak.
Related
This is what real progress feels like.