1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Why do token counts matter in practice, and why do languages like Hindi often use more tokens?
30-second answerSay your answer out loud first, then reveal.
Where tokens show up
- Cost: price = input tokens × input rate + output tokens × output rate (output is usually several times more expensive).
- Latency: generation speed is measured in tokens/second, so a 1,000-token answer takes longer than a 100-token one.
- Context limits: system prompt + history + retrieved documents + output must all fit.
- Rate limits: often expressed in tokens per minute.
Why non-English text costs more
- BPE merges are learned from training data frequency. If Devanagari text is rare in that data, few Devanagari merges exist, so words split into many small pieces, sometimes down to individual bytes. (Each Devanagari character is 3 bytes in UTF-8.)
- Newer tokenizers with larger, more multilingual vocabularies have narrowed the gap considerably, but it still exists for many languages.
What to do about it
- Measure: count tokens for your actual multilingual traffic with each candidate model's tokenizer.
- Choose models with better tokenizers for your languages.
- Romanised Hindi (Hinglish) may tokenize differently from Devanagari. Measure both.
Interview signal: connecting tokenization to business impact (cost, latency, fairness across languages) shows practical maturity.
Related
Slow is fine. Stopping is the only problem.