1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Why do LLMs struggle with tasks like counting letters ("how many r's in strawberry"), arithmetic, or reversing facts?
30-second answerSay your answer out loud first, then reveal.
1. Character-level tasks
- "strawberry" may tokenize as ["str", "aw", "berry"], so letters aren't directly visible as units.
- The model must have memorised the spelling of each token. Workarounds: ask it to spell the word out letter by letter first, or use code.
2. Arithmetic
- Numbers are tokenized into chunks inconsistently (e.g. "12345" → ["123", "45"]), which misaligns place values.
- Multi-digit multiplication needs many sequential steps. Without writing intermediate steps (CoT), the fixed per-token computation isn't enough.
- Mitigations: chain-of-thought, digit-level tokenization (some models split numbers into individual digits), or a calculator / code tool, which is the best option in production.
3. Reversal curse (Berglund et al. 2023)
- Trained on "X's mother is Y", models often fail "Who is Y's son?" Learned associations are directional.
- Relevant for knowledge retrieval: an entity may be "known" from only one direction.
4. Other characteristic weaknesses: precise counting over long lists; exact recall of rare facts; consistency across long outputs; planning tasks requiring search.
Interview framing. "These aren't random bugs. They follow from tokenization, autoregressive computation, and statistical learning. In production I route such sub-tasks to deterministic tools."
Related
Little by little, you're building something great.