Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q41HardConcept

Why do LLMs struggle with tasks like counting letters ("how many r's in strawberry"), arithmetic, or reversing facts?

30-second answerSay your answer out loud first, then reveal.

1. Character-level tasks

  • "strawberry" may tokenize as ["str", "aw", "berry"], so letters aren't directly visible as units.
  • The model must have memorised the spelling of each token. Workarounds: ask it to spell the word out letter by letter first, or use code.

2. Arithmetic

  • Numbers are tokenized into chunks inconsistently (e.g. "12345" → ["123", "45"]), which misaligns place values.
  • Multi-digit multiplication needs many sequential steps. Without writing intermediate steps (CoT), the fixed per-token computation isn't enough.
  • Mitigations: chain-of-thought, digit-level tokenization (some models split numbers into individual digits), or a calculator / code tool, which is the best option in production.

3. Reversal curse (Berglund et al. 2023)

  • Trained on "X's mother is Y", models often fail "Who is Y's son?" Learned associations are directional.
  • Relevant for knowledge retrieval: an entity may be "known" from only one direction.

4. Other characteristic weaknesses: precise counting over long lists; exact recall of rare facts; consistency across long outputs; planning tasks requiring search.

Interview framing. "These aren't random bugs. They follow from tokenization, autoregressive computation, and statistical learning. In production I route such sub-tasks to deterministic tools."

Little by little, you're building something great.