Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q11EasyConcept

What is hallucination, and why do LLMs hallucinate?

30-second answerSay your answer out loud first, then reveal.

Main causes

  1. Objective mismatch: next-token prediction rewards plausibility. Nothing in pretraining checks truth.
  2. Lossy, compressed knowledge: rare facts (a small company's founding year) are poorly memorised, and the model fills gaps with plausible patterns.
  3. Knowledge cutoff: events after training aren't known, but the model may still answer.
  4. Training incentives: evaluations and feedback often reward a confident answer over abstaining, which encourages guessing.
  5. Prompt pressure: leading questions ("Why did X win the Nobel Prize?" when X didn't) and requests for specifics (citations, numbers).
  6. Decoding: high temperature, or long generations where an early error snowballs.

Types: factual errors, fabricated references, wrong reasoning steps, unfaithfulness to provided context (in RAG), invented code APIs.

Mitigations

  • Ground with retrieval (RAG) and tools (search, calculators, code execution).
  • Ask for citations and verify them; allow and encourage "I don't know".
  • Lower temperature for factual tasks.
  • Post-hoc verification (a second model, NLI checks, self-consistency across samples).
  • Use stronger models and models trained to abstain.

Common mistakes

  • Claiming RAG or a bigger model eliminates hallucination. They reduce it, so you still need verification for high-stakes outputs.

Little by little, you're building something great.