1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your model's outputs are repetitive (looping phrases) or degenerate. How do you diagnose and fix it?
30-second answerSay your answer out loud first, then reveal.
Diagnosis checklist
- Decoding:
• Greedy / temperature 0 → deterministic loops ("I think that I think that ..."). Try temperature 0.6–0.8 with top-p 0.9.
• Addrepetition_penalty(~1.05–1.2) or frequency/presence penalties. Too strong and the model avoids necessary words (names, code keywords). - Chat template mismatch (very common with open models): instruct models expect specific special tokens and role markers. Raw text prompts make the model behave like a base model and ramble. Use the tokenizer's
apply_chat_template. - Stop conditions: the EOS / stop token isn't configured, so generation runs until max_tokens and degrades.
- Fine-tuning issues:
• EOS token not appended to training targets → the model never learned to stop.
• Training data with repetitive patterns or duplicates.
• Overfitting (too many epochs, learning rate too high) → collapse to frequent phrases. - Context issues: the context is full of repeated text (e.g. a chat history echo), which the model continues.
- Quantization: aggressive low-bit quantization can increase degeneration. Compare with the full-precision model.
Fix and verify: build a small set of prompts that trigger the issue, adjust one variable at a time, and measure (e.g. the rate of repeated n-grams).
Related
This is what real progress feels like.