Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q30IntermediateScenario

After fine-tuning on your domain data, the model got better at your task but worse at general instructions and reasoning. What happened, and how do you fix it?

30-second answerSay your answer out loud first, then reveal.

Why it happens: gradient updates optimise only the new data's loss. Nothing preserves performance on the original distribution.

Mitigations

  1. PEFT (LoRA): fewer trainable parameters, so less drift. "LoRA learns less and forgets less."
  2. Data mixing / replay: mix in 10–30% general instruction data (or self-generated data from the original model).
  3. Conservative hyperparameters: a lower learning rate (e.g. 1e-5 to 2e-5 for full fine-tuning; LoRA tolerates higher), 1–3 epochs, early stopping.
  4. Keep the format consistent: use the base model's chat template and system-prompt conventions.
  5. Regularisation: a KL penalty against the original model's outputs, or weight-space regularisers.
  6. Model merging: interpolate fine-tuned and original weights (or merge LoRA adapters at a reduced scale) to trade domain gains for general ability.
  7. Evaluate both: track a domain eval and a general eval (instruction following, reasoning, safety) for every checkpoint, and choose the best balance, not the lowest training loss.

Also check: safety behaviours can degrade after fine-tuning even on benign data. Re-run safety and red-team evals.

Every expert started right here.