Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q22IntermediateConcept

Explain LoRA and QLoRA. Why are they so popular?

30-second answerSay your answer out loud first, then reveal.
The input x goes through the frozen pretrained weight W and, in parallel, through the trainable matrices A then B (B initialised to zero, scaled by α/r); the two results are added to give h = Wx + BAx.

Parameter savings example: a 4096 × 4096 matrix has 16.8M parameters. LoRA with r = 8 has 2 × 4096 × 8 = 65,536 trainable parameters, about 0.4%.

Key hyperparameters

  • Rank r: capacity of the update (8–64 common; higher for harder or domain-shifting tasks).
  • α (alpha): scaling of the update (effective scale α/r).
  • Target modules: attention projections (q, k, v, o) and often the MLP layers. Targeting all linear layers usually works best.
  • B is initialised to zero, so training starts exactly at the base model.

Benefits

  1. Much less GPU memory (no optimiser state for the frozen weights).
  2. Tiny artifacts (MBs instead of GBs), so you can keep many task adapters and hot-swap or serve many adapters on one base model.
  3. After training, B·A can be merged into W, adding zero inference latency.

QLoRA (Dettmers et al. 2023): 4-bit NF4 base weights + double quantization + paged optimizers. The paper fine-tuned a 65B model on a single 48 GB GPU.

Limitations: for large domain shifts or teaching lots of new knowledge, full fine-tuning (or high-rank LoRA) can perform better. LoRA tends to learn less but also forget less.

You understood something today that you didn't yesterday.