1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Explain LoRA and QLoRA. Why are they so popular?
30-second answerSay your answer out loud first, then reveal.

Parameter savings example: a 4096 × 4096 matrix has 16.8M parameters. LoRA with r = 8 has 2 × 4096 × 8 = 65,536 trainable parameters, about 0.4%.
Key hyperparameters
- Rank r: capacity of the update (8–64 common; higher for harder or domain-shifting tasks).
- α (alpha): scaling of the update (effective scale α/r).
- Target modules: attention projections (q, k, v, o) and often the MLP layers. Targeting all linear layers usually works best.
- B is initialised to zero, so training starts exactly at the base model.
Benefits
- Much less GPU memory (no optimiser state for the frozen weights).
- Tiny artifacts (MBs instead of GBs), so you can keep many task adapters and hot-swap or serve many adapters on one base model.
- After training, B·A can be merged into W, adding zero inference latency.
QLoRA (Dettmers et al. 2023): 4-bit NF4 base weights + double quantization + paged optimizers. The paper fine-tuned a 65B model on a single 48 GB GPU.
Limitations: for large domain shifts or teaching lots of new knowledge, full fine-tuning (or high-rank LoRA) can perform better. LoRA tends to learn less but also forget less.
Related
You understood something today that you didn't yesterday.