1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is quantization? Compare INT8, INT4, GPTQ, AWQ, and GGUF.
30-second answerSay your answer out loud first, then reveal.
Basic idea: map floating-point values to integers using a scale (and zero-point): w ≈ scale × int_value. Use a separate scale per channel or per group (e.g. per 128 weights) to keep error low.
Types
| Method | What | Notes |
|---|---|---|
| FP8 / INT8 weights (+ activations) | 8-bit | Usually near-lossless; FP8 is well supported on modern GPUs |
| GPTQ | 4-bit (also 3/8) post-training; quantizes layer by layer using second-order information to compensate errors | Popular for GPU inference |
| AWQ | Activation-aware: protects the small fraction of weights that matter most (by activation magnitude) via scaling | Good 4-bit quality, fast kernels |
| GGUF (k-quants like Q4_K_M) | File format + quantization schemes for llama.cpp | CPU/Mac/edge; mixed bit-widths per layer |
| bitsandbytes NF4 | 4-bit NormalFloat | Used in QLoRA training (Q22) |
| QAT (quantization-aware training) | Simulate quantization during training | Best quality at low bits; costs training |
Trade-offs
- 8-bit: minimal quality impact for most models.
- 4-bit: usually acceptable for chat; noticeable degradation on maths, code, long reasoning and non-English for some models. Evaluate on your tasks.
- Below 4 bits: significant degradation unless specially trained.
- Speed gains depend on kernel support. Some quantized formats are smaller but not faster.
Rule of thumb: a larger model at 4-bit often beats a smaller model at 16-bit with the same memory budget, but verify on your workload.
Related
Little by little, you're building something great.