Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q21IntermediateConcept

What is quantization? Compare INT8, INT4, GPTQ, AWQ, and GGUF.

30-second answerSay your answer out loud first, then reveal.

Basic idea: map floating-point values to integers using a scale (and zero-point): w ≈ scale × int_value. Use a separate scale per channel or per group (e.g. per 128 weights) to keep error low.

Types

MethodWhatNotes
FP8 / INT8 weights (+ activations)8-bitUsually near-lossless; FP8 is well supported on modern GPUs
GPTQ4-bit (also 3/8) post-training; quantizes layer by layer using second-order information to compensate errorsPopular for GPU inference
AWQActivation-aware: protects the small fraction of weights that matter most (by activation magnitude) via scalingGood 4-bit quality, fast kernels
GGUF (k-quants like Q4_K_M)File format + quantization schemes for llama.cppCPU/Mac/edge; mixed bit-widths per layer
bitsandbytes NF44-bit NormalFloatUsed in QLoRA training (Q22)
QAT (quantization-aware training)Simulate quantization during trainingBest quality at low bits; costs training

Trade-offs

  • 8-bit: minimal quality impact for most models.
  • 4-bit: usually acceptable for chat; noticeable degradation on maths, code, long reasoning and non-English for some models. Evaluate on your tasks.
  • Below 4 bits: significant degradation unless specially trained.
  • Speed gains depend on kernel support. Some quantized formats are smaller but not faster.

Rule of thumb: a larger model at 4-bit often beats a smaller model at 16-bit with the same memory budget, but verify on your workload.

Little by little, you're building something great.