1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Encoder-only, decoder-only, encoder-decoder: what's the difference, and when is each used?
30-second answerSay your answer out loud first, then reveal.
| Encoder-only | Decoder-only | Encoder-decoder | |
|---|---|---|---|
| Attention | Bidirectional | Causal (left-to-right) | Encoder bidirectional; decoder causal + cross-attention |
| Training objective | Masked language modelling | Next-token prediction | Span corruption / seq2seq |
| Examples | BERT, RoBERTa, DeBERTa | GPT, Llama, Mistral, Qwen, Claude, Gemini | T5, BART, original Transformer, Whisper |
| Best for | Classification, embeddings, extraction | General generation, chat, code, agents | Translation, summarisation, speech-to-text |
Why decoder-only won for general LLMs
- Simple, uniform objective that scales well on raw text.
- Any task can be phrased as "continue this text."
- Efficient generation with a KV cache (Q18).
Encoder models are still everywhere: embedding models for RAG and rerankers are often encoder-based, because bidirectional attention gives better representations for understanding and retrieval.
Related
You understood something today that you didn't yesterday.