1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is in-context learning, and what do we know about how it works?
30-second answerSay your answer out loud first, then reveal.
Observations
- ICL improves with model scale.
- The label format and the input distribution in the examples matter a lot. Some studies found models still perform reasonably even when example labels are randomised, which suggests the examples often help the model locate the task rather than teach it.
- Example order and selection affect results (recency bias, majority-label bias).
Mechanistic insight: induction heads (Olsson et al. 2022)
- Two-step circuit: a "previous token" head plus an induction head implement pattern completion: "[A] [B] ... [A] → predict [B]".
- They appear during training around the same time as a jump in ICL ability.
Theoretical views
- Task inference / Bayesian: the prompt narrows down which latent "task" or "concept" from pretraining applies.
- Implicit optimisation: for simple settings (e.g. linear regression), transformers can implement gradient-descent-like updates in their activations.
Practical implications
- Choose diverse, representative examples; retrieve examples similar to the query (dynamic few-shot).
- Keep the format consistent; balance labels.
- For many examples (many-shot ICL with long context), performance can keep improving, sometimes rivalling fine-tuning for some tasks.
Related
This is what real progress feels like.