Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q44HardConcept

What is in-context learning, and what do we know about how it works?

30-second answerSay your answer out loud first, then reveal.

Observations

  • ICL improves with model scale.
  • The label format and the input distribution in the examples matter a lot. Some studies found models still perform reasonably even when example labels are randomised, which suggests the examples often help the model locate the task rather than teach it.
  • Example order and selection affect results (recency bias, majority-label bias).

Mechanistic insight: induction heads (Olsson et al. 2022)

  • Two-step circuit: a "previous token" head plus an induction head implement pattern completion: "[A] [B] ... [A] → predict [B]".
  • They appear during training around the same time as a jump in ICL ability.

Theoretical views

  1. Task inference / Bayesian: the prompt narrows down which latent "task" or "concept" from pretraining applies.
  2. Implicit optimisation: for simple settings (e.g. linear regression), transformers can implement gradient-descent-like updates in their activations.

Practical implications

  • Choose diverse, representative examples; retrieve examples similar to the query (dynamic few-shot).
  • Keep the format consistent; balance labels.
  • For many examples (many-shot ICL with long context), performance can keep improving, sometimes rivalling fine-tuning for some tasks.

This is what real progress feels like.