Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q40HardConcept

How do you extend a model's context window (e.g. from 8K to 128K)? Explain RoPE scaling methods.

30-second answerSay your answer out loud first, then reveal.

Why naive extrapolation fails: RoPE dimensions rotate at different frequencies. Low-frequency dimensions encode long distances, and at unseen positions they produce out-of-distribution angles.

Methods

MethodIdeaNotes
Position Interpolation (Chen et al. 2023)Scale positions by trained_len / new_len, so position 16K looks like 8KSimple; hurts short-range resolution slightly; needs fine-tuning
NTK-aware scalingIncrease the RoPE base (θ) so low frequencies stretch while high frequencies are preservedWorks partially even without fine-tuning
YaRN (Peng et al. 2023)Different scaling per frequency band + attention temperature adjustmentEfficient; widely used
Larger base θ during trainingTrain with a big RoPE base from the start (e.g. 500K)Common in newer models
Progressive long-context trainingTrain at 8K, then 32K, then 128K with long documents in mid-trainingData with genuine long-range dependencies matters

Other requirements

  • Long training data: books, code repositories, long documents. Synthetic long tasks help.
  • Infrastructure: FlashAttention, context parallelism, KV cache memory planning.
  • Evaluation beyond needle-in-a-haystack: multi-needle retrieval, aggregation, reasoning over long documents (RULER-style tasks), and your real use case. Models often support a long context nominally but use it unevenly.

Every expert started right here.