1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you extend a model's context window (e.g. from 8K to 128K)? Explain RoPE scaling methods.
30-second answerSay your answer out loud first, then reveal.
Why naive extrapolation fails: RoPE dimensions rotate at different frequencies. Low-frequency dimensions encode long distances, and at unseen positions they produce out-of-distribution angles.
Methods
| Method | Idea | Notes |
|---|---|---|
| Position Interpolation (Chen et al. 2023) | Scale positions by trained_len / new_len, so position 16K looks like 8K | Simple; hurts short-range resolution slightly; needs fine-tuning |
| NTK-aware scaling | Increase the RoPE base (θ) so low frequencies stretch while high frequencies are preserved | Works partially even without fine-tuning |
| YaRN (Peng et al. 2023) | Different scaling per frequency band + attention temperature adjustment | Efficient; widely used |
| Larger base θ during training | Train with a big RoPE base from the start (e.g. 500K) | Common in newer models |
| Progressive long-context training | Train at 8K, then 32K, then 128K with long documents in mid-training | Data with genuine long-range dependencies matters |
Other requirements
- Long training data: books, code repositories, long documents. Synthetic long tasks help.
- Infrastructure: FlashAttention, context parallelism, KV cache memory planning.
- Evaluation beyond needle-in-a-haystack: multi-needle retrieval, aggregation, reasoning over long documents (RULER-style tasks), and your real use case. Models often support a long context nominally but use it unevenly.
Related
Every expert started right here.