Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q10EasyConcept

Walk through the LLM training lifecycle: pretraining, SFT, and preference alignment.

30-second answerSay your answer out loud first, then reveal.
The LLM training lifecycle: web, books and code feed pretraining to give a base model, SFT on instruction-response pairs gives an instruction model, preference or RL tuning gives a chat assistant, then evals and red-teaming lead to deployment.
StageDataComputeWhat it adds
PretrainingTrillions of tokens, mostly unlabelledHuge (thousands of GPUs for weeks or months)Language, knowledge, reasoning patterns
Mid-training / continued pretraining (optional)Domain data, long documents, code, mathsMediumDomain strength, longer context
SFT~10K–1M high-quality demonstrationsSmallFollowing instructions, formats, tool use
Preference / RLComparisons (A better than B), reward models, verifiable tasksSmall–mediumHelpfulness, safety, style, reasoning

Base vs instruct models: ask a base model "What is the capital of France?" and it may continue with more quiz questions. An instruct model answers.

Follow-ups to expect

  • How does RLHF work in detail? (Q24.) What is DPO? (Q25.)

Every expert started right here.