Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

LLM Fundamentals Interview Questions logoLLM Fundamentals Interview Questions

50 questions that come up again and again in AI Engineer, ML Engineer and Forward Deployed Engineer interviews, from tokenization and attention to KV caches, fine-tuning, alignment and GPU serving maths.

Last updated: 06 Oct, 2026

The questions get harder as you go. The easy questions cover how LLMs work: tokens, attention, transformer blocks, sampling, training stages. The intermediate questions cover the mechanics interviewers probe: KV cache, quantization, LoRA, RLHF and DPO, MoE, serving optimisations. The hard questions cover capacity maths, distributed training, reasoning-model RL, and end-to-end system design.

How to use these questions

  • 30-second answer first. Every question opens with a short answer. Say it out loud before reading on. In a real interview, lead with this, then go deeper if the interviewer asks.
  • Then the depth. The detailed answer is what a senior interviewer listens for: trade-offs, failure modes, and how you'd actually build it.
  • Common mistakes. The answers that make interviewers lose interest. Knowing them helps as much as knowing the right answer.
  • Follow-ups to expect. Interviewers rarely stop at one question. Prepare these and the conversation stays on your ground.

Question types

TypeQuestionsWhat it tests
Concept41How things work
Scenario6"This broke in production, what do you do?"
System design3Whiteboard rounds

Tips for LLM fundamentals interviews

  • Do the maths out loud. Memory for weights, KV cache per token, and training bytes per parameter come up constantly. Practise them until they're quick.
  • Connect concepts to production. When you explain attention or the KV cache, add what it means for latency, cost and concurrency.
  • Know the 'why', not just the 'what'. Why divide by √d_k? Why GQA? Why pre-norm? Interviewers follow up on reasons.
  • Be precise with terms. Open-weight vs open source, prefill vs decode, perplexity vs accuracy. Precision signals depth.
  • Admit uncertainty about fast-moving details. Model names and benchmark scores change monthly; principles don't.

Easy: foundations

Intermediate: building and debugging

#QuestionType
Q16Write the scaled dot-product attention formula. Why divide by √d_k, and what is the causal mask?Concept
Q17What is multi-head attention? What are Multi-Query Attention (MQA) and Grouped-Query Attention (GQA), and why do modern LLMs use GQA?Concept
Q18What is the KV cache? Why is it needed, and how do you calculate its memory?Concept
Q19Why is attention O(n²)? What are FlashAttention and other long-context techniques?Concept
Q20Explain the prefill and decode phases of LLM inference. What are TTFT and TPOT, and which phase is compute-bound vs memory-bound?Concept
Q21What is quantization? Compare INT8, INT4, GPTQ, AWQ, and GGUF.Concept
Q22Explain LoRA and QLoRA. Why are they so popular?Concept
Q23How do you decide between prompting, RAG, PEFT (LoRA), and full fine-tuning?Concept
Q24Explain RLHF: the reward model and PPO stages.Concept
Q25What is DPO, and how does it compare to RLHF with PPO?Concept
Q26What are scaling laws? What did the Chinchilla paper show?Concept
Q27What is a Mixture of Experts (MoE) model?Concept
Q28How are LLMs evaluated? Explain perplexity, common benchmarks, and data contamination.Concept
Q29Your model's outputs are repetitive (looping phrases) or degenerate. How do you diagnose and fix it?Scenario
Q30After fine-tuning on your domain data, the model got better at your task but worse at general instructions and reasoning. What happened, and how do you fix it?Scenario
Q31How does constrained decoding guarantee structured outputs such as JSON?Concept
Q32What is speculative decoding, and why does it speed up inference without changing outputs?Concept
Q33What are continuous batching and PagedAttention (vLLM), and why do they matter for serving?Concept
Q34What is knowledge distillation for LLMs, and when would you use it?Concept
Q35Your application sends the same prompt with temperature 0 but gets slightly different outputs. Why, and how do you get reproducibility?Scenario

Hard: production and design

#QuestionType
Q36Estimate the GPUs needed to serve a 70B model to 32 concurrent users with 8K-token contexts. Walk through the math.Scenario
Q37Design an LLM inference platform serving 1,000 requests per second across several models.System design
Q38Outline how you'd pretrain an LLM from scratch. What are the main stages and decisions?System design
Q39Explain data parallelism, tensor parallelism, pipeline parallelism, and ZeRO/FSDP.Concept
Q40How do you extend a model's context window (e.g. from 8K to 128K)? Explain RoPE scaling methods.Concept
Q41Why do LLMs struggle with tasks like counting letters ("how many r's in strawberry"), arithmetic, or reversing facts?Concept
Q42You're asked to build a model for Hindi legal Q&A. Walk through your end-to-end plan, from base model choice to deployment.Scenario
Q43How are reasoning models trained with reinforcement learning? Explain RL with verifiable rewards, GRPO, and reward hacking.Concept
Q44What is in-context learning, and what do we know about how it works?Concept
Q45How are LLMs made safe? Discuss alignment techniques, jailbreaks, and red-teaming.Concept
Q46A model ranks top on public benchmarks but performs poorly on your company's task. Explain why, and how you'd choose a model properly.Scenario
Q47Explain the architecture choices of modern LLMs: pre-norm vs post-norm, RMSNorm, SwiGLU, RoPE, and GQA. Why did they win?Concept
Q48An enterprise asks: should we use LLM APIs or self-host open models? How do you analyse it?System design
Q49How do multimodal LLMs process images? Describe the typical architecture.Concept
Q50Walk through everything that happens from the moment a user sends a prompt to a hosted LLM until the response streams back.Concept
Back toInterview prep

Every expert started right here.