50 questions that come up again and again in AI Engineer, ML Engineer and Forward Deployed Engineer interviews, from tokenization and attention to KV caches, fine-tuning, alignment and GPU serving maths.
Last updated: 06 Oct, 2026
The questions get harder as you go. The easy questions cover how LLMs work: tokens, attention, transformer blocks, sampling, training stages. The intermediate questions cover the mechanics interviewers probe: KV cache, quantization, LoRA, RLHF and DPO, MoE, serving optimisations. The hard questions cover capacity maths, distributed training, reasoning-model RL, and end-to-end system design.
How to use these questions
- 30-second answer first. Every question opens with a short answer. Say it out loud before reading on. In a real interview, lead with this, then go deeper if the interviewer asks.
- Then the depth. The detailed answer is what a senior interviewer listens for: trade-offs, failure modes, and how you'd actually build it.
- Common mistakes. The answers that make interviewers lose interest. Knowing them helps as much as knowing the right answer.
- Follow-ups to expect. Interviewers rarely stop at one question. Prepare these and the conversation stays on your ground.
Question types
| Type | Questions | What it tests |
|---|
| Concept | 41 | How things work |
| Scenario | 6 | "This broke in production, what do you do?" |
| System design | 3 | Whiteboard rounds |
Tips for LLM fundamentals interviews
- Do the maths out loud. Memory for weights, KV cache per token, and training bytes per parameter come up constantly. Practise them until they're quick.
- Connect concepts to production. When you explain attention or the KV cache, add what it means for latency, cost and concurrency.
- Know the 'why', not just the 'what'. Why divide by √d_k? Why GQA? Why pre-norm? Interviewers follow up on reasons.
- Be precise with terms. Open-weight vs open source, prefill vs decode, perplexity vs accuracy. Precision signals depth.
- Admit uncertainty about fast-moving details. Model names and benchmark scores change monthly; principles don't.
Easy: foundations
| # | Question | Type |
|---|
| Q1 | What is a Large Language Model, and how does "predicting the next token" lead to useful behaviour? | Concept |
| Q2 | What is tokenization, and why do LLMs use subword tokens? | Concept |
| Q3 | Why do token counts matter in practice, and why do languages like Hindi often use more tokens? | Concept |
| Q4 | What are token embeddings and positional encodings? Why does a transformer need position information? | Concept |
| Q5 | Explain self-attention intuitively. What are queries, keys and values? | Concept |
| Q6 | What are the components of a transformer block in a modern LLM? | Concept |
| Q7 | Encoder-only, decoder-only, encoder-decoder: what's the difference, and when is each used? | Concept |
| Q8 | Explain temperature, top-k and top-p (nucleus) sampling. | Concept |
| Q9 | What is the context window, and what happens when you exceed it? | Concept |
| Q10 | Walk through the LLM training lifecycle: pretraining, SFT, and preference alignment. | Concept |
| Q11 | What is hallucination, and why do LLMs hallucinate? | Concept |
| Q12 | What are zero-shot, few-shot, and chain-of-thought prompting? | Concept |
| Q13 | What does "7B" or "70B parameters" mean, and how much memory does a model need? | Concept |
| Q14 | Open-weight vs closed (API) models: what are the trade-offs? | Concept |
| Q15 | What are reasoning models, and what is test-time compute? | Concept |
| # | Question | Type |
|---|
| Q16 | Write the scaled dot-product attention formula. Why divide by √d_k, and what is the causal mask? | Concept |
| Q17 | What is multi-head attention? What are Multi-Query Attention (MQA) and Grouped-Query Attention (GQA), and why do modern LLMs use GQA? | Concept |
| Q18 | What is the KV cache? Why is it needed, and how do you calculate its memory? | Concept |
| Q19 | Why is attention O(n²)? What are FlashAttention and other long-context techniques? | Concept |
| Q20 | Explain the prefill and decode phases of LLM inference. What are TTFT and TPOT, and which phase is compute-bound vs memory-bound? | Concept |
| Q21 | What is quantization? Compare INT8, INT4, GPTQ, AWQ, and GGUF. | Concept |
| Q22 | Explain LoRA and QLoRA. Why are they so popular? | Concept |
| Q23 | How do you decide between prompting, RAG, PEFT (LoRA), and full fine-tuning? | Concept |
| Q24 | Explain RLHF: the reward model and PPO stages. | Concept |
| Q25 | What is DPO, and how does it compare to RLHF with PPO? | Concept |
| Q26 | What are scaling laws? What did the Chinchilla paper show? | Concept |
| Q27 | What is a Mixture of Experts (MoE) model? | Concept |
| Q28 | How are LLMs evaluated? Explain perplexity, common benchmarks, and data contamination. | Concept |
| Q29 | Your model's outputs are repetitive (looping phrases) or degenerate. How do you diagnose and fix it? | Scenario |
| Q30 | After fine-tuning on your domain data, the model got better at your task but worse at general instructions and reasoning. What happened, and how do you fix it? | Scenario |
| Q31 | How does constrained decoding guarantee structured outputs such as JSON? | Concept |
| Q32 | What is speculative decoding, and why does it speed up inference without changing outputs? | Concept |
| Q33 | What are continuous batching and PagedAttention (vLLM), and why do they matter for serving? | Concept |
| Q34 | What is knowledge distillation for LLMs, and when would you use it? | Concept |
| Q35 | Your application sends the same prompt with temperature 0 but gets slightly different outputs. Why, and how do you get reproducibility? | Scenario |
Hard: production and design
| # | Question | Type |
|---|
| Q36 | Estimate the GPUs needed to serve a 70B model to 32 concurrent users with 8K-token contexts. Walk through the math. | Scenario |
| Q37 | Design an LLM inference platform serving 1,000 requests per second across several models. | System design |
| Q38 | Outline how you'd pretrain an LLM from scratch. What are the main stages and decisions? | System design |
| Q39 | Explain data parallelism, tensor parallelism, pipeline parallelism, and ZeRO/FSDP. | Concept |
| Q40 | How do you extend a model's context window (e.g. from 8K to 128K)? Explain RoPE scaling methods. | Concept |
| Q41 | Why do LLMs struggle with tasks like counting letters ("how many r's in strawberry"), arithmetic, or reversing facts? | Concept |
| Q42 | You're asked to build a model for Hindi legal Q&A. Walk through your end-to-end plan, from base model choice to deployment. | Scenario |
| Q43 | How are reasoning models trained with reinforcement learning? Explain RL with verifiable rewards, GRPO, and reward hacking. | Concept |
| Q44 | What is in-context learning, and what do we know about how it works? | Concept |
| Q45 | How are LLMs made safe? Discuss alignment techniques, jailbreaks, and red-teaming. | Concept |
| Q46 | A model ranks top on public benchmarks but performs poorly on your company's task. Explain why, and how you'd choose a model properly. | Scenario |
| Q47 | Explain the architecture choices of modern LLMs: pre-norm vs post-norm, RMSNorm, SwiGLU, RoPE, and GQA. Why did they win? | Concept |
| Q48 | An enterprise asks: should we use LLM APIs or self-host open models? How do you analyse it? | System design |
| Q49 | How do multimodal LLMs process images? Describe the typical architecture. | Concept |
| Q50 | Walk through everything that happens from the moment a user sends a prompt to a hosted LLM until the response streams back. | Concept |
Every expert started right here.