50 questions that come up again and again in RAG Engineer, AI Engineer and Forward Deployed Engineer interviews, from chunking and hybrid search to permission-aware retrieval and production system design.
Last updated: 06 Oct, 2026
The questions get harder as you go. The easy questions cover the RAG pipeline and the vocabulary interviewers expect: embeddings, chunking, vector search, reranking, evaluation. The intermediate questions cover how retrieval and generation fail and how you fix them. The hard questions cover enterprise-scale system design, security, and high-stakes reliability.
How to use these questions
- 30-second answer first. Every question opens with a short answer. Say it out loud before reading on. In a real interview, lead with this, then go deeper if the interviewer asks.
- Then the depth. The detailed answer is what a senior interviewer listens for: trade-offs, failure modes, and how you'd actually build it.
- Common mistakes. The answers that make interviewers lose interest. Knowing them helps as much as knowing the right answer.
- Follow-ups to expect. Interviewers rarely stop at one question. Prepare these and the conversation stays on your ground.
Question types
| Type | Questions | What it tests |
|---|
| Concept | 36 | How things work |
| Scenario | 10 | "This broke in production, what do you do?" |
| System design | 4 | Whiteboard rounds |
Tips for RAG interviews
- Separate retrieval from generation. When an answer is wrong, first say how you'd find out which stage failed: was the right chunk retrieved, and was it used?
- Parsing caps everything. Mention document parsing and chunk quality early. Most real RAG failures start at ingestion.
- Hybrid + rerank is the strong default. Be ready to explain BM25, dense retrieval, RRF and cross-encoders, and when each helps.
- Know when RAG is the wrong tool. Aggregations, counts and filters belong in SQL; behaviour belongs in fine-tuning or prompting.
- Permissions are not a prompt. In enterprise designs, access control happens in the retrieval engine, before the model sees anything.
Easy: foundations
| # | Question | Type |
|---|
| Q16 | Explain query transformation techniques: rewriting, multi-query, HyDE, step-back, and decomposition. | Concept |
| Q17 | How do you choose an embedding model for a RAG system? | Concept |
| Q18 | The answer exists in your documents, but the retriever never returns it. How do you debug? | Scenario |
| Q19 | Retrieval returns the right chunks, but the LLM still gives wrong or hallucinated answers. What do you do? | Scenario |
| Q20 | How do you fuse results in hybrid search? Explain Reciprocal Rank Fusion (RRF). | Concept |
| Q21 | Explain parent-child (small-to-big) retrieval and sentence-window retrieval. | Concept |
| Q22 | What is contextual retrieval, or contextual chunk enrichment? | Concept |
| Q23 | How do you handle tables and structured data in RAG? | Concept |
| Q24 | How do you handle follow-up questions in a conversational RAG chatbot? | Concept |
| Q25 | How do you build an evaluation dataset for RAG? | Concept |
| Q26 | Explain the core RAGAS metrics: faithfulness, answer relevancy, context precision, and context recall. | Concept |
| Q27 | Users complain the chatbot gives outdated answers even though documents were updated. How do you fix it? | Scenario |
| Q28 | Your users ask questions in Hindi, English and Hinglish, but most documents are in English. How do you design retrieval? | Scenario |
| Q29 | Break down the latency of a RAG request. How do you make it faster? | Concept |
| Q30 | What drives the cost of a RAG system, and how do you reduce it? | Concept |
| Q31 | Explain vector quantization and Matryoshka embeddings. What are the trade-offs? | Concept |
| Q32 | Retrieved documents contradict each other (an old vs new policy, two teams' docs). How should the system handle it? | Scenario |
| Q33 | What is late interaction (ColBERT), and how does it compare to bi-encoders and cross-encoders? | Concept |
| Q34 | How do you build multimodal RAG over documents with images, charts, and scanned pages? | Concept |
| Q35 | What are Corrective RAG, Self-RAG, and Adaptive RAG? | Concept |
Hard: production and design
| # | Question | Type |
|---|
| Q36 | Design an enterprise knowledge assistant over Confluence, Google Drive, SharePoint and Slack for 10,000 employees. | System design |
| Q37 | How do you implement permission-aware (ACL-aware) retrieval correctly? | Concept |
| Q38 | How would you scale RAG to 100 million+ documents (billions of chunks)? | System design |
| Q39 | What is GraphRAG? When would you use a knowledge graph with RAG? | Concept |
| Q40 | You're building RAG for a legal or medical setting where a wrong answer is very costly. How do you design for maximum trustworthiness? | Scenario |
| Q41 | An attacker uploads a document containing hidden instructions that get retrieved into the prompt. What are the risks, and how do you defend? | Scenario |
| Q42 | When and how would you fine-tune an embedding model or reranker for your domain? | Concept |
| Q43 | With 1M+ token context windows, is RAG dead? Argue both sides. | Concept |
| Q44 | Design a documentation assistant that answers coding questions using the latest library docs and source code (version-aware). | System design |
| Q45 | You switched to a new embedding model and a new chunking strategy, and production quality dropped. How should you have managed this migration, and what now? | Scenario |
| Q46 | How do you monitor a RAG system in production and improve it continuously? | Concept |
| Q47 | Users ask "How many of our vendor contracts have an auto-renewal clause?" and RAG gives wrong answers. Why, and what would you build? | Scenario |
| Q48 | Design a RAG system over real-time data, e.g. a news or market-intelligence assistant that must reflect events from the last few minutes. | System design |
| Q49 | How do you handle privacy and compliance in RAG: PII, the right to be forgotten, and multi-tenancy? | Concept |
| Q50 | A stakeholder says "the RAG bot's accuracy is bad." You have two weeks. How do you approach it end to end? | Scenario |
Every expert started right here.