50 questions that come up again and again in Agentic AI, AI Engineer and Forward Deployed Engineer interviews, with detailed answers, scenario walkthroughs and architecture diagrams.
Last updated: 06 Oct, 2026
The questions get harder as you go. The easy questions cover what an agent is and the vocabulary interviewers expect. The intermediate questions cover how agents break and how you fix them. The hard questions cover whiteboard system design, security and production reliability.
How to use these questions
- 30-second answer first. Every question opens with a short answer. Say it out loud before reading on. In a real interview, lead with this, then go deeper if the interviewer asks.
- Then the depth. The detailed answer is what a senior interviewer listens for: trade-offs, failure modes, and how you'd actually build it.
- Common mistakes. The answers that make interviewers lose interest. Knowing them helps as much as knowing the right answer.
- Follow-ups to expect. Interviewers rarely stop at one question. Prepare these and the conversation stays on your ground.
Question types
| Type | Questions | What it tests |
|---|
| Concept | 35 | How things work |
| Scenario | 9 | "This broke in production, what do you do?" |
| System design | 6 | Whiteboard rounds |
Tips for agentic AI interviews
- Measure before you fix. In a scenario question, the first sentence should be about traces, error analysis or evals, not a new prompt.
- Prompts guide, code enforces. Any rule that matters (refund limits, permissions, approvals) is checked at the tool boundary.
- Simplest thing that works. Show you know when not to use an agent, or multiple agents.
- In system design, clarify first. Ask about users, actions, risk, scale and SLAs before drawing boxes.
- Frameworks change; fundamentals don't. Name tools when it helps, but argue from principles: context, tools, evals, safety.
Easy: foundations
| # | Question | Type |
|---|
| Q16 | Compare planning strategies: ReAct, Plan-and-Execute, and Reflection. | Concept |
| Q17 | How do you manage the context window in a long-running agent? | Concept |
| Q18 | Your agent keeps calling the same tool with the same arguments in a loop. How do you debug and fix it? | Scenario |
| Q19 | Your agent has 40+ tools and often picks the wrong one. What do you do? | Scenario |
| Q20 | How should an agent handle tool failures and retries? | Concept |
| Q21 | Explain common multi-agent architectures. | Concept |
| Q22 | How do you evaluate an AI agent? | Concept |
| Q23 | What does observability look like for agents, and why is it harder than for normal apps? | Concept |
| Q24 | Why do agents need state persistence and checkpointing? | Concept |
| Q25 | What is prompt injection in the context of agents? Direct vs indirect? | Concept |
| Q26 | Your customer-support agent issued a refund that violated policy. What happened, and how do you prevent it? | Scenario |
| Q27 | How do you reduce an agent's latency? | Concept |
| Q28 | How do you reduce an agent's cost? | Concept |
| Q29 | Design long-term memory for an agent: what to store, when to write, how to retrieve, and how to forget. | Concept |
| Q30 | When is it safe to run tool calls in parallel? | Concept |
| Q31 | Your agent worked great in the demo, but fails about 30% of the time in production. How do you approach it? | Scenario |
| Q32 | What is idempotency, and why does it matter for agent tools? | Concept |
| Q33 | What are the pitfalls of using LLM-as-judge to evaluate agents, and how do you mitigate them? | Concept |
| Q34 | What are sub-agents, and why do coding and research agents use them? | Concept |
| Q35 | How would you choose between LangGraph, CrewAI, OpenAI Agents SDK, Google ADK, or writing your own loop? | Concept |
Hard: production and design
| # | Question | Type |
|---|
| Q36 | Design a customer-support agent for an e-commerce company with 1M monthly users. | System design |
| Q37 | Design a coding agent like Claude Code or Cursor's agent mode. | System design |
| Q38 | Design a "deep research" agent that produces a cited report on any topic. | System design |
| Q39 | How do you design authorization for an agent that acts on behalf of users across systems (email, CRM, drive)? | System design |
| Q40 | An email-assistant agent was tricked by a malicious email into forwarding confidential documents to an attacker. Explain the attack and design defences. | Scenario |
| Q41 | Your team built a 6-agent system. It costs 10x more than the old single agent, with no quality improvement. What do you do? | Scenario |
| Q42 | If each step of an agent is 95% reliable, what happens on a 20-step task? How do you build reliable agents anyway? | Concept |
| Q43 | You need to switch the agent's model to a newer or cheaper one. After switching, some behaviours regress. How do you manage model migrations? | Scenario |
| Q44 | Design an evaluation harness and CI/CD process for an agent. | System design |
| Q45 | How is designing tools for agents (the agent-computer interface) different from designing APIs for developers? | Concept |
| Q46 | How do you keep an agent effective on long-horizon tasks that run for hours (e.g. a large migration)? | Concept |
| Q47 | Your agent reports "Done! The record has been updated," but the record was never updated. How do you detect and prevent this? | Scenario |
| Q48 | Design a multi-tenant platform where customers can deploy their own agents (code execution, tools, memory). | System design |
| Q49 | When would you fine-tune a model for agentic behaviour instead of improving prompts and tools? | Concept |
| Q50 | A finance team asks you to build an "autonomous" agent that processes and approves vendor invoices. How do you scope, guard, and roll it out? | Scenario |
Every expert started right here.