Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q43HardSystem design

How do you operate long-running autonomous agents in production?

30-second answerSay your answer out loud first, then reveal.
A task request flows through a task queue to a durable orchestrator, which drives an agent worker, human approval interrupts and a traces store; the agent worker runs in a sandbox, reaches business systems through a tool policy gateway, and is capped by a budget enforcer, while traces feed a fleet dashboard.

Operational concerns

  1. Durability: a worker crash shouldn't lose 40 minutes of progress. Checkpoint after each step (Temporal, LangGraph checkpointers, or similar).
  2. Idempotency: tool calls carry idempotency keys; resumed runs don't duplicate side effects.
  3. Stuck-run detection: no progress for N minutes or repeated actions leads to an auto-pause and alert.
  4. Budgets: hard stops with partial results; per-tenant budget limits.
  5. Security: short-lived credentials per run; sandbox teardown after completion; egress allow-lists.
  6. Concurrency control: limits per tenant and per downstream system (don't overwhelm the CRM API).
  7. Observability: per-run trace trees; fleet metrics; sampling of runs for quality review.
  8. Versioning: in-flight runs keep their agent version; new runs pick up new versions via the rollout.

Slow is fine. Stopping is the only problem.