1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you operate long-running autonomous agents in production?
30-second answerSay your answer out loud first, then reveal.

Operational concerns
- Durability: a worker crash shouldn't lose 40 minutes of progress. Checkpoint after each step (Temporal, LangGraph checkpointers, or similar).
- Idempotency: tool calls carry idempotency keys; resumed runs don't duplicate side effects.
- Stuck-run detection: no progress for N minutes or repeated actions leads to an auto-pause and alert.
- Budgets: hard stops with partial results; per-tenant budget limits.
- Security: short-lived credentials per run; sandbox teardown after completion; egress allow-lists.
- Concurrency control: limits per tenant and per downstream system (don't overwhelm the CRM API).
- Observability: per-run trace trees; fleet metrics; sampling of runs for quality review.
- Versioning: in-flight runs keep their agent version; new runs pick up new versions via the rollout.
Related
Slow is fine. Stopping is the only problem.