1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
When do you use synchronous, streaming, or asynchronous/batch processing for LLM features?
30-second answerSay your answer out loud first, then reveal.
| Mode | Latency need | Example | Notes |
|---|---|---|---|
| Streaming | Interactive, perceived speed | Chat, drafting emails | SSE/WebSockets; handle cancellation |
| Sync request/response | Short, needed before next step | Intent classification, JSON extraction in a workflow | Set timeouts; small fast models |
| Async job | Seconds to hours | Report generation, video summarisation, agent tasks | Queue + workers + status API + webhook/notification |
| Batch | Hours acceptable | Tag 10M products, nightly summaries, embeddings backfill | Batch APIs, spot GPUs, max throughput |
Async pattern
POST /jobs → 202 Accepted {job_id}
Worker pulls from queue → processes (with checkpoints) → stores result
Client polls GET /jobs/{id} or receives webhook / push notificationDesign considerations
- Timeouts: gateways and load balancers often cut connections at 30–60s, so long tasks must be async.
- Idempotency: retries of job submission must not duplicate work.
- Progress UX: stream intermediate steps ("Reading 12 documents...").
- Priority queues: interactive traffic must not be starved by batch jobs sharing the same model capacity.
Related
Every expert started right here.