Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q5EasyConcept

When do you use synchronous, streaming, or asynchronous/batch processing for LLM features?

30-second answerSay your answer out loud first, then reveal.
ModeLatency needExampleNotes
StreamingInteractive, perceived speedChat, drafting emailsSSE/WebSockets; handle cancellation
Sync request/responseShort, needed before next stepIntent classification, JSON extraction in a workflowSet timeouts; small fast models
Async jobSeconds to hoursReport generation, video summarisation, agent tasksQueue + workers + status API + webhook/notification
BatchHours acceptableTag 10M products, nightly summaries, embeddings backfillBatch APIs, spot GPUs, max throughput

Async pattern

text
POST /jobs → 202 Accepted {job_id}
Worker pulls from queue → processes (with checkpoints) → stores result
Client polls GET /jobs/{id} or receives webhook / push notification

Design considerations

  • Timeouts: gateways and load balancers often cut connections at 30–60s, so long tasks must be async.
  • Idempotency: retries of job submission must not duplicate work.
  • Progress UX: stream intermediate steps ("Reading 12 documents...").
  • Priority queues: interactive traffic must not be starved by batch jobs sharing the same model capacity.

Every expert started right here.