Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q28IntermediateConcept

How do you implement token streaming for a chat application end to end?

30-second answerSay your answer out loud first, then reveal.
Sequence chart: the client posts to the backend over SSE, the backend calls the LLM API with stream=true, and in a loop each delta chunk is buffered, moderation-checked and sent as a token event; then citations and tool status, done plus usage, persistence and a done event follow, and a client disconnect cancels the LLM request.

Details that show experience

  1. SSE vs WebSockets: SSE suits server→client token streams and works over HTTP/2 with automatic reconnect. WebSockets suit bidirectional real-time (voice, collaborative editing).
  2. Infrastructure: disable response buffering (e.g. Nginx proxy_buffering off), increase idle timeouts, and send heartbeats/keep-alive comments.
  3. Event protocol: typed events (token, tool_start, tool_result, citation, error, done) so the UI can show progress.
  4. Cancellation: propagate client disconnect to abort the upstream LLM request, which saves tokens.
  5. Moderation while streaming: buffer a small window or check sentence by sentence; if a violation appears mid-stream, stop and replace the message.
  6. Persistence: save partial content periodically for long generations, so a resume after reconnect is possible.
  7. Rendering: incremental Markdown rendering; handle partially streamed code blocks gracefully.

Slow is fine. Stopping is the only problem.