Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q35IntermediateScenario

Your application sends the same prompt with temperature 0 but gets slightly different outputs. Why, and how do you get reproducibility?

30-second answerSay your answer out loud first, then reveal.

Sources of non-determinism

  1. Floating-point non-associativity: (a + b) + c ≠ a + (b + c) in finite precision. Parallel reductions in matrix multiplies and attention can change order.
  2. Batch-dependent kernels: the same request batched with different other requests can use different kernel strategies or tile splits, which changes numerics. This is a major source in served LLMs.
  3. MoE routing: capacity limits and token dropping can depend on what else is in the batch.
  4. Near ties: two tokens with almost equal logits, so a tiny difference flips the choice.
  5. Different hardware / software versions across a provider's fleet.
  6. Model updates behind an alias such as "latest".

Getting more reproducibility

  • Pin exact model versions; use seed parameters where offered (usually "best effort").
  • Self-hosted: deterministic / batch-invariant kernels (at some throughput cost), fixed batch size of 1 for critical paths, same hardware and software versions.
  • Cache responses for identical inputs where business logic needs consistency.

Design for variance

  • Evals with multiple runs and pass rates, not single runs.
  • Structured outputs + validation.
  • Don't build logic that depends on exact string equality of free text.

Every expert started right here.