1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your application sends the same prompt with temperature 0 but gets slightly different outputs. Why, and how do you get reproducibility?
30-second answerSay your answer out loud first, then reveal.
Sources of non-determinism
- Floating-point non-associativity: (a + b) + c ≠ a + (b + c) in finite precision. Parallel reductions in matrix multiplies and attention can change order.
- Batch-dependent kernels: the same request batched with different other requests can use different kernel strategies or tile splits, which changes numerics. This is a major source in served LLMs.
- MoE routing: capacity limits and token dropping can depend on what else is in the batch.
- Near ties: two tokens with almost equal logits, so a tiny difference flips the choice.
- Different hardware / software versions across a provider's fleet.
- Model updates behind an alias such as "latest".
Getting more reproducibility
- Pin exact model versions; use seed parameters where offered (usually "best effort").
- Self-hosted: deterministic / batch-invariant kernels (at some throughput cost), fixed batch size of 1 for critical paths, same hardware and software versions.
- Cache responses for identical inputs where business logic needs consistency.
Design for variance
- Evals with multiple runs and pass rates, not single runs.
- Structured outputs + validation.
- Don't build logic that depends on exact string equality of free text.
Related
Every expert started right here.