Why Temperature Zero Is Not Reproducible: Floating-Point Arithmetic and Batching in Language Model Inference
Setting a language model’s temperature to zero is the usual way to ask it for the same answer every time. It does not reliably produce one, and the providers of hosted models say so in their own documentation.