Why Temperature Zero Is Not Reproducible: Floating-Point Arithmetic and Batching in Language Model Inference

22-minute read

Setting a language model’s temperature to zero is the usual way to ask it for the same answer every time. It does not reliably produce one, and the providers of hosted models say so in their own documentation. Anthropic’s API reference states that “even with temperature of 0.0, the results will not be fully deterministic.” OpenAI describes its seed parameter as “a best effort to sample deterministically” and adds that “determinism is not guaranteed.”

The cause lies neither in the model nor in the sampling rule. It is an artifact of how computers add numbers, combined with how inference servers share hardware among many users. This article follows that artifact from a single rounded addition to a model that gives two different answers to the same question, and measures each step along the way: a CPU and a GPU for the arithmetic, and an open-weight model, Qwen3-8B, served by the open-source inference engine vLLM, for the text. The later sections describe what it takes to make a model’s output reproducible down to the last bit, what that costs, and what to do when it is not possible.

What temperature zero promises

A language model writes one token at a time, where a token is a word or a fragment of one. At each step it assigns a score, called a logit, to every token in its vocabulary, and the scores are converted into probabilities. Temperature controls how strongly those probabilities favour the leading candidates: a high temperature spreads the choice across many plausible tokens, and a low one concentrates it on the best few. Temperature zero is the limit in which the model always takes the single highest-scoring token, a rule called greedy decoding.

Greedy decoding involves no randomness. Given the same scores, it chooses the same token every time. If a model at temperature zero gives different answers to the same prompt, then the scores themselves must have changed between the runs. The rest of this article explains how they change.

Arithmetic that rounds

Computers store real numbers in floating-point format: a fixed number of significant digits together with an exponent that places the decimal point, much as scientific notation does. Because the number of digits is fixed, most results of arithmetic cannot be stored exactly, and each is rounded to the nearest number the format can represent. The formats used in machine learning keep very different numbers of digits, and the ones used to run language models keep very few.

Two panels. The top panel is a horizontal bar chart of significant decimal digits: float64 keeps 15.9, float32 7.2, fp16 3.3, bf16 2.4 and fp8 E4M3 1.2. The bottom panel is a number line of every value bf16 can store between 248 and 266: consecutive integers up to 256, then only even numbers, so 257 has no representation and is stored as 256.

Figure 1: Digits are each format’s significand bits expressed in decimal. The lower panel enumerates every bf16 value between 248 and 266.

bf16 is a 16-bit format widely used to train and serve language models, and fp8 is increasingly used for low-precision serving. Because the number of digits is fixed, the gaps between storable numbers widen as the numbers grow, and the narrower the format, the sooner a gap becomes larger than the quantity being added.

Why the order of additions matters

In ordinary arithmetic, addition is associative: (a + b) + c equals a + (b + c). In floating point it is not, because each addition rounds its result before the next one begins. The same three numbers, grouped differently, can therefore give different totals:

float64   (0.1 + 0.2) + 0.3  =  0.6000000000000001     0.1 + (0.2 + 0.3)  =  0.6
float32   (1e8 + 1) - 1e8    =  0                      (1e8 - 1e8) + 1    =  1
bf16      (256 + 1) + 1      =  256                    256 + (1 + 1)      =  258

In the bf16 line, each 1 added to 256 is rounded away, whereas two 1s added to each other first survive as a 2. Over a long sum the effect accumulates. The next figure adds the same 100,000 numbers in 10,005 different orders.

Histogram of the float32 total of the same 100,000 positive numbers added in 10,000 random orders. The totals fall between 5,157 and 3,881 below the exact sum of 53,581,585.46 and take 261 distinct values. Five systematic orders are marked: largest first is 7,965.5 short, right to left 4,773.5 short, left to right 4,481.5 short, pairwise summation as NumPy performs it 1.5 short, and smallest first 42.5 over.

Figure 2: The exact sum was computed without rounding. Each random order is an independent shuffle of the same 100,000 values.

Two lessons follow. The order of the additions changes the answer, here by as much as 0.015 per cent of the total. And some orders are far more accurate than others: once a running total has grown large, it discards the small numbers added to it, whereas adding the small numbers first, or adding in pairs as NumPy does, preserves their digits. Numerical libraries prefer such pairwise, tree-shaped sums for that reason.

Rounding is not randomness, however. The IEEE 754 standard specifies every floating-point operation exactly, so the same numbers added in the same order give the same bits on every run; repeating the left-to-right sum 100 times produced a single result. The useful question is therefore not why the computer is random, but why the order of the additions changes.

Why the order changes

Parallel hardware divides the work

A GPU has thousands of arithmetic units, and a server CPU has dozens of cores. To keep them all busy, a numerical library divides a long sum into blocks, adds each block separately, and then combines the partial totals. The way it divides the work determines how the additions are grouped, and therefore how they round.

Line chart, on logarithmic axes, of how far the float32 total of the same 100,000 numbers falls short of the exact sum when the work is divided among 1 to 1,024 workers. One worker is 4,481.5 short; the shortfall falls steadily with more workers to 1.5 at 384 workers, then rises again to 29.5 at 1,024.

Figure 3: Each split is a fixed sequence of float32 additions, so rerunning it returns the same total.

No split involves any randomness, and each gives a different answer. Here more workers also meant a more accurate total, because each worker’s running total stayed small. A library chooses its split for speed, according to the shape of the problem in front of it. That is a sound engineering decision, and it has one consequence that matters here: the grouping of the additions depends on the shape of the input.

Shared servers batch requests together

An inference server does not process one request at a time. It gathers requests from many users into a batch and computes them together, because the hardware is far more efficient that way, and the size of the batch changes from moment to moment with the load. The shape of every large computation inside the model therefore depends on how busy the server is.

The next figure isolates the effect in the operation that dominates a language model’s work, a matrix multiplication. One row of inputs is multiplied by a fixed 4,096 × 4,096 matrix, first on its own and then as one row of batches of up to 2,048 rows, and the results are compared bit for bit.

Four horizontal strips, one per way of computing the product, coloured by whether the row's result matches the row computed alone, across batch sizes from 1 to 2,048. On the CPU in float32, every batch larger than one differs, in 3,421 outputs at two or three rows and 3,910 from four rows up. On the GPU in bf16 with default settings, batches up to 64 rows match and larger ones differ, in 1,600, 9, 1,507 or 9 outputs depending on the size. With split-K disabled, batches of 17 to 64 rows still differ in 9 outputs. A batch-invariant kernel matches at every size.

Figure 4: Numbers count the row’s outputs, of 4,096, that differ. Batch sizes 1 to 512 were all tested, then 16 more up to 2,048. Repeating any batch size 100 times gave identical bits in every mode.

The row did not change; only its neighbours did. On the CPU, every batch larger than one altered most of the row’s outputs. On the GPU, every batch of up to 64 rows reproduced the single-row result exactly, and no larger batch did. The differences are confined to the last digit or two that each format can hold, but they are not zero, and at any one batch size they never varied.

That last observation refutes a common explanation, that GPU threads finish in an unpredictable order and so produce random results. Thinking Machines Lab, whose 2025 analysis traced this problem in detail, makes the same point: “running the same matrix multiplication on the same data repeatedly will always provide bitwise equal results.” What varies is the batch. In their words, “the primary reason nearly all LLM inference endpoints are nondeterministic is that the load (and thus batch-size) nondeterministically varies!” They note that endpoints served from CPUs or TPUs share the problem, which the CPU strip above illustrates.

The two lower strips anticipate the remedy discussed later. A kernel, the routine that carries out one operation on the hardware, can be written so that each row is always summed in the same order whatever else is in the batch. The bottom strip is such a batch-invariant kernel, and it matched the single-row result at every size. The third strip shows a cheaper measure that disables only one batch-dependent strategy, split-K, described below; it removed most of the differences, but not all of them.

Where this happens inside a model

Three kinds of operation in a transformer, the architecture behind current language models, contain long sums whose division can depend on the batch:

  • Normalization layers (RMSNorm), which sum the squares of every value in a token’s internal representation.
  • Matrix multiplications, whose inner sums a library may split across cores and recombine, a strategy called split-K.
  • Attention, which sums over every earlier token in the text. Here the division can also depend on how the prompt was processed: in one piece or in chunks, and with earlier tokens recomputed or read back from the KV cache, the store of intermediate results a server keeps so that it need not compute them again.

Other sources of variation

Batching is the main cause, but not the only one.

  • Thread-order effects. Some operations let parallel threads add into a shared total in whatever order they finish, using what are called atomic additions. These are rare in inference but appear in training, notably in the backward pass of FlashAttention, so even an identical training setup can vary from run to run.
  • Different machines and software. A hosted service may route a request to a different GPU generation, parallel layout or library version, each with its own order of operations. PyTorch’s documentation is explicit: “Completely reproducible results are not guaranteed across PyTorch releases, individual commits, or different platforms.”
  • Changes behind the same model name. A provider can update a model or its serving stack without changing the name a client calls. OpenAI’s system_fingerprint field existed to signal such changes; both it and the seed parameter are now marked deprecated in OpenAI’s API specification.

From the last digit to a different answer

A difference in the last digit of a score usually changes nothing. At most steps one token leads clearly, and a tiny change in its score cannot unseat it. When two candidates are nearly tied, however, the noise decides the winner, and because every later token is generated from the text so far, a single flip changes the context for everything after it. From that point the two runs write different answers.

To measure how often this happens, I served Qwen3-8B, an open-weight model, with vLLM on one NVIDIA A100 GPU and asked it four questions at temperature zero, allowing up to 512 tokens per answer. Each question was submitted alone, and then together with between 1 and 127 other requests from a fixed pool, as if other users were sharing the server at that moment. Every configuration was run twice.

Three tile grids, one per serving configuration, with four questions as rows and fifteen batch sizes from 1 to 128 as columns. Each tile carries a letter naming a distinct answer, A being the answer when the question is asked alone. With default kernels, the Feynman, vaccine and lighthouse questions receive a different answer at nearly every batch size, and the pencil question at four sizes. With VLLM_BATCH_INVARIANT=1 as shipped the picture is much the same. With the Triton matrix multiplication forced on and compilation off, every question receives answer A at every batch size from 1 to 12, and other answers from 16 upward.

Figure 5: Each answer is up to 512 tokens. An amber dot marks a configuration whose second run gave a different answer from its first; the bottom configuration was run once.

Only the batch of one is reproducible here. Each question asked alone gave the same answer every time, yet with default kernels three of the four questions received a different answer at nearly every batch size, and the earliest split came at the ninth token. Even an identical batch was no guarantee: an amber dot marks a configuration whose second run differed from its first.

A further test isolated the cause of those reruns. Five runs of the same 96- or 128-request batch gave between two and five different answers with vLLM’s default engine, which receives requests from a separate process while it is already computing, so that the grouping of requests into computation steps depends on when each one arrives. With the engine run in-process, so that every request was queued before computation began, each configuration gave one answer every time, and the same answer on a second machine with the same GPU model and software. On a live server, arrival times vary on every request.

Where two answers part, the text on either side is ordinary. Three of the splits in the top panel:

"Vaccines train the immune system by"
    alone                    " mimicking an infection without causing the actual disease."
    with 4 other requests    " simulating an infection without causing the actual disease."

"... a lighthouse keeper named Elias Grant. He had lived alone in the tower for over"
    alone                    " a decade, tending to the light that guided ships"
    with 1 other request     " twenty years, tending to the light that guided ships"

"... an American theoretical physicist, Nobel laureate, and one of the most"
    alone                    " celebrated scientists of the 20th century."
    with 7 other requests    " influential and celebrated scientists of the 20th century."

In the arithmetic question, all three versions reached the same final answer, $10.00; they differed in wording, such as “one pencil” against “1 pencil”. The story, by contrast, changed a fact about its own character.

Bar chart of the gap between the two highest-scoring tokens at each of 1,768 generation steps of the four answers computed alone. Most steps have a clear leader: 804 have a gap over 5 and 415 between 2 and 5. Only 26 are exact ties and 84 have gaps of 0.125 or 0.25. All 15 points at which an answer split from the one computed alone fell at a tie (7) or a gap of 0.25 (8).

Figure 6: Scores are computed in bf16, so gaps below 1 come in steps of 0.125. Across batch sizes, the gap at a given step moved by 0.25 or more at about one step in five.

The scores themselves are rounded, which is why ties are common. vLLM computes them in bf16, the format of Figure 1, and every gap below 1 in these answers is a multiple of 0.125, the spacing between bf16 numbers from 16 to 32. At a few steps in every hundred, two candidates are therefore tied outright or one rounding step apart, and the last digits of all the arithmetic before them decide which side of a rounding boundary each falls on. A 512-token answer passes through dozens of such steps, and any one of them can become a fork. The story was the most exposed of the four answers: about one step in ten was a tie or a single rounding step from one, against one in thirty for the arithmetic, which had no exact ties at all. Open-ended text offers many continuations that are almost equally good.

The batch is not the only thing a real server varies. It also keeps the KV cache of recent requests, so that a repeated prefix, such as a long system prompt, need not be computed again. Serving the same 1,000-token prompt cold and then from a warm cache changed the answer to all four questions, at tokens 205, 9, 399 and 192 of the answers; the two warm runs agreed with each other. A prompt cache is a correct optimization, and it is also a change in the order of the arithmetic.

Thinking Machines Lab report the same behaviour at a larger scale. They sent the prompt “Tell me about Richard Feynman” to Qwen3-235B-A22B-Instruct-2507 1,000 times at temperature zero and received 80 different completions; all 1,000 agreed for the first 102 tokens, and at token 103, 992 continued “Queens, New York” while 8 wrote “New York City”.

Reproducible inference when you run the model yourself

Reproducibility requires every operation to add in the same order whatever else is in the batch. Kernels with that property are called batch-invariant, and both of the main open-source serving engines now offer them:

  • vLLM: set the environment variable VLLM_BATCH_INVARIANT=1. The feature is in beta, and its documentation lists NVIDIA GPUs of compute capability 8.0 or higher, a class that includes the A100, as well as Intel XPUs.
  • SGLang: start the server with --enable-deterministic-inference. Its batch-invariant attention supports the FlashInfer, FlashAttention 3 and Triton backends, and it remains compatible with chunked prefill, CUDA graphs and the radix cache.

On the GPU used here, the switch alone did not make the answers reproducible. In the middle panel of Figure 5, run with VLLM_BATCH_INVARIANT=1, the answers changed with the batch nearly as often as with default kernels. Figure 4 shows one reason. With PyTorch 2.9 on this GPU, the mode leaves matrix multiplications to cuBLAS with split-K disabled, which still depends on the batch size, and uses vLLM’s own batch-invariant Triton kernel only on newer GPUs; its source code refers to Hopper and Blackwell GPUs when it makes that choice.

Routing every matrix product through the Triton kernel made no difference while the model was compiled. With compilation off as well, every question received the same answer in every batch of up to 12 requests (bottom panel), and exactly the same answers again when the run was repeated on a second machine with the same GPU model. From 16 requests upward the answers still parted, though only after 120 tokens or more, and they did so even with every request queued before computation began, so some other operation still depends on the size of the batch. vLLM labels the feature beta, and on this hardware it is not yet a guarantee.

Two panels of paired timing results against default kernels. With VLLM_BATCH_INVARIANT=1 as shipped, one request of 256 tokens takes 1.3 percent less time (95 percent interval 1.0 to 1.9 percent less) and 1,000 requests of 100 tokens take 3.2 percent more (3.0 to 3.7). With the Triton matrix multiplication forced on and compilation off, one request takes 169 percent more time (169 to 176) and 1,000 requests take 45 percent more (45 to 46).

Figure 7: Faint points are pairs of fresh engine sessions, with the order of the two modes alternating; bars are 95 per cent bootstrap intervals on the median. As shipped: 10 pairs. With the Triton multiplication: 15 pairs, on a second machine with the same GPU model.

The switch as shipped costs little on this GPU because it changes little. The configuration that came closest to invariance switches off compilation and CUDA graphs and routes every matrix product through a Triton kernel, and it is far slower, most of all for a single stream of tokens. Published figures for complete implementations are of the same order. In Thinking Machines’ test, serving 1,000 sequences of 90 to 110 tokens with Qwen3-8B on one GPU took 26 seconds with default vLLM, 55 seconds with an unoptimized deterministic version and 42 seconds once they improved the attention kernel. LMSYS report an average slowdown of 34.35 per cent for SGLang’s deterministic mode, against 61.5 per cent for the Thinking Machines implementation.

Batch-invariant kernels remove the main cause, but full reproducibility also requires everything else to be held fixed:

  • the same model weights, engine version and library versions;
  • the same hardware type and parallel layout;
  • for temperatures above zero, a fixed sampling seed, which SGLang’s deterministic mode accepts per request.

None of this is available through a hosted API, where the client controls neither the batch nor the hardware. Anthropic’s newest models no longer accept a temperature other than the default at all: the same API reference marks the parameter deprecated, noting that “models released after Claude Opus 4.6 do not support setting temperature” and that any value other than 1.0 is rejected.

Where else the artifact matters

Reinforcement learning

In reinforcement learning for language models, one program generates practice answers, called rollouts, and another computes from them the gradients that improve the model. The first is usually an inference engine such as vLLM, chosen for speed, and the second a training framework. Each computes the probability of the same tokens with its own kernels, its own order of additions and its own choices about where to round, so the two disagree slightly. Training that is meant to be on-policy, learning from the model’s own current behaviour, becomes quietly off-policy.

Two panels. The left panel is a line chart of the share of 25,585 sampled tokens whose probability differs between vLLM and a Hugging Face forward pass by more than a threshold: 55.1 percent differ by more than 0.01 percent, 25.9 percent by more than 1 percent, 5.5 percent by more than 10 percent and 0.1 percent by more than 30 percent. The right panel is a histogram of the ratio of the trainer's probability of each whole answer to the sampler's, over 128 answers, spread from about 0.24 to 2.8 around a median of 0.90, with a dashed line at 1, where every answer would sit if the two agreed.

Figure 8: Answers were sampled at temperature 1 from 128 prompts of the fixed pool, up to 256 tokens each. The trainer’s pass uses PyTorch’s scaled-dot-product attention in bf16.

Per token the disagreement is small, with an estimated divergence (KL) of about 0.001 between the two programs’ distributions, but it compounds over an answer, and the right panel is what a reinforcement-learning update actually sees: in a correctly on-policy update, the ratio of the trainer’s probability to the sampler’s would be exactly 1 for every answer.

Thinking Machines measured the consequence in an RLVR setup, reinforcement learning with verifiable rewards, on the Bigmath dataset, starting from Qwen2.5-VL Instruct 8B. Without a correction for the mismatch, reward collapsed partway through training, together with a spike in the divergence between sampler and trainer. With an importance-weighting correction, which reweights each sample by the ratio in the right panel, the divergence stayed around 0.001 with occasional spikes and training proceeded smoothly. With sampler and trainer made bitwise identical, it stayed at exactly zero, and training also proceeded smoothly.

Training

The same random seed does not reproduce a training run if anything about its layout changes. A run on a different number of GPUs adds gradients across devices in a different order, and the weights drift apart as training proceeds. Atomic additions in backward passes add further variation even on an identical layout. PyTorch’s torch.use_deterministic_algorithms(True) switches to deterministic implementations where they exist and raises an error for operations that have none; its documentation warns that “deterministic operations are often slower than nondeterministic operations.”

Scientific computing

High-performance computing has lived with the same artifact for decades. A parallel sum over 64 processes and the same sum over 128 combine their partial results in different orders, so a distributed reduction can return slightly different totals at different scales, which is the experiment of Figure 3 on a larger machine. Intel’s oneMKL library offers a Conditional Numerical Reproducibility mode, selected with the MKL_CBWR environment variable, that returns bitwise identical results from run to run provided the number of threads does not change. Intel’s documentation warns that fixing the order of operations can cost 10 to 20 per cent of performance, and in some circumstances more than half. It is the same trade-off that language model serving now faces.

When a run cannot be replayed, record it

On a hosted API, and on any shared server without batch-invariant kernels, a run cannot be reproduced on demand. Rerunning a bad answer usually produces a different path through the text, often a correct one, and the failure appears never to have happened. The only reliable record of what happened is the one taken at the time, and that is the central idea of observability for language model systems.

  • Record each request as a trace. A trace is a structured record of one request, divided into spans for each model call, tool call and document retrieval. Each span keeps the model name and version, the parameters, the full input and output, the documents retrieved, and the tokens, time and cost consumed, together with any error, retry or fallback. OpenTelemetry, the common open standard for tracing, defines semantic conventions for generative AI so that such records look alike across tools.
  • Score the content, not only the call. A wrong answer still returns successfully, with no error code and a normal latency, so conventional monitoring sees nothing amiss. Quality has to be checked on the content itself: whether each claim is supported by the source it cites, whether a second model acting as a judge accepts it against a rubric, whether simple rules hold (the answer cites something, or declines when it should), and what readers think of it.
  • Measure rates rather than single runs. Every run is one sample from a distribution, so quality is a rate to be tracked over time, and a slow drift, such as a provider updating a model behind an unchanged name, appears as a trend rather than a single complaint. The same holds offline: a test that runs once and passes proves little about an answer that can change between runs, so each case should be run several times and reported as a pass rate.
  • Turn failures into tests. A failure caught in a trace becomes a case in the offline test set, so that the fix can be checked against it and later versions cannot quietly reintroduce it.
  • Decide what to store. A full trace contains people’s questions and the system’s answers. What to keep, for how long and who may read it should be settled before tracing is enabled; storing identifiers, counts and scores for every request, and full text only where it is needed and permitted, is a sound default.

Practical guidance

When using a hosted API:

  • Do not rely on temperature zero for reproducibility; the newest models may not accept it at all.
  • Record every request as a trace, including the model version and any backend identifier the provider returns.
  • Score answers automatically, and track the scores as rates over time.
  • Evaluate with repeated runs, and report pass rates rather than single outcomes.

When running a model yourself:

  • Enable batch-invariant kernels (VLLM_BATCH_INVARIANT=1 in vLLM, --enable-deterministic-inference in SGLang) when a run must be replayable, and budget for the slowdown.
  • Pin the weights, the engine and library versions, the hardware type and the parallel layout.
  • Fix the sampling seed whenever the temperature is above zero.
  • Verify invariance on your own hardware before relying on it: run one request alone and inside batches of several sizes, and compare the outputs token for token.

When training:

  • Use torch.use_deterministic_algorithms(True) when debugging requires exact replays.
  • Keep the GPU count and layout fixed when comparing runs, and expect divergence otherwise.
  • In reinforcement learning, correct for the mismatch between sampler and trainer, or eliminate it.

Methods

All measurements were made on October 7, 2026, with Python 3.11, NumPy 2.2.6 and PyTorch 2.9.1.

  • Arithmetic (Figures 1 to 3). The 100,000 values are 10 raised to powers drawn uniformly between −4 and 4 with NumPy’s default generator (seed 20261007), stored in float32. Every float32 total is a strict left-to-right accumulation except NumPy’s built-in sum; the exact sum of the stored values is computed with Python’s math.fsum.
  • Kernels (Figure 4). Inputs are standard normal (seed 0) and the weights are scaled by 1/64. The CPU run uses float32 on a 64-core AMD EPYC 7702 with 8 threads; the GPU runs use bf16 on one NVIDIA A100-SXM4 with 40 GB. “Split-K disabled” applies the cuBLAS workspace limits that vLLM 0.11.2 sets in batch-invariant mode on this GPU with this PyTorch version, which are meant to rule out split-K; the batch-invariant kernel is the Triton matrix multiplication that vLLM ships for the same purpose.
  • Text (Figures 5 to 8). vLLM 0.11.2 and Transformers 5.6.2 on one A100-SXM4, serving Qwen3-8B in bf16 with its thinking mode disabled. Greedy decoding with up to 512 new tokens per answer and prefix caching off, except in the cache test, which adds a 1,000-token system prompt. The other requests in a batch are the first of 256 prompts built from 32 topics and 8 templates, each with a fixed output budget between 16 and 512 tokens, also decoded greedily. Each question and batch size was run twice, in shuffled order. The rerun test repeats one batch five times with vLLM’s engine core in its default separate process and five times in-process (VLLM_ENABLE_V1_MULTIPROCESSING=0), on each of two machines. “With the Triton matrix multiplication” registers vLLM’s batch-invariant Triton kernel for every matrix product, which vLLM 0.11.2 itself does only on compute capability 10.0 with PyTorch 2.9, and runs without compilation or CUDA graphs (enforce_eager=True); it ran on a second machine with the same GPU model, and once more there with the engine in-process. The timing comparisons alternate fresh engine sessions in pairs, switching which mode runs first; each session discards two warm-up repetitions of each workload and times five (10 pairs, as shipped) or six (15 pairs, with the Triton multiplication), with output lengths forced so that both modes do the same work. Within each comparison, every pair moved in the same direction. Intervals are percentile bootstrap intervals (10,000 resamples) on the median ratio of paired session medians, and p-values come from the Wilcoxon signed-rank test on the same pairs. The sampler-trainer comparison samples 128 answers at temperature 1 (seed 1234, up to 256 tokens) and scores them with a Transformers forward pass in bf16 using PyTorch’s scaled-dot-product attention.

Sources