Evaluating AI Agents When Every Run Differs: Repeated Trials, Paired Comparisons and Graders
An agent that completes a task today may fail the same task tomorrow, so a single run of an evaluation is one draw from a distribution rather than a measurement of the agent. Comparing two versions of an agent therefore needs the tools of an experiment: repeated trials, comparisons made on the same tasks, an interval around every difference, and graders whose own errors are known.