Latest articles

14-minute read

Evaluating AI Agents When Every Run Differs: Repeated Trials, Paired Comparisons and Graders

An agent that completes a task today may fail the same task tomorrow, so a single run of an evaluation is one draw from a distribution rather than a measurement of the agent. Comparing two versions of an agent therefore needs the tools of an experiment: repeated trials, comparisons made on the same tasks, an interval around every difference, and graders whose own errors are known.

All 129 articles

Browse by topic

R analyses, 2021–2022

Shorter analyses in R from 2021 and 2022, most of them on TidyTuesday datasets.

108 posts

  1. Using LASSO Bootstraps to Predict Animal Crossing Rating
  2. LASSO Model on Predicting GDPR Fines
  3. PCA on Cocktail Ingredients (with recipes package used)
  4. All 108 in the archive