I am a computational scientist at the University of Chicago. I write long-form articles on training language models from scratch, on open-source tools for high-performance computing, and on what tens of millions of job postings and rental listings can and cannot show.

Latest articles

View all

11-minute read

Tuning the GPU Interconnect in Multi-Node Language Model Pretraining

Training a language model across several machines costs something that training on one machine does not. At every optimizer step, each GPU computes gradients, the corrections the model learns from, and every GPU’s copy has to be added together.

Browse by topic

R analyses, 2021–2022

Shorter analyses in R from 2021 and 2022, most of them on TidyTuesday datasets.

108 posts

  1. Using LASSO Bootstraps to Predict Animal Crossing Rating
  2. LASSO Model on Predicting GDPR Fines
  3. PCA on Cocktail Ingredients (with recipes package used)
  4. All 108 in the archive