Benchmarking studies to enhance rigor in prediction modeling
At the SOUP Lab, we care a lot about getting the “recipe” right when it comes to evaluating prediction models. Our benchmarking work is all about creating fair, transparent, and reproducible ways to compare models and design choices—so we’re not just tasting results, but really understanding how they were made. We design benchmarking studies that go beyond simple performance metrics, carefully accounting for data quirks, real-world complexity, and what happens when models are served in new settings. In other words, we try to bring consistency to a field that can sometimes feel like everyone is cooking their own version of the same dish. By building principled benchmarking frameworks, we aim to raise the standard for prediction modeling research and ensure that what looks good in the lab actually holds up when it’s served in practice.