Most organizations running experiments on people — A/B tests, user studies, surveys, model evaluations — are making expensive decisions on designs that don't hold up. Underpowered samples. Broken randomization. Metrics that measure the wrong thing. Results that won't survive a second look.
I'm a methodologist who teaches this at the graduate level and writes the code to execute it. I review your design, tell you whether the result is real, and show you how to fix what isn't. The same rigor that goes into peer-reviewed experimental research — applied to the decisions your team is making this quarter.
A free, dependency-sequenced curriculum in research methods and statistics — from my graduate course at Texas A&M.
Browse the lessons →Applied pieces on measurement, evaluation, and AI — where the methods meet the leaderboard.
Read →Design reviews for model evals, experiments, and surveys. I tell you whether the result is real.
Services & rates →Graduate-school applications reviewed by someone who teaches at the level you're applying to.
How it works →Tell me what you're trying to measure.