Writing

Writing

Applied pieces on measurement, evaluation, and AI.

Where the methods meet the leaderboard. These pieces take the foundations from the Methods Library and run them at live problems — model evaluations, preference data, the numbers companies ship decisions on.

What Your LLM Eval Got Wrong: A Methodologist's Checklist

Seven failures that quietly invalidate model evaluations — reliability, blinding, sampling, power — and what the fix looks like. June 2026.