Book 42 · Patriola’s Guide to Claude
Designing Valid Evals
What your benchmark is and isn't measuring. Construct validity, confound control, and decomposing aggregate scores back into the failures they hide.
The passing score was a guess wearing the costume of a result
An abstract said 162 items. The data set held 157. Nobody caught the gap across four revisions of the paper before it was marked ready to submit — an audit found it in the first hour. That was the smallest of three failures. The methodology section made directional claims about which subgroup showed wider capability, and every one of those claims pointed the wrong way once the numbers were rerun. The measurements were accurate to the decimal; they just came from a different comparison than the one the sentences described. The third failure was quietest of all: cited means that reproduced perfectly, and effect sizes derived from a denominator that contradicted them. None of it was sabotage — a competent team produced careful work and passed every review their process required, because reading a paper and running its numbers are two different acts.
Most evals carry exactly this kind of hidden defect, and most never get audited, so the defect ships. Teams build a benchmark, run a model against it, read a score, and report the score as evidence. What they actually have is a hypothesis: this number reflects the capability we care about. Until somebody pressure-tests the instrument — confirms it measures what it claims, controls for what else could move the score, and breaks the aggregate back down into the failure shapes it papers over — the score is a guess wearing the costume of a result.
This book turns that audit into a repeatable methodology: seven chapters that take you from a concrete audit failure through construct validity, confound control, the statistical masking effect of aggregate scores, and wiring the resulting gate into a pipeline so it runs on every merge and fails loudly when something drifts.
What you’ll learnA methodology, not a statistics lecture
- Construct validity — defining what an eval actually measures before the items are written, so the target can't drift out from under the benchmark.
- Confound control — isolating the conditions that could move a score without touching the capability you're claiming to test.
- Aggregate scores as statistical masking — why a clean summary number hides the failure rows that built it, and how to decompose one back apart.
- A metric built from scratch — a worked example: an AI-tell detection engine built from first principles, starting from the failure and naming the behaviors before writing a single detection rule.
- Setting a passing score — where the threshold comes from, and why an arbitrary cutoff is just a confound wearing a number.
- Wiring the eval into the pipeline — running the gate automatically on every merge, so a drifting score fails loudly instead of shipping quietly.
A preview
A passing review tells you the eval looked right to a reader. Validity is a different property, and a harder one: the eval has to measure what it says it measures, survive recomputation, and hold up when its summary numbers are taken apart.
Operators whose quality signals need to actually signal something
This book is for anyone building or reviewing benchmarks, capability evals, or automated quality gates for AI systems run at scale. It assumes no statistics background beyond what the worked examples build as they go — an AI-tell detection engine constructed from first principles, and a white paper audit that surfaced three failures after the paper had already been marked ready to submit. It does assume you're past the point of trusting a passing score just because the code ran and the number came back above the threshold.
A longer excerpt is available to newsletter subscribers.
More from Patriola
New books in this series
One short email per book launch.