// book 42
patriola.com

Book 42 · Patriola’s Guide to Claude

Designing Valid Evals


What your benchmark is and isn't measuring. Construct validity, confound control, and decomposing aggregate scores back into the failures they hide.

Buy Ebook on Amazon

Patriola's Guide to Claude — Designing Valid Evals: What Your Benchmark Is and Isn't Measuring
What this book is

The passing score was a guess wearing the costume of a result

An abstract said 162 items. The data set held 157. Nobody caught the gap across four revisions of the paper before it was marked ready to submit — an audit found it in the first hour. That was the smallest of three failures. The methodology section made directional claims about which subgroup showed wider capability, and every one of those claims pointed the wrong way once the numbers were rerun. The measurements were accurate to the decimal; they just came from a different comparison than the one the sentences described. The third failure was quietest of all: cited means that reproduced perfectly, and effect sizes derived from a denominator that contradicted them. None of it was sabotage — a competent team produced careful work and passed every review their process required, because reading a paper and running its numbers are two different acts.

Most evals carry exactly this kind of hidden defect, and most never get audited, so the defect ships. Teams build a benchmark, run a model against it, read a score, and report the score as evidence. What they actually have is a hypothesis: this number reflects the capability we care about. Until somebody pressure-tests the instrument — confirms it measures what it claims, controls for what else could move the score, and breaks the aggregate back down into the failure shapes it papers over — the score is a guess wearing the costume of a result.

This book turns that audit into a repeatable methodology: seven chapters that take you from a concrete audit failure through construct validity, confound control, the statistical masking effect of aggregate scores, and wiring the resulting gate into a pipeline so it runs on every merge and fails loudly when something drifts.

What you’ll learn

A methodology, not a statistics lecture

  • Construct validity — defining what an eval actually measures before the items are written, so the target can't drift out from under the benchmark.
  • Confound control — isolating the conditions that could move a score without touching the capability you're claiming to test.
  • Aggregate scores as statistical masking — why a clean summary number hides the failure rows that built it, and how to decompose one back apart.
  • A metric built from scratch — a worked example: an AI-tell detection engine built from first principles, starting from the failure and naming the behaviors before writing a single detection rule.
  • Setting a passing score — where the threshold comes from, and why an arbitrary cutoff is just a confound wearing a number.
  • Wiring the eval into the pipeline — running the gate automatically on every merge, so a drifting score fails loudly instead of shipping quietly.
From the book

A preview

A passing review tells you the eval looked right to a reader. Validity is a different property, and a harder one: the eval has to measure what it says it measures, survive recomputation, and hold up when its summary numbers are taken apart.
You The aggregate score passed. Before I report it — what does the score actually stop meaning if I never decompose it?
Claude An eval that finds nothing doesn't earn a clean record by default — either the system genuinely has no failures, which is possible but rare, or the eval is measuring nothing at all. A summary score papers over the rows that built it, and a passing aggregate can sit directly on top of a real failure pattern the mean cancels out. Decomposing it isn't optional diligence; it's the step that tells you which of those two you actually have.
Who it’s for

Operators whose quality signals need to actually signal something

This book is for anyone building or reviewing benchmarks, capability evals, or automated quality gates for AI systems run at scale. It assumes no statistics background beyond what the worked examples build as they go — an AI-tell detection engine constructed from first principles, and a white paper audit that surfaced three failures after the paper had already been marked ready to submit. It does assume you're past the point of trusting a passing score just because the code ran and the number came back above the threshold.

A longer excerpt is available to newsletter subscribers.

Buy Ebook on Amazon

Stay current

New books in this series

One short email per book launch.