// book 47
patriola.com

Book 47 · Patriola’s Guide to Claude

Preference Data and Comparative Judgment


Rater A looked at two outputs and picked one. That fact tells you less than it appears to. Most A/B ratings are noisy not because the rater was careless but because the comparison was designed wrong. This book teaches how to build comparisons that hold up.

Buy Ebook on Amazon

Patriola's Guide to Claude — Preference Data and Comparative Judgment: Design Comparisons That Produce Useful Training Signal
What this book is

What a preference label actually carries

A preference label carries three things: the rater’s perceived difference, the axis they used, and the conditions under which the rating occurred. It does not carry whether the axis was the right one, whether conditions were equivalent across items, or whether another rater would agree. That gap is where most comparison designs fail quietly — clean numbers, consistent labels, sound statistics, and a result that measures something other than what the study claims.

This book builds comparisons that hold up: the A/B instrument from the axis down, condition control, systematic labeling across before/after data, order-effect and fatigue detection, inter-rater reliability for comparison pairs, automated judgment at scale, and the export format that turns a label into a reusable training artifact.

What you’ll learn

Eight chapters, one discipline

  • what-a-comparison-measures — The three things a preference label actually carries — perceived difference, axis, and conditions — and why a label isn’t a finding on its own.
  • building-the-ab-instrument — Structuring the trial: axis of variation, anchor design, paired vs. ranked formats, and how instrument shape determines what the label means.
  • condition-control — What “controlled conditions” means operationally, and the failure modes — a different room, a different time of day — that quietly corrupt a comparison.
  • systematic-labeling — Applying a label scheme across before/after data so the comparison is directly readable instead of two separate annotation passes.
  • order-effects-and-rater-fatigue — Anchoring, contrast effects, and the fatigue slope across a rating session — detected by regressing score on position.
  • irr-for-comparisons — Inter-rater reliability when the unit is a comparison pair, not a single label — a structurally different calculation with different failure signals.
  • automated-comparison-at-scale — An inter-batch consistency scorer and a transcription-quality signal, reframed as comparative judgment instruments rather than pass/fail gates.
  • from-comparison-to-training-signal — The export format that makes a preference label reusable: comparison ID, axis, scores, conditions, rater metadata, confidence.
From the book

A preview

Rater A looked at two outputs and picked one. That fact tells you less than it appears to. A preference label is the atomic unit of comparative evaluation — and the temptation is to read it as a finding. Resist it. A label is the starting material for a finding, not the finding itself.
You — reviewing an A/B session Rater A picked B in 8 of 10 pairs. Can I call B the better output?
Claude Only if you can also answer three questions: what axis did the rater use, were the conditions held constant across the ten pairs, and could the rater actually perceive the difference you care about? An 8-of-10 split with an undocumented axis and no condition log tells you a rater had a preference. It doesn’t yet tell you B is better at the thing you’re measuring.
Who it’s for

Anyone whose comparisons need to survive scrutiny

Researchers and practitioners running A/B tests, before/after studies, or ranked preference collection who need the resulting labels to hold up as training signal, not just as data. This book builds on Designing Valid Evals (Book 42) and Annotation Scheme Design (Book 45) — readers should have both under their belt first.

A longer excerpt is available to newsletter subscribers.

Buy Ebook on Amazon

Stay current

New books in this series

One short email per book launch.