Book 47 · Patriola’s Guide to Claude
Preference Data and Comparative Judgment
Rater A looked at two outputs and picked one. That fact tells you less than it appears to. Most A/B ratings are noisy not because the rater was careless but because the comparison was designed wrong. This book teaches how to build comparisons that hold up.
What a preference label actually carries
A preference label carries three things: the rater’s perceived difference, the axis they used, and the conditions under which the rating occurred. It does not carry whether the axis was the right one, whether conditions were equivalent across items, or whether another rater would agree. That gap is where most comparison designs fail quietly — clean numbers, consistent labels, sound statistics, and a result that measures something other than what the study claims.
This book builds comparisons that hold up: the A/B instrument from the axis down, condition control, systematic labeling across before/after data, order-effect and fatigue detection, inter-rater reliability for comparison pairs, automated judgment at scale, and the export format that turns a label into a reusable training artifact.
What you’ll learnEight chapters, one discipline
- what-a-comparison-measures — The three things a preference label actually carries — perceived difference, axis, and conditions — and why a label isn’t a finding on its own.
- building-the-ab-instrument — Structuring the trial: axis of variation, anchor design, paired vs. ranked formats, and how instrument shape determines what the label means.
- condition-control — What “controlled conditions” means operationally, and the failure modes — a different room, a different time of day — that quietly corrupt a comparison.
- systematic-labeling — Applying a label scheme across before/after data so the comparison is directly readable instead of two separate annotation passes.
- order-effects-and-rater-fatigue — Anchoring, contrast effects, and the fatigue slope across a rating session — detected by regressing score on position.
- irr-for-comparisons — Inter-rater reliability when the unit is a comparison pair, not a single label — a structurally different calculation with different failure signals.
- automated-comparison-at-scale — An inter-batch consistency scorer and a transcription-quality signal, reframed as comparative judgment instruments rather than pass/fail gates.
- from-comparison-to-training-signal — The export format that makes a preference label reusable: comparison ID, axis, scores, conditions, rater metadata, confidence.
A preview
Rater A looked at two outputs and picked one. That fact tells you less than it appears to. A preference label is the atomic unit of comparative evaluation — and the temptation is to read it as a finding. Resist it. A label is the starting material for a finding, not the finding itself.
Anyone whose comparisons need to survive scrutiny
Researchers and practitioners running A/B tests, before/after studies, or ranked preference collection who need the resulting labels to hold up as training signal, not just as data. This book builds on Designing Valid Evals (Book 42) and Annotation Scheme Design (Book 45) — readers should have both under their belt first.
A longer excerpt is available to newsletter subscribers.
More from Patriola
New books in this series
One short email per book launch.