← ERS STAGE 1 · EVIDENCE BOOK34 occasions · 5 raters / 3 complete

Inter-rater agreement

PLAIN ANSWER

The three complete raters often chose differently. With only ten shared claims, the agreement numbers are too unstable to support a broad conclusion.

We can compare two raters only on claims both people saw. The three complete raters form 3 pairs, and each pair shares all ten claims. We apply the same check to every pair instead of selecting the pair that looks most interesting.

ALL COMPLETE-RATER PAIRS · 3 PAIRS · MARGINAL-INDEPENDENCE BENCHMARK

R05 × R08 · 10 shared claims

ChoiceObservedMarginal expectedUniform referenceκ
MOST0%6%25%-0.064
LEAST40%21%25%0.241

R05 × R14 · 10 shared claims

ChoiceObservedMarginal expectedUniform referenceκ
MOST20%17%25%0.036
LEAST20%17%25%0.036

R08 × R14 · 10 shared claims

ChoiceObservedMarginal expectedUniform referenceκ
MOST30%43%25%-0.228
LEAST30%32%25%-0.029

The chance comparison is one reasonable reference for a sample this small. It is not a pass/fail line, and another reasonable chance model could give a different number.

With only ten shared claims, changing one choice moves the result noticeably. A below-zero score in such a small sample does not prove that the raters systematically disagree.

These rows answer one narrow question: did a pair agree more than its own choice habits would suggest? They do not settle that question, and they say nothing about which method is better.