ERS STAGE 1 · EVIDENCE BOOK24 occasions · 4 raters / 2 complete

Agreement is measurable only where two raters saw the same claim. In this pilot that is the 10 claims the two complete raters share, and nowhere else. The reference plotted against it is not the naive one-in-four line but the agreement each rater’s own selection habits already produce by chance, derived from their observed marginals.

INTER-RATER AGREEMENT · R08 × R14 · 10 SHARED CLAIMS
observedexpected from their own habitsuniform 1 in 4

The objection to the ordering is that two people simply agreed with each other. If that were what produced it, agreement between these two would have to sit above what each rater's own selection habits already generate by chance. Here it lands at or below that line on MOST and LEAST.

0.000.250.50
MOST3 of 10 agreed
0.300.43
obs−exp−0.13κ−0.228
LEAST3 of 10 agreed
0.300.32
obs−exp−0.02κ−0.029
proportion of the 10 shared claims · scale 0–0.50 of 1.00
arms they agreed on
most
B structured scope3/10
least
D ablation2/10
A plain1/10
selection habits · benchmark input
ABCDR08 most1810R08 least2215R14 most3502R14 least2035
what the benchmark is

Expected agreement under independence, conditional on the two observed marginal distributions — an appropriate chance correction, not uniquely “the correct” one at 10 shared claims. Read the hollow mark as one defensible reference among several, not as the answer. src/ers_evidence/agreement.py

what this does not establish

2 raters over 10 shared claims is a very small base. A κ below zero here is not evidence of systematic disagreement; it is what small samples do. This weakens one alternative explanation of the ordering — that it came from these two agreeing item by item — and nothing more. It says nothing about the arms.

That benchmark is one defensible chance correction at this size, not uniquely the correct one, and src/ers_evidence/agreement.py says so where the quantity is computed. The hollow mark is a reference, not a threshold.

At 10 shared claims from 2 raters the quantity is unstable by construction — a single changed selection moves it. A kappa below zero at this size is what small samples do, not evidence that these raters systematically disagree.

The comparison bears on one objection and no others. If the aggregate ordering were an artifact of two people agreeing item by item, their agreement would have to clear what their own habits already generate. On this pair it does not. That weakens the objection; it does not settle it, and it says nothing about the arms.