R05 × R08 · 10 shared claims
| Choice | Observed | Marginal expected | Uniform reference | κ |
|---|---|---|---|---|
| MOST | 0% | 6% | 25% | -0.064 |
| LEAST | 40% | 21% | 25% | 0.241 |
The three complete raters often chose differently. With only ten shared claims, the agreement numbers are too unstable to support a broad conclusion.
We can compare two raters only on claims both people saw. The three complete raters form 3 pairs, and each pair shares all ten claims. We apply the same check to every pair instead of selecting the pair that looks most interesting.
| Choice | Observed | Marginal expected | Uniform reference | κ |
|---|---|---|---|---|
| MOST | 0% | 6% | 25% | -0.064 |
| LEAST | 40% | 21% | 25% | 0.241 |
| Choice | Observed | Marginal expected | Uniform reference | κ |
|---|---|---|---|---|
| MOST | 20% | 17% | 25% | 0.036 |
| LEAST | 20% | 17% | 25% | 0.036 |
| Choice | Observed | Marginal expected | Uniform reference | κ |
|---|---|---|---|---|
| MOST | 30% | 43% | 25% | -0.228 |
| LEAST | 30% | 32% | 25% | -0.029 |
The chance comparison is one reasonable reference for a sample this small. It is not a pass/fail line, and another reasonable chance model could give a different number.
With only ten shared claims, changing one choice moves the result noticeably. A below-zero score in such a small sample does not prove that the raters systematically disagree.
These rows answer one narrow question: did a pair agree more than its own choice habits would suggest? They do not settle that question, and they say nothing about which method is better.