Exact counts are a direct sample fact. The ordering asserted from them is a separate, pilot-restricted claim — the default record-level aggregation is a descriptive target, not a population estimator.
A received 4 MOST and 4 LEAST; B 14 and 4; C 2 and 6; D 4 and 10.
In this pilot, the default record-level tally ordered the arms B > A > C > D.
Default descriptive aggregation, not a population estimator. Selected after collection.
The cohort behind those counts is uneven, and the unevenness governs everything after it.
Two raters reached every claim. The other two contributed a handful of occasions between them, and one of those profiles has no single modal choice at all — its selections tie across three arms. Any weighting rule is therefore a decision about how much those two short profiles ought to count.
Weighting sensitivities do not all agree, and the disagreement is the point.
Restricted to the two complete raters, the ordering is unchanged at B > A > C > D.
Under equal-rater weighting the ordering reverses to D > B > A > C.
Different unit; not count-comparable. A one-observation rater carries equal weight.
B leads under all four leave-one-rater-out drops; the lower-arm ordering is not stable.
Omitting R14 yields B > A=C > D; omitting R08 yields B > A > D > C.