The tally is recomputed 10 times, once per claim, each run omitting that claim’s occasions. Every row is the whole analysis rerun — not a subgroup read off the full one.
The same arm leads every drop, with no tie at the top. Beneath it the ordering does not hold: the bottom pair collapses into a tie in some drops and separates in the rest, so the sequence below the leader is not something this sample fixes.
Leave-one-claim-out keeps B first and A second in 10 of 10 drops.
C exceeds D in 7 of 10 and ties D in 3 of 10.
Dropping a rater is the harder test, because two raters supplied most of the 24 occasions between them. Across the 4 rater drops the leader again does not change, and again the arms beneath it do — one drop reverses the bottom two, another leaves second place shared.
B leads under all four leave-one-rater-out drops; the lower-arm ordering is not stable.
Omitting R14 yields B > A=C > D; omitting R08 yields B > A > D > C.
Eligibility is a third lever. The same tally recomputed under four different admission rules is not four experiments; it is one pilot counted four ways over heavily overlapping occasions.
The rank order does not move across those rules, though the margins do. That is the weakest kind of agreement available: the rows share most of their data, so they were never free to disagree by much.
Weighting is a choice of unit, not a correction. The default aggregation weights judgment occasions, so a rater who rated every claim carries more of the total than a rater who rated one. Weight raters equally instead and a rater with a single occasion carries as much as a rater with ten. The ordering does not survive that change, and both units are defensible; the first is the one reported.
Under equal-rater weighting the ordering reverses to D > B > A > C.
Different unit; not count-comparable. A one-observation rater carries equal weight.
None of these are replications. A drop removes occasions and never adds any, so every row is a re-tally of a subset of the same pilot and adjacent rows share nearly all of their data. Checks like these can show an ordering coming apart, which is what happens below the leader. They cannot show one holding up. That would take a second collection.