B leads all 10 drops, no tie at the top · shaded band = B range +3 to +8
bottom pair A/C/D is not stably ordered — tied in 3 of 10 drops, and the order reverses between drops
n 29–31 per drop against a full sample of 34 · adjacent rows share nearly all of their data
B stays first in every claim-removal check. The order below B changes. That means these checks
support only a narrow statement about B under this one counting rule.
Removing one claim at a time keeps B first in all 10 checks. The rest of the ranking still changes.
C is above D in 4 checks, D is above C in 5, and they tie in 1.
AUDIT DETAILS · D · 10 drops · stage1.sensitivity.leave_one_claim
Removing one rater is the harder check, and the result fails it. Across the
5 checks, the leader changes once. Removing the rater who supplied
the most of the 34 completed ratings gives the lead to a different
method. This is why the headline describes saved ratings instead of declaring a better method.
Removing one rater at a time changes the leader: B leads in four checks and D in one.
Removing R08 puts D first. The other four checks leave B first.
AUDIT DETAILS · D · 5 drops · stage1.sensitivity.leave_one_rater
We also count the same pilot under four rules for which ratings to include. These are four views of
one small dataset, not four separate experiments.
ELIGIBILITY RULES · NET BEST−WORST RECOMPUTED 4× ON ONE PILOT
A plain · B structured scope · C full ERS · D ablation — n = judgment occasions the rule keeps, of 34 · overlapping re-tallies of one sample, not replications
all judgment occasions (default descriptive aggregation) · full sample
34
0+6-3-3
ABC=D
≠ B > A > C=D
complete raters only
30
0+7-2-5
ABCD
B > A > C > D
excluding recovered records
31
0+6-2-4
ABCD
B > A > C > D
excluding the defect claim
30
0+6-1-5
ABCD
B > A > C > D
the rules do not agree — 2 distinct orderings across 4 rules: B > A > C=D · B > A > C > D
the arm totals still move under the rules (B +6…+7, C -3…-1, D -5…-3)
n 30–34 of 34 occasions from 5 raters · agreement across overlapping slices of one small pilot is a weak stability check, not evidence about anyone outside this sample
B leads in all four views, but C and D change places or tie. The rows share most of the same data,
so this agreement is weaker than repeating the study with new people.
Weighting is a choice about what gets one vote; it is not a repair. If each completed rating gets
one vote, a person who finished ten claims counts more than a person who finished one. If each
person gets the same total weight, those people count equally. The ranking changes. The report
shows both views instead of choosing the one that makes a preferred method look best.
Giving each rater the same total weight changes the ranking to D > A > B > C.
Under this rule, a rater who completed one claim gets the same total weight as a rater who completed ten.
AUDIT DETAILS · D · mean of per-rater net rates across 5 raters · stage1.sensitivity.equal_rater
None of these checks repeats the study. Each one removes data from the same pilot. They can reveal a
fragile ranking, but they cannot prove a ranking will hold. That requires a new, independent study.