ERS STAGE 1 · EVIDENCE BOOK24 occasions · 4 raters / 2 complete

The tally is recomputed 10 times, once per claim, each run omitting that claim’s occasions. Every row is the whole analysis rerun — not a subgroup read off the full one.

LEAVE ONE CLAIM OUT · TALLY RECOMPUTED 10× · FULL SAMPLE B > A > C > D AT n=24
A plain · B structured scope · C full ERS · D ablation — n = occasions surviving the drop, not independent replications
omitted claimnABCDnet best−worst · shared scaleorderingbottom pair
eu_non_authorized_not_substantiated-13200+12-5-7
ABCD
B > A > C > DC > D
eu_non_authorized_not_sufficiently_defined-4210+9-2-7
ABCD
B > A > C > DC > D
eu_authorized_conditioned-144210+7-2-5
ABCD
B > A > C > DC > D
eu_authorized_conditioned-489-2899220+8-4-4
ABC=D
B > A > C=DC=D
foodkg-activated-charcoal22-1+9-3-5
ABCD
B > A > C > DC > D
healthfc-13220+8-3-5
ABCD
B > A > C > DC > D
healthfc-2122-1+9-4-4
ABC=D
B > A > C=DC=D
public-guidance-eatright-seed-oils22+2+9-4-7
ABCD
B > A > C > DC > D
fda_d_tagatose_denial220+10-5-5
ABC=D
B > A > C=DC=D
fda_oleic_acid_discretion220+9-4-5
ABCD
B > A > C > DC > D
B leads all 10 drops, no tie at the top · shaded band = B range +7 to +12
bottom pair C/D is not stably ordered — tied in 3 of 10 drops, C > D in the rest
n 20–22 per drop against a full sample of 24 · adjacent rows share nearly all of their data

The same arm leads every drop, with no tie at the top. Beneath it the ordering does not hold: the bottom pair collapses into a tie in some drops and separates in the rest, so the sequence below the leader is not something this sample fixes.

Leave-one-claim-out keeps B first and A second in 10 of 10 drops.

C exceeds D in 7 of 10 and ties D in 3 of 10.

D · 10 drops · stage1.sensitivity.leave_one_claim

Dropping a rater is the harder test, because two raters supplied most of the 24 occasions between them. Across the 4 rater drops the leader again does not change, and again the arms beneath it do — one drop reverses the bottom two, another leaves second place shared.

B leads under all four leave-one-rater-out drops; the lower-arm ordering is not stable.

Omitting R14 yields B > A=C > D; omitting R08 yields B > A > D > C.

D · 4 drops · stage1.sensitivity.leave_one_rater

Eligibility is a third lever. The same tally recomputed under four different admission rules is not four experiments; it is one pilot counted four ways over heavily overlapping occasions.

ELIGIBILITY RULES · NET BEST−WORST RECOMPUTED 4× ON ONE PILOT
A plain · B structured scope · C full ERS · D ablation — n = judgment occasions the rule keeps, of 24 · overlapping re-tallies of one sample, not replications
eligibility rulenABCDnet best−worst · shared scaleordering
all judgment occasions (default descriptive aggregation) · full sample
24
0+10-4-6
ABCD
B > A > C > D
complete raters only
20
0+11-3-8
ABCD
B > A > C > D
excluding recovered records
21
0+10-3-7
ABCD
B > A > C > D
excluding the defect claim
21
0+9-2-7
ABCD
B > A > C > D
all 4 rules return the same ordering — B > A > C > D
the arm totals still move under the rules (B +9…+11, C -4…-2, D -8…-6) — the numbers shift, the rank order does not
n 20–24 of 24 occasions from 4 raters · agreement across overlapping slices of one small pilot is a weak stability check, not evidence about anyone outside this sample

The rank order does not move across those rules, though the margins do. That is the weakest kind of agreement available: the rows share most of their data, so they were never free to disagree by much.

Weighting is a choice of unit, not a correction. The default aggregation weights judgment occasions, so a rater who rated every claim carries more of the total than a rater who rated one. Weight raters equally instead and a rater with a single occasion carries as much as a rater with ten. The ordering does not survive that change, and both units are defensible; the first is the one reported.

Under equal-rater weighting the ordering reverses to D > B > A > C.

Different unit; not count-comparable. A one-observation rater carries equal weight.

D · mean of per-rater net rates across 4 raters · stage1.sensitivity.equal_rater

None of these are replications. A drop removes occasions and never adds any, so every row is a re-tally of a subset of the same pilot and adjacent rows share nearly all of their data. Checks like these can show an ordering coming apart, which is what happens below the leader. They cannot show one holding up. That would take a second collection.