ERS STAGE 1 · EVIDENCE BOOK24 occasions · 4 raters / 2 complete

NET BEST-WORST · 24 JUDGMENT OCCASIONS
A plain
0
B structured scope
+10
C full ERS
-4
D ablation
-6
MOST − LEAST · ORDERING B > A > C > D

Exact counts are a direct sample fact. The ordering asserted from them is a separate, pilot-restricted claim — the default record-level aggregation is a descriptive target, not a population estimator.

A received 4 MOST and 4 LEAST; B 14 and 4; C 2 and 6; D 4 and 10.

V · 24 judgment occasions · stage1.primary.raw_counts

In this pilot, the default record-level tally ordered the arms B > A > C > D.

Default descriptive aggregation, not a population estimator. Selected after collection.

D · 24 judgment occasions; 4 raters; 2 complete · stage1.primary.default_ordering

The cohort behind those counts is uneven, and the unevenness governs everything after it.

PER-RATER SELECTION PROFILE · ALL 4 RATERS · 2 COMPLETE ·2 PARTIAL · 24 JUDGMENT OCCASIONS · PANEL WIDTH ∝ n · ORDERED BY n
R08COMPLETE
n10/ 24
MOST
A
1
B
8
C
1
D
0
LEAST
A
2
B
2
C
1
D
5
BESTB
WORSTD
R14COMPLETE
n10/ 24
MOST
A
3
B
5
C
0
D
2
LEAST
A
2
B
0
C
3
D
5
BESTB
WORSTD
R11PARTIAL
n3/ 24
MOST
A
0
B
1
C
1
D
1
LEAST
A
0
B
1
C
2
D
0
BESTTIE B=C=D
WORSTC
R06PARTIAL
n1/ 24
MOST
A
0
B
0
C
0
D
1
LEAST
A
0
B
1
C
0
D
0
BESTD
WORSTB
ONE CELL = ONE SELECTION · DASHED CELL = 0 · PARTIAL = DID NOT COMPLETE · BEST/WORST = MODAL PICK, TIES SHOWN AS TIES · A PLAIN · B STRUCTURED SCOPE · C FULL ERS · D ABLATION
A SEPARATE SUBSET, EXCLUDING RECOVERED RECORDS, RUNS ON n 21 OF 24 · NOT ALL RECORDS ARRIVED THE SAME WAY

Two raters reached every claim. The other two contributed a handful of occasions between them, and one of those profiles has no single modal choice at all — its selections tie across three arms. Any weighting rule is therefore a decision about how much those two short profiles ought to count.

Weighting sensitivities do not all agree, and the disagreement is the point.

Restricted to the two complete raters, the ordering is unchanged at B > A > C > D.

D · 20 judgment occasions; 2 raters · stage1.sensitivity.complete_panel

Under equal-rater weighting the ordering reverses to D > B > A > C.

Different unit; not count-comparable. A one-observation rater carries equal weight.

D · mean of per-rater net rates across 4 raters · stage1.sensitivity.equal_rater

B leads under all four leave-one-rater-out drops; the lower-arm ordering is not stable.

Omitting R14 yields B > A=C > D; omitting R08 yields B > A > D > C.

D · 4 drops · stage1.sensitivity.leave_one_rater