← ERS STAGE 1 · EVIDENCE BOOK34 occasions · 5 raters / 3 complete

Primary results

PLAIN ANSWER

B leads under the simplest count, but the result changes when people are weighted equally. The pilot therefore does not name a winner.

MOST CHOICES MINUS LEAST CHOICES · 34 COMPLETED RATINGS
A plain
0
B structured scope
+6
C full ERS
-3
D ablation
-3
POSITIVE MEANS MORE “MOST” THAN “LEAST” · RANKING B > A > C=D

The counts below are exact facts about these saved ratings. Turning them into a ranking takes a counting rule. The first rule counts every completed rating once, so people who finished more claims have more influence. It describes this pilot, not experts in general.

The pilot saved 34 completed claim ratings from 5 raters. Three raters finished all ten claims.

Thirty-four is the number of completed claim ratings, not people. Each rating includes one MOST and one LEAST choice.

AUDIT DETAILS · V · 34 judgment occasions; 5 raters; 3 complete · stage1.unit.judgment_occasions

A received 7 MOST and 7 LEAST choices; B received 14 and 8; C received 5 and 8; D received 8 and 11.

AUDIT DETAILS · V · 34 judgment occasions · stage1.primary.raw_counts

When each completed claim rating counts once, this pilot gives B > A > C=D.

This describes these saved ratings only. The counting rule was chosen after collection and does not describe experts in general.

AUDIT DETAILS · D · 34 judgment occasions; 5 raters; 3 complete · stage1.primary.default_ordering

The five raters did not contribute the same amount of data.

EACH PERSON’S CHOICES · ALL 5 RATERS · 3 FINISHED ALL CLAIMS ·2 STOPPED EARLY · 34 COMPLETED RATINGS · WIDER CARD = MORE RATINGS
R05COMPLETE
n10/ 34
MOST
A
3
B
0
C
3
D
4
LEAST
A
3
B
4
C
2
D
1
BESTD
WORSTB
R08COMPLETE
n10/ 34
MOST
A
1
B
8
C
1
D
0
LEAST
A
2
B
2
C
1
D
5
BESTB
WORSTD
R14COMPLETE
n10/ 34
MOST
A
3
B
5
C
0
D
2
LEAST
A
2
B
0
C
3
D
5
BESTB
WORSTD
R11PARTIAL
n3/ 34
MOST
A
0
B
1
C
1
D
1
LEAST
A
0
B
1
C
2
D
0
BESTTIE B=C=D
WORSTC
R06PARTIAL
n1/ 34
MOST
A
0
B
0
C
0
D
1
LEAST
A
0
B
1
C
0
D
0
BESTD
WORSTB
ONE CELL = ONE SELECTION · DASHED CELL = 0 · PARTIAL = DID NOT COMPLETE · BEST/WORST = MODAL PICK, TIES SHOWN AS TIES · A PLAIN · B STRUCTURED SCOPE · C FULL ERS · D ABLATION
A SEPARATE SUBSET, EXCLUDING RECOVERED RECORDS, RUNS ON n 31 OF 34 · NOT ALL RECORDS ARRIVED THE SAME WAY

Three raters finished all ten claims. The other two completed four claims between them. Choosing a counting rule therefore means choosing how much influence complete and partial raters receive. Even the three complete raters do not share one ranking.

Weighting sensitivities do not all agree, and the disagreement is the point.

Using only the three raters who finished all ten claims gives B > A > C > D.

AUDIT DETAILS · D · 30 judgment occasions; 3 complete raters · stage1.sensitivity.complete_panel

Giving each rater the same total weight changes the ranking to D > A > B > C.

Under this rule, a rater who completed one claim gets the same total weight as a rater who completed ten.

AUDIT DETAILS · D · mean of per-rater net rates across 5 raters · stage1.sensitivity.equal_rater

Removing one rater at a time changes the leader: B leads in four checks and D in one.

Removing R08 puts D first. The other four checks leave B first.

AUDIT DETAILS · D · 5 drops · stage1.sensitivity.leave_one_rater

Two of the three complete raters put D in the bottom half. One put D first, so the missing-step check is only partial.

Bottom half means third or fourth after adding MOST choices and subtracting LEAST choices within each rater.

AUDIT DETAILS · D · 3 complete raters · stage1.ablation.complete_raters

R05 completed all ten claims, chose D as MOST four times and LEAST once, and marked all ten MOST choices as specific.

The tags make specificity a useful clue. They do not prove why R05 made the choices or justify removing the record.

AUDIT DETAILS · D · 10 completed claim ratings from R05 · stage1.ablation.r05_specificity_clue

The evidence report shows every required counting view and keeps the failed position check visible.

This checks whether the report is complete. It does not fix the study design or the rater disagreement.

AUDIT DETAILS · V · 5 raters; 4 declared estimand lenses; 5 rater drops; 3 complete-rater pairs · stage1.balance.reporting_gate