ERS STAGE 1 · Epistemic Reasoning Substrate

Can expert taste be measured?

Brandon Gottshall

A small pilot asking whether experts can give one dependable ranking of question sets they cannot trace back to a method.

DATA 4990 · a self-guided study

This Stage 1 pilot is coursework for DATA 4990 at Valdosta State University — a self-guided directed study, not a taught lecture course.

WHEN
Summer short session 2

Six weeks of design, build, collection, and analysis — compressed into one summer term.

WHAT IT IS
Self-guided study

The student owns the protocol, the software, the recruitment, and the readout. Faculty supervision covers the course; it does not co-author the work.

WHO WAS ASKED
Faculty on break

Expert raters were solicited while other professors were between teaching terms — when a short review was more possible than during a full semester.

COURSE CONTAINER · NOT A GRANT · NOT A LAB PRODUCT · A SIX-WEEK PILOT UNDER DATA 4990

What this is

Stage 1 asks a simple question: can this rating task give us one dependable ranking? The answer is no — not yet.

Research programs turn on which questions get treated as worth investigating. Stage 1 asks whether that judgement can be written down and measured at all. It is a design diagnosis — a check of the measurement setup — not a verdict on which method is better.

WHAT ERS IS

The Epistemic Reasoning Substrate is a research project on representing research questions so experts can judge them without seeing which method wrote them. Later beats compare four ways of writing those questions. One is the full ERS method — that is why "full ERS" appears as arm C. The project is the line of work; arm C is one thing it produced.

THE OPEN PROBLEM

Fields decide what to study by deciding which questions matter. Right now that choice is made informally — by taste, by seniority, by who is in the room — and it leaves no record anyone can examine. If it can be written down and measured, it becomes something a department can teach, hand on, and check. This is a first test of whether it can.

A MAP OF SEPARATE PROGRAMS · NOT A RESULT

If measurement works, later work could investigate links along this path. Treat each arrow as its own research program. This pilot does not establish any arrow.

better question representationadmissible expert judgmentuseful expert leveragechanged research workflowcumulative knowledge effectsinstitutional use

What it does with a literature

The three-system architecture

ERS turns explicit gaps into structured question groups on a provenance-aware graph—not a tree. Models propose; validators, review, and transition rules decide what is committed. The three systems are coupled.

PROPOSAL · NOT ESTABLISHED BY STAGE 1

CANONICAL EPISTEMIC GRAPH
stores state
RESOLUTION POLICY GRAPH
admissible moves
SUPERVISOR AGENT
without owning truth

committed MEANS A RECORD PASSED PROCESS GATES · IT DOES NOT MEAN true

After validated questions exist, experts contribute informed taste about which are most worth pursuing — the component Stage 1 first attempted. What it produced was a repair plan, not a dependable ranking.

Four pilot judgments on this claim

PILOT · THIS CLAIM · 4 JUDGMENT OCCASIONS
3/4
CHOSE THE SAME SET AS MOST WORTH INVESTIGATING
3/4
CHOSE THE SAME SET AS LEAST
Show the evidence unit
DATA: VERIFIED · CLAIM: DIRECT FACT

The pilot saved 34 completed claim ratings from 5 raters. Three raters finished all ten claims.

Thirty-four is the number of completed claim ratings, not people. Each rating includes one MOST and one LEAST choice.

34 judgment occasions; 5 raters; 3 complete · DEPENDENCIES S3 · outputs/stage1_results.json#unit

The pilot found disagreement, not a winner

WHAT WE SAW

Five raters made 34 completed choices. Counting each completed choice once gave B > A > C=D.

BOTTOM LINE

The pilot did not find one dependable ranking.

DECISION

Do not choose a winning method or say this rating task has been proven to work.

WHAT TO FIX

Balance screen positions. Choose the counting rule and pass/fail checks before collecting data. Test clear scope and causal reasoning as separate parts.

The simple count puts B first

MOST CHOICES MINUS LEAST CHOICES · 34 COMPLETED RATINGS
A plain
0
B structured scope
+6
C full ERS
-3
D ablation
-3
POSITIVE MEANS MORE “MOST” THAN “LEAST” · RANKING B > A > C=D
Show weighting and sensitivities
DATA: VERIFIED · CLAIM: DESCRIPTIVE

When each completed claim rating counts once, this pilot gives B > A > C=D.

This describes these saved ratings only. The counting rule was chosen after collection and does not describe experts in general.

34 judgment occasions; 5 raters; 3 complete · DEPENDENCIES S1, S2, S3 · outputs/stage1_results.json#primary.ordering

Change what gets one vote, and the ranking changes

The crossing below is real for this full sample. It is not robust evidence of instability on its own: among the 3 raters who finished every claim, both weightings keep the same order — and that check does not justify dropping anyone. What holds separately: these raters did not share one dependable ordering, and removing one complete rater can change which arm sits first.

WHAT GETS ONE VOTE? · 34 COMPLETED RATINGS ·5 RATERS · 10 CLAIMS
Each completed rating gets one vote
MOST MINUS LEAST
34 completed ratings; frequent raters count more
B > A > C=D
Each claim gets one vote
AVERAGE WITHIN CLAIM
10 claims receive equal weight
B > A > C > D
Each person gets one vote
AVERAGE WITHIN PERSON
5 raters receive equal total weight
D > A > B > C
1ST2ND3RD4TH
A plain0
A+0.013
0.000A plain
B structured scope+6
B+0.202
-0.060B structured scope
C full ERS-3
C-0.072
-0.107C full ERS
D ablation-3
D-0.143
+0.167D ablation

B structured scope sits first when each judgment occasion counts once and third when each rater counts once. D ablation sits third when each judgment occasion counts once and first when each rater counts once. Same records, same selections — only the weighting changes.

Low confidence before and after, for different reasons. Before: this small pilot already had no dependable ranking. After: still low — same uneven ratings, counted so each person weighs the same.

COMPARE RANKS, NOT THE RAW NUMBERS · EQUAL VALUES SHARE A RANK ·Under this rule, a rater who completed one claim gets the same total weight as a rater who completed ten. · PER PERSON R05 n=10 · R06 n=1 · R08 n=10 · R11 n=3 · R14 n=10 · 3/5 FINISHED ALL 10

Screen position and question content are mixed together

HOW OFTEN EACH METHOD APPEARED IN EACH SCREEN POSITION
POSITION 1POSITION 2POSITION 3POSITION 4A10897B910123C107512D59812
SHOWN EQUALLY ACROSS POSITIONS: FALSE · CAN SEPARATE POSITION FROM CHOICE: FALSE
SCREEN-BALANCE CHECK · FAILED

The screen-position check failed. We added this check after collection, without using the choices to set its rule. It was not planned in advance. The check warns us about the design; it does not fix it or produce a corrected ranking.

Assignment support fails; no confirmatory position-adjusted arm ordering may be promoted. Preserve raw and exploratory summaries.

ADDED post collection · PLANNED IN ADVANCE FALSE · MAXIMUM WITHIN-ARM EXPOSURE SPREAD 9

Show the balance-gate contract
DATA: VERIFIED · CLAIM: DIRECT FACT

The saved screen assignments fail the position-balance check for the full pilot and for the complete raters.

This check was added after collection and was not planned in advance. Failure blocks a corrected ranking but does not make the saved choices invalid.

34 judgment occasions; 16 aggregate arm-position cells; 3 complete raters · DEPENDENCIES S1, S2, S3 · outputs/balance_gate.json#position_balance_gate

The ablation check does not hold for everyone

D REMOVED THE FINAL “DOES THIS TEST THE CLAIM?” CHECK · 3 COMPLETE RATERS ·2 PUT D IN THE BOTTOM HALF · 1 PUT D FIRST
R05 CLUE

R05 chose D as MOST 4 times and LEAST 1 time. R05 marked 10 of 10 MOST choices as specific. That is a clue, not proof of why R05 chose them.

A plainB structured scopeC full ERSD ablationR05 · 10 claims
R08 · 10 claims
R14 · 10 claims
Show per-rater sensitivity
DATA: VERIFIED · CLAIM: DESCRIPTIVE

Removing one rater at a time changes the leader: B leads in four checks and D in one.

Removing R08 puts D first. The other four checks leave B first.

5 drops · DEPENDENCIES S1, S2, S3 · outputs/tables/leave_one_rater_out.csv

Every number traces back to locked source files

1
QUESTION BANK

Generated and hash-locked before a single rating was collected.

2
KEYS, SEALED

The arm decode key is substrate-only. It never reaches a rater-facing surface.

4
SETS PER CLAIM

Raters saw four unlabeled sets. Nothing on screen named the method.

34
RATING RECORDS

Written once, never edited. One per judgment occasion, plus 7 session objects.

44
ONE MANIFEST

Every input re-hashed and compared on each build. Drift fails the build.

KEYS LOCKED BEFORE COLLECTION · ANALYSIS RULES CHOSEN AFTER COLLECTION · MANIFEST 9cee29bfc6d0ee8dc9df7ecb8e90196c8fa884e396dcfe83029fc0e2a069b174

Show the unit of analysis
DATA: VERIFIED · CLAIM: DIRECT FACT

The pilot saved 34 completed claim ratings from 5 raters. Three raters finished all ten claims.

Thirty-four is the number of completed claim ratings, not people. Each rating includes one MOST and one LEAST choice.

34 judgment occasions; 5 raters; 3 complete · DEPENDENCIES S3 · outputs/stage1_results.json#unit

What the pilot supports

Stage 1 produced records we can inspect, but it did not find one dependable ranking. B led when each completed choice counted once. D led when each rater received the same total weight. The leader also changed when one rater was removed, and screen positions were not balanced. The missing-step check was only partial. This result gives us a repair plan for Stage 2. It does not name a winning method, show what experts in general prefer, or prove that the rating task works.

WHAT THESE RATINGS SHOW

When each completed choice counts once, the order is B > A > C=D. This describes this small pilot only.

When each rater gets the same total weight, the order changes to D > A > B > C. Removing one rater at a time changes the leader in 1 of 5 checks.

SAFE TO SAY

The task produced records we can inspect.

The raters did not share one dependable ranking.

The next study needs balanced screens and rules chosen in advance.

NOT SUPPORTED

A winning method.

What experts in general prefer.

Proof that the rating task works.

Why the observed ranking happened.

A claim that screen position caused the disagreement.

The next study tests the missing parts separately

SCOPE ABSENT
SCOPE SUPPLIED
CAUSAL REASONING ABSENT
S-C-
no scope · no causal
S+C-
scope only
THE ARM THAT LED
CAUSAL REASONING PRESENT
S-C+
causal only
STAGE 1 FULL METHOD
S+C+
scope + causal

NEXT-STUDY RULES NOT YET FINAL

Show the design contract
DATA: NOT-APPLICABLE · CLAIM: PROTOCOL

The proposed Stage 2 tests clear scope and causal reasoning as separate parts in four combinations.

This is a proposed design. Its rules and sample size are not final.

· DEPENDENCIES

Stage 1 does not say which method is better; it establishes what a study must control in order to say. The question we can now ask — does supplying structure change which questions a field treats as worth investigating? — is one nobody can answer today.

PRESENTER · NEEDS YOUR LINK34 occasions · 5 raters / 3 complete