Brandon Gottshall
A small pilot asking whether experts can give one dependable ranking of question sets they cannot trace back to a method.
This Stage 1 pilot is coursework for DATA 4990 at Valdosta State University — a self-guided directed study, not a taught lecture course.
Six weeks of design, build, collection, and analysis — compressed into one summer term.
The student owns the protocol, the software, the recruitment, and the readout. Faculty supervision covers the course; it does not co-author the work.
Expert raters were solicited while other professors were between teaching terms — when a short review was more possible than during a full semester.
COURSE CONTAINER · NOT A GRANT · NOT A LAB PRODUCT · A SIX-WEEK PILOT UNDER DATA 4990
Stage 1 asks a simple question: can this rating task give us one dependable ranking? The answer is no — not yet.
Research programs turn on which questions get treated as worth investigating. Stage 1 asks whether that judgement can be written down and measured at all. It is a design diagnosis — a check of the measurement setup — not a verdict on which method is better.
The Epistemic Reasoning Substrate is a research project on representing research questions so experts can judge them without seeing which method wrote them. Later beats compare four ways of writing those questions. One is the full ERS method — that is why "full ERS" appears as arm C. The project is the line of work; arm C is one thing it produced.
Fields decide what to study by deciding which questions matter. Right now that choice is made informally — by taste, by seniority, by who is in the room — and it leaves no record anyone can examine. If it can be written down and measured, it becomes something a department can teach, hand on, and check. This is a first test of whether it can.
If measurement works, later work could investigate links along this path. Treat each arrow as its own research program. This pilot does not establish any arrow.
Papers in. Better questions out — disagreement still attached.
NO VERDICT · PROS, CONS, AND NUANCE TRAVEL WITH EACH QUESTION
ERS turns explicit gaps into structured question groups on a provenance-aware graph—not a tree. Models propose; validators, review, and transition rules decide what is committed. The three systems are coupled.
PROPOSAL · NOT ESTABLISHED BY STAGE 1
committed MEANS A RECORD PASSED PROCESS GATES · IT DOES NOT MEAN true
After validated questions exist, experts contribute informed taste about which are most worth pursuing — the component Stage 1 first attempted. What it produced was a repair plan, not a dependable ranking.
The pilot saved 34 completed claim ratings from 5 raters. Three raters finished all ten claims.
Thirty-four is the number of completed claim ratings, not people. Each rating includes one MOST and one LEAST choice.
Five raters made 34 completed choices. Counting each completed choice once gave B > A > C=D.
The pilot did not find one dependable ranking.
Do not choose a winning method or say this rating task has been proven to work.
Balance screen positions. Choose the counting rule and pass/fail checks before collecting data. Test clear scope and causal reasoning as separate parts.
When each completed claim rating counts once, this pilot gives B > A > C=D.
This describes these saved ratings only. The counting rule was chosen after collection and does not describe experts in general.
The crossing below is real for this full sample. It is not robust evidence of instability on its own: among the 3 raters who finished every claim, both weightings keep the same order — and that check does not justify dropping anyone. What holds separately: these raters did not share one dependable ordering, and removing one complete rater can change which arm sits first.
B structured scope sits first when each judgment occasion counts once and third when each rater counts once. D ablation sits third when each judgment occasion counts once and first when each rater counts once. Same records, same selections — only the weighting changes.
Low confidence before and after, for different reasons. Before: this small pilot already had no dependable ranking. After: still low — same uneven ratings, counted so each person weighs the same.
The screen-position check failed. We added this check after collection, without using the choices to set its rule. It was not planned in advance. The check warns us about the design; it does not fix it or produce a corrected ranking.
Assignment support fails; no confirmatory position-adjusted arm ordering may be promoted. Preserve raw and exploratory summaries.
ADDED post collection · PLANNED IN ADVANCE FALSE · MAXIMUM WITHIN-ARM EXPOSURE SPREAD 9
The saved screen assignments fail the position-balance check for the full pilot and for the complete raters.
This check was added after collection and was not planned in advance. Failure blocks a corrected ranking but does not make the saved choices invalid.
R05 chose D as MOST 4 times and LEAST 1 time. R05 marked 10 of 10 MOST choices as specific. That is a clue, not proof of why R05 chose them.
Removing one rater at a time changes the leader: B leads in four checks and D in one.
Removing R08 puts D first. The other four checks leave B first.
Generated and hash-locked before a single rating was collected.
The arm decode key is substrate-only. It never reaches a rater-facing surface.
Raters saw four unlabeled sets. Nothing on screen named the method.
Written once, never edited. One per judgment occasion, plus 7 session objects.
Every input re-hashed and compared on each build. Drift fails the build.
KEYS LOCKED BEFORE COLLECTION · ANALYSIS RULES CHOSEN AFTER COLLECTION · MANIFEST 9cee29bfc6d0ee8dc9df7ecb8e90196c8fa884e396dcfe83029fc0e2a069b174
The pilot saved 34 completed claim ratings from 5 raters. Three raters finished all ten claims.
Thirty-four is the number of completed claim ratings, not people. Each rating includes one MOST and one LEAST choice.
Stage 1 produced records we can inspect, but it did not find one dependable ranking. B led when each completed choice counted once. D led when each rater received the same total weight. The leader also changed when one rater was removed, and screen positions were not balanced. The missing-step check was only partial. This result gives us a repair plan for Stage 2. It does not name a winning method, show what experts in general prefer, or prove that the rating task works.
When each completed choice counts once, the order is B > A > C=D. This describes this small pilot only.
When each rater gets the same total weight, the order changes to D > A > B > C. Removing one rater at a time changes the leader in 1 of 5 checks.
The task produced records we can inspect.
The raters did not share one dependable ranking.
The next study needs balanced screens and rules chosen in advance.
A winning method.
What experts in general prefer.
Proof that the rating task works.
Why the observed ranking happened.
A claim that screen position caused the disagreement.
NEXT-STUDY RULES NOT YET FINAL
The proposed Stage 2 tests clear scope and causal reasoning as separate parts in four combinations.
This is a proposed design. Its rules and sample size are not final.
Stage 1 does not say which method is better; it establishes what a study must control in order to say. The question we can now ask — does supplying structure change which questions a field treats as worth investigating? — is one nobody can answer today.