Contents · 1 / 6
Blinded Judge-Panel Qualification: LLM Judges Sit an Exam Before They Score Anything
Some of what JackBench measures takes judgement, and the judges are themselves language models, so the obvious question is why anyone should trust them. The answer here is an exam. Every judge candidate makes 112 blind predictions against cases whose answers were locked by two independent reviewers before the exam existed, in both candidate orders, repeated. In the first sitting, GPT-5.6 Sol and Gemini 3.1 Pro scored 112/112 and Claude Opus 5 scored 111/112. A later extension put two more candidates through the identical exam: Qwen 3.8 Max and Kimi K3 both scored 112/112 with zero instability, bringing the qualified pool to five judges from five laboratories. The exam also refuses: the studio's own local Qwen 3.8 scored 96.4% and was declined for answer instability. Total exam spend so far: about 2 USD. The takeaway: don't let a language model grade anything until it has passed a blind exam with locked answers, and be suspicious of any judging setup whose exam can't say no. None of this says anything about how good these models are at real work; it qualifies a measuring instrument, and only that.
Why judges need an exam
Some of what JackBench measures can't be decided by an automated check. Whether work subtly missed the point, or built more than anyone asked for, takes judgement, and the judges are themselves language models, which means they come with the same failure modes as the systems they judge.
So a judge has to earn the job. It gets cases where the right answer is already known and locked, and it doesn't get to score anything real until it has shown it can tell those cases apart, blind, repeatedly, in both candidate orders. A judge that can't pass the exam doesn't sit on the panel.
The exam
Fourteen pairwise cases, each independently labelled by two separate reviewers before the exam existed, covering seven kinds of situation a judge must tell apart: clean exact completion, subtle semantic failure, overbuilt but plausible work, truncated work, fabricated verification claims, boundary violations, and cases where the evidence genuinely can't decide.
Each judge saw every case in both candidate orders, under two rubric layouts, twice over: 112 predictions per judge, 336 in total. One case per request, nothing identifying either candidate, answers as strict verdicts with a short reason. No human graded anything; the labels were locked before the first prediction was made.
The result
| Judge | Lab | Correct | Stability |
|---|---|---|---|
| GPT-5.6 Sol (high) | OpenAI | 112/112 | no instability |
| Gemini 3.1 Pro (medium) | 112/112 | no instability | |
| Claude Opus 5 (high) | Anthropic | 111/112 | 0.018, under the 0.05 ceiling |
Provider-pinned seats, no fallbacks. Accepted exam spend: 1.42 USD across 336 predictions.
Opus's single miss is worth telling. On one repetition of a deliberately subtle case, it read changed file contents as evidence against a reference that had in fact been resolved earlier. A defensible reading of a case built to be hard, but the label was locked by two independent reviewers before the exam, and a label doesn't move after seeing a result. The miss stands, and so does the qualification.
The exam also earned its keep before the panel did: preparing it surfaced six real defects in the exam harness itself, from a broken schema to timestamps set in the future, all caught and fixed before they could contaminate a single prediction.
The pool didn't stop at three. A later extension ran the identical, unchanged exam on two more candidates: Qwen 3.8 Max and Kimi K3 both scored 112/112 with no instability of any kind, for 0.60 USD of provider spend, every generation reconciled against the provider's own records. That makes five judges from five laboratories, and at study time a judge never scores work from its own model family. The exam has refused a candidate too: the studio's own local Qwen 3.8 27B reached 96.4% accuracy and was declined, because it changed its answers between repeats too often. Instability is disqualifying no matter whose hardware the judge runs on, including mine.
The rules at study time
Qualification is the licence, and the licence has conditions. At study time the panel only sees cases that automated checks can't decide, and an automated check that contradicts a judge wins outright. Judges see two candidates and a frozen rubric, with everything identifying stripped, in an order set by a sealed random seed. Two aligned votes make a verdict. A tie, a malformed response or an unresolved disagreement is recorded as withheld, never rounded up into a pass.
And the boundary that matters most: this page qualifies a measuring instrument. It licenses no claim about how good Sol, Opus or Gemini are at actual work, any more than calibrating a scale tells you what anything weighs.
Status and citation
The panel is qualified and waiting: applying it to live study output needs its own approved policy first, including how to handle a judge scoring work made by its own model family. Until then, no semantic verdict from this panel appears in any study on this site.
Tyler, J. (2026). Blinded Judge-Panel Qualification. Cold Anvil Studios. https://coldanvil.com/research/blinded-judge-panel-qualification/
@misc{tyler2026judgepanel,
author = {Tyler, Jack},
title = {Blinded Judge-Panel Qualification},
year = {2026},
publisher = {Cold Anvil Studios},
howpublished = {\url{https://coldanvil.com/research/blinded-judge-panel-qualification/}}
}