Exactly two numbers usually make it into a report on the quality of automated evaluation: how closely a model's verdicts match human ones, and how robust they are to minor changes like reordering options or rewording the prompt. A new paper argues that this is not enough — and proposes looking at the judge differently.
Agreement answers the wrong question
Agreement and robustness to superficial perturbations are about reliability. A reliable measuring instrument produces a stable result, but stability alone says nothing about whether it measures the right quantity. A thermometer that always reads 20 degrees is exceptionally reliable and completely useless.

The authors of "A Judge Should Know What Changed: Construct Validity for LLM-as-a-Judge Evaluation" (Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong; arXiv:2608.24419, cs.AI, submitted August 25, 2026; 39 pages, 10 figures, 11 tables) move the discussion into the realm of construct validity: does the evaluation correspond to the concept it claims to measure? For an LLM judge, this means a simple thing — the judge should respond to meaning, not to form.
A two-dimensional profile
Instead of a single metric, a profile of two probabilities is proposed.
S — invariance. How often the verdict stays the same when an edit preserves meaning: paragraphs reordered, style changed, filler removed. A good judge should not flinch here.
R — construct sensitivity. How often the verdict changes when an edit is minimal but meaning is affected: a caveat added, a claim weakened, scope narrowed. A good judge must react here.
The paper's key claim: S and R are independent, and no scalar summary preserves all relevant comparisons. Put simply, you cannot average two numbers and get one honest one — a judge with high S and low R and a judge with low S and high R can produce the same "average score" while remaining fundamentally different instruments.
How it was tested
The experimental design is noticeably stricter than is usual for such work.
- 7 judges and 4 domains — that is, not a single model and not a single task type were tested.
- 7 types of construct-changing interventions versus 5 controls that touch only register and do not change meaning.
- The direction of the edit — whether it changes the construct or not — was determined by human annotators, not by default labeling.
- Generation, verification, and judging were assigned to non-overlapping model families. This is an important detail: the judge does not evaluate text it produced itself, nor does it check edits it invented.
The last point deserves special attention. When the generator, verifier, and judge are the same model, high agreement may be an artifact of shared biases rather than a sign of quality.
Invariance is nearly perfect, sensitivity is not
Here is where it gets most interesting. At a threshold of consistent invariance S ≥ 0.90, judges on average showed S = 0.945. The figure that reporting usually aims for.
Sensitivity, meanwhile, is R = 0.319. A rough interpretation: in about two out of three cases, the judge failed to notice an edit that changes meaning. A formally impeccable result on invariance coexists with a very weak response to content.

Scale and strength are not the same thing
Weak sensitivity is distributed unevenly. When an edit changes the scale of a claim, the response is higher: R_scope = 0.383. When it changes strength, it is weaker: R_strength = 0.262. The gap is +0.121, and its sign is the same for all seven judges.
So this is not random variation between models but a systematic feature. Judges read widening and narrowing of scope better than shifts along the confidence scale — "probably" versus "definitely," "may help" versus "helps." For practice, this is unpleasant news: it is precisely such gradations that most often need to be distinguished.
Human labels are worth checking too
A separate storyline is the five public label sets used as ground truth. Predictors relying solely on superficial text features reproduce 55–67% of the labels in pairwise mode. In the case of MT-Bench, a superficial heuristic matched 67.4% of human votes.
The conclusion is an uncomfortable one. If a simple rule that understands nothing about content reproduces two-thirds of human decisions, then such labels poorly separate meaning from form — and it is not only the judge that needs validating, but also the very set it is tested on.

What to do about it
The authors do not propose replacing LLM judges with something else — they propose changing reporting and the validation procedure.
Joint reporting. S and R are published side by side rather than collapsed into a single number. The reader of the report immediately sees where the judge falls short.
Auditing the validation set. Before trusting agreement with human labels, it is worth checking how predictable those labels are without understanding the text. If they are highly predictable, trust in them is inflated.
Separation of roles. Generation, verification, and judging by different model families reduce the risk that high invariance is a consequence of shared blind spots.
The main takeaway
High agreement is about stability, not correctness. A judge that delivers its verdict with equal confidence both before and after a meaning-changing edit is not "robust" but inattentive. As long as reports contain only a single number, these two cases cannot be told apart — and that is exactly why a two-dimensional profile looks not like a complication but like the minimum necessary honesty.



