Evaluation, Benchmarking & Observability · 7 min
LLM-as-a-Judge as a Measurement Instrument
Using a model to grade a model is fine — as long as you treat the judge like any instrument: calibrate it against human labels, probe its biases, version it into every number it produces, and re-check it when it drifts.
A grading harness once told a team their new model beat GPT-4 on 62% of prompts. They drafted the announcement. Then someone swapped the order the two answers were shown to the judge and re-ran it. Now their model won 38% of the time. Same answers, same judge, same rubric — the only thing that changed was which response came first. The judge wasn't measuring quality. It was measuring position, and quality was noise on top.
That is the whole problem with LLM-as-a-judge in one story. Using a model to grade a model is a legitimate, well-studied technique. It is also an instrument, and an uncalibrated instrument that reads 62 or 38 depending on how you hold it is not telling you about the world. It's telling you about itself.
Two shapes of judge, and why rubrics beat vibes
There are two dominant patterns. In pairwise comparison you show the judge a prompt and two answers, A and B, and ask which is better (or if they tie). In rubric scoring (also called single-answer grading or reference-guided grading) you show one answer and ask for a score against explicit criteria, often 1–10 or a set of yes/no checks.
Pairwise is more reliable when you're ranking systems, because relative judgments are easier than absolute ones — for a model and for a person. Rubric scoring is what you want when you need a standalone number per response, for a dashboard or a regression gate.
Either way, the rubric is the instrument's calibration. Vague criteria produce vague, unstable readings. "Rate the helpfulness from 1 to 10" invites the judge to average a fog. Decompose it:
Score each dimension 0 or 1, then sum: [ ] Directly answers the question that was asked [ ] Every factual claim is supported by the provided context [ ] No instruction from the prompt is ignored [ ] Code (if any) runs without modification [ ] Stays under the requested length
Binary, observable checks push the judge toward reading the text instead of emitting a mood. Zheng et al., who built MT-Bench and ran the large-scale study behind this technique, found that GPT-4's agreement with human preferences reaches roughly 85% under a good pairwise setup — higher than the agreement between two independent humans (about 81%). That is the encouraging headline. The rest of their paper is the fine print, and the fine print is where you live.
Position bias, shown flipping a verdict
The most documented failure is position bias: the judge systematically prefers whichever answer sits in a particular slot. It is not subtle. Feed a judge two answers, then feed it the same two swapped, and count how often the winner changes. Here is the check, and it is the single most valuable thing in this lesson:
def position_bias_rate(judge, prompt, answers):
flips = 0
for a, b in answers: # a, b are two candidate responses
v1 = judge(prompt, first=a, second=b) # "A", "B", or "tie"
v2 = judge(prompt, first=b, second=a) # order swapped
# normalize v2 back to A/B identity
v2 = {"A": "B", "B": "A", "tie": "tie"}[v2]
if v1 != v2 and "tie" not in (v1, v2):
flips += 1
return flips / len(answers)Concretely: the prompt is "Explain why quicksort is O(n log n) on average." Answer X is a crisp, correct three-sentence explanation. Answer Y is longer, correct, and slightly padded. Ask the judge "which is better, the first or the second?" With X first, GPT-3.5-class judges often say "first." Swap them so Y is first, and the same judge says "first" again — it just reversed its verdict on X versus Y without a word of the answers changing. Wang et al. measured exactly this and showed the ranking of candidates can be flipped by reordering them alone; they named the effect and proposed calibration fixes. Zheng et al. saw it strongly in weaker judges and still measurably in GPT-4.
The mitigation is cheap and non-negotiable: run every pair in both orders and only count a win when the verdict survives the swap. Disagreements become ties. Your position_bias_rate above is now a live diagnostic — if it's above ~5%, your judge or rubric is too weak to trust for that task, and no amount of averaging fixes a biased instrument.
The biases that don't announce themselves
Position bias is easy to catch because it has an obvious symmetry to test. The others hide.
Length / verbosity bias. Judges tend to reward longer answers even when the extra length adds nothing, and can be talked into preferring a verbose wrong answer over a terse right one. Test for it: take a set of answers, produce a padded variant of each that adds no real information, and check whether the scores rise for the longer version. If they do, your rubric needs an explicit "does not pad" check and a hard length cap.
Self-preference / judge-model dependence. A judge tends to favor text that looks like its own output. If GPT-4 is your judge and one of the systems you're comparing is also GPT-4, that system gets a quiet thumb on the scale. This is why the judge model is part of your result, not a neutral detail. "System A scored 7.8" is meaningless; "System A scored 7.8 as graded by GPT-4o with rubric v3" is a measurement. Change the judge and you have to recalibrate.
Reference-answer dependence. For anything with a right answer — math, factual QA, "did it follow the schema" — give the judge a reference (gold answer or the source context) and ask it to grade against that, not against its own belief. Reference-guided grading sharply cut the judge's error rate on math and reasoning questions in the MT-Bench study, because the judge stops re-deriving the answer badly and starts checking. But now the reference is part of the instrument: a wrong gold answer produces confidently wrong grades at scale.
Calibrate against humans, or you're guessing
Here is the line that separates an instrument from a vibe: you do not trust a judge you have not checked against human labels. Take 100–200 examples, have humans label them (pairwise preferences or rubric scores), then run your judge on the same set and measure agreement.
Use a real agreement statistic, not raw percent. For categorical verdicts, Cohen's kappa corrects for the agreement you'd get by chance:
kappa = (p_observed - p_chance) / (1 - p_chance) kappa < 0.2 poor — the judge is barely better than a coin 0.2 – 0.4 fair 0.4 – 0.6 moderate — usable for coarse gating, not fine ranking 0.6 – 0.8 substantial — trustworthy for most eval work > 0.8 near-human
If your judge lands at kappa 0.3, you have not built an eval; you have built a random number generator with a system prompt. Fix the rubric, add a reference, or pick a stronger judge model, then re-measure. And keep the human set: it's your calibration standard, the thing you re-run whenever you change the judge model, the prompt, or the provider silently updates the model under you. Judges drift exactly like any hosted dependency drifts.
Disagreement is signal, not an error to suppress
When the judge and your humans disagree, resist the urge to "fix" the judge until it matches. Read the disagreements. Often they're telling you the rubric is ambiguous — two reasonable readers, one human and one model, split because the criterion was underspecified. That's a rubric bug, and fixing it improves the instrument for everyone.
When two judges disagree with each other, treat it like conflicting sensors. A cheap ensemble — grade with two different judge models, keep the cases where they agree, and route the rest to a human — gives you a high-precision automatic signal plus a bounded human queue. LMSYS's Chatbot Arena takes the complementary route: instead of trusting one judge, it aggregates hundreds of thousands of pairwise human votes into Elo-style ratings, and leans on a model-judge only where it has been shown to track those humans.
None of this is "ask another LLM if the first one was right." It's metrology. You have an instrument with known, measurable biases. You test each bias with a specific probe, you calibrate against a human standard with a real statistic, you version the judge as part of every number it produces, and you re-run the calibration when anything upstream changes. Do that, and a model grading a model is a genuine measurement. Skip it, and you'll ship the 62% press release right before someone swaps the order.
Sources
- Zheng, L., et al. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." NeurIPS 2023. https://arxiv.org/abs/2306.05685
- Wang, P., et al. "Large Language Models are not Fair Evaluators." 2023. https://arxiv.org/abs/2305.17926
- Liu, Y., et al. "G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment." EMNLP 2023. https://arxiv.org/abs/2303.16634
- Kim, S., et al. "Prometheus: Inducing Fine-grained Evaluation Capability in Language Models." 2023. https://arxiv.org/abs/2310.08491
- Cohen, J. "A Coefficient of Agreement for Nominal Scales." Educational and Psychological Measurement, 1960. https://doi.org/10.1177/001316446002000104
- Chiang, W.-L., et al. "Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference." 2024. https://arxiv.org/abs/2403.04132