321AI
← 321 AI Labs/321 AI Labs Technical Report TR#8

Measuring and Correcting Systematic Bias in LLM-as-Judge Panels

A preregistered audit of position, verbosity, confidence, anchoring, and self-preference bias in three open-weight judges.

Kiyoshi Casey · 321 AI Labs · 2026-09-24

In one paragraph

A preregistered 741-call battery (0 errors) audited three open-weight LLM judges - glm-4-9b-chat, Llama-3.1-8B-Instruct, and qwen3-8b - for five systematic biases. All three judges showed position bias (flip rates 0.17-0.50) and verbosity bias (padded-preference 0.15-0.42) above chance. Llama-3.1-8B-Instruct additionally preferred confident-but-wrong answers (0.075). Anchoring was not significant at n=12. In a 44-pair preregistered self-preference round, glm and qwen over-graded their own answers (+3.53, +4.53 on a 0-10 scale) while Llama under-graded its own (-5.71). We ship the battery, the harness, and a significance-gated correction model, s_corrected = s - coefficient.

0.50

Llama-3.1-8B position flip rate, CI95 [0.34, 0.66]

-5.71

Llama self-under-grading on a 0-10 scale, CI95 [-8.14, -3.14]

741

battery calls executed, 0 errors, 0 scored-class parse failures

What did the audit find?

  • Position bias: all three judges flipped verdicts between AB/BA orders significantly above zero (glm 0.25, Llama 0.50, qwen 0.17).
  • Verbosity bias: all three preferred the padded, incorrect answer above chance (0.15-0.42).
  • Confident-but-wrong: only Llama showed a detectable wrong-answer preference (0.075); glm and qwen at exactly 0.
  • Anchoring: no significant drift for any judge at exploratory n=12 - unresolved, not absent.
  • Self-preference (44 paired blind rounds): glm +3.53 and qwen +4.53 over-graded their own answers; Llama -5.71 under-graded its own. All significant.
  • Correction model: s_corrected = s - coefficient per bias class; coefficients whose 95% CI crosses zero ship significant=false and may not be used to claim a correction.

What does this not claim?

  • Anchoring underpowered at n=12; the null is unresolved, not absent.
  • Judges are 8-9B open-weight models; magnitudes do not transfer to frontier judges.

References

  1. 01Zheng et al. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023. arXiv:2306.05685
  2. 02Wang et al. 2024. Large Language Models are not Fair Evaluators. ACL 2024, pp. 9440-9450. arXiv:2305.17926
  3. 03Panickssery, Bowman, Feng 2024. LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS 2024. arXiv:2404.13076
  4. 04Wataoka, Takahashi, Ri 2024. Self-Preference Bias in LLM-as-a-Judge. NeurIPS SafeGenAI Workshop. arXiv:2410.21819
  5. 05Saito, Wachi, Wataoka, Akimoto 2023. Verbosity Bias in Preference Labeling by Large Language Models. arXiv:2310.10076
  6. 06Gu et al. 2024. A Survey on LLM-as-a-Judge. arXiv:2411.15594

Sealed Artifact

PDF sha256: 9a61e2b916a0cc6fc4207c13ea56acf091425bc56993f8ae42a9c883a1aed5e7

© 2026 321 AI Labs · Measurement and governance research; no capability claims.