A preregistered audit of position, verbosity, confidence, anchoring, and self-preference bias in three open-weight judges.
Kiyoshi Casey · 321 AI Labs · 2026-09-24
In one paragraph
A preregistered 741-call battery (0 errors) audited three open-weight LLM judges - glm-4-9b-chat, Llama-3.1-8B-Instruct, and qwen3-8b - for five systematic biases. All three judges showed position bias (flip rates 0.17-0.50) and verbosity bias (padded-preference 0.15-0.42) above chance. Llama-3.1-8B-Instruct additionally preferred confident-but-wrong answers (0.075). Anchoring was not significant at n=12. In a 44-pair preregistered self-preference round, glm and qwen over-graded their own answers (+3.53, +4.53 on a 0-10 scale) while Llama under-graded its own (-5.71). We ship the battery, the harness, and a significance-gated correction model, s_corrected = s - coefficient.
0.50
Llama-3.1-8B position flip rate, CI95 [0.34, 0.66]
-5.71
Llama self-under-grading on a 0-10 scale, CI95 [-8.14, -3.14]
741
battery calls executed, 0 errors, 0 scored-class parse failures
Sealed Artifact
PDF sha256: 9a61e2b916a0cc6fc4207c13ea56acf091425bc56993f8ae42a9c883a1aed5e7
© 2026 321 AI Labs · Measurement and governance research; no capability claims.