Measuring and Correcting Systematic Bias in LLM-as-Judge Panels
A preregistered 741-call battery (0 errors) audited three open-weight LLM judges for five systematic biases. All three showed position and verbosity bias above chance; glm and qwen over-graded their own answers while Llama under-graded its own. Ships with a significance-gated correction model.

