A production text-to-SQL team found that its automated evaluator was performing far worse than expected when compared with human judgments. On a disagreement-heavy test set, the deployed GPT-4o-mini judge achieved a Cohen’s kappa of just 0.04, a measure of agreement beyond chance, and incorrectly flagged 77.1% of examples humans considered faithful.

The audit traced most false alarms to one behavior the researchers call “grade hallucination.” Replacing the judge with a self-hosted Qwen3.6-27B model raised kappa to 0.72, close to the 0.71 measured for Claude Opus 4.7 in a small 96-example comparison. The paper estimates Qwen’s per-call cost at roughly one three-hundredth of Claude’s.

Simply combining judges did not reliably fix the problem: pairing a weak judge with a strong one reduced agreement. Three strong judges using unanimous decisions reached 0.79 kappa while automatically covering 89.7% of cases. The exact figures come from one pipeline and some comparisons are underpowered, but the lesson is broad: teams should measure an LLM judge against human labels before treating its score as ground truth. The authors released code and a preregistration for the audit.