A study of multi-agent code evaluation found that an elaborate verification pipeline often had no evidence for choosing between two solutions. When applied unchanged to two coding benchmarks, the MARCH framework rated both candidates equally good in 78% to 95% of comparisons and achieved just 4.4% accuracy in one setting, versus 43.7% for a direct model judgment.

The problem was structural rather than simply a weak model. Verification methods developed for retrieved documents expect evidence that is independent of the answer and differs between candidates. In code judging, both solutions can generate the same apparent evidence, leaving the judge with no grounded basis for a preference even though it still produces confident reasoning.

The researchers derived two warning measurements from the evaluator’s own logs, so they do not require human correctness labels. Using one signal as a gate, the system declined unsupported comparisons and improved accuracy from 20.7% to 36.9% while still answering half of them. That remains below the direct-judgment result and is not a new state-of-the-art judge. Its practical value is calibration: an evaluation system should expose when its evidence cannot distinguish candidates rather than dressing a guess in a detailed explanation.