Researchers have developed an adaptive evidence router for language models that judge pairs of answers. Instead of forcing every judge to use the same reasoning protocol, BAER selects among three symmetric methods according to the benchmark and underlying judge model.
The methods combine evidence directly, route among experts by estimated reliability or verify against a reference without looking at candidate identity. Symmetry is important: swapping answer A and answer B may reverse which one wins, but it should not change the strength of the judgment. Development data select one method for each benchmark-model pairing, and that selection is frozen before testing.
Across four benchmarks and two eight-billion-parameter judge models, BAER achieved the highest reported test accuracy in all eight conditions. Gains over the strongest external baseline ranged from 0.87 to 7.32 percentage points, with predictions on every example. The study is limited to two backbones and four benchmarks, so the method still needs testing with larger judges and different evaluation tasks before being treated as universally reliable.