A new cs.AI paper, “Life After Benchmark Saturation,” argues that retiring a benchmark once accuracy saturates can throw away useful measurement opportunities. The authors use CORE-Bench as a case study.

The central point is that accuracy is only one dimension of agent performance. Even when systems score highly, researchers can still study shortcuts, out-of-distribution behavior, efficiency, reliability, and other traits that matter in deployment.

The paper is timely because AI evaluation is under pressure from rapid benchmark saturation, contamination concerns, and growing interest in agent systems that need richer measurement than pass-or-fail scores.