A new arXiv paper asks whether AI can evaluate AI systems that generate scientific research. The study focuses on autonomous research generation systems, sometimes described as AI scientists, and explores automated multi-model review as a way to benchmark their output.

The problem is important because the promise of AI-generated research depends on evaluation. Producing a plausible paper is not the same as producing a correct, novel, and useful result. Human peer review is slow and inconsistent, but fully automated review can inherit the weaknesses of the models doing the judging.

The authors propose and implement a benchmarking approach using multiple models as reviewers. That can make review more scalable, but it still leaves open questions about ground truth, novelty, reproducibility, and whether models reward style over substance.

The paper should be read as research infrastructure rather than a verdict that AI scientists are ready. Its main contribution is identifying evaluation as a bottleneck for autonomous scientific discovery.