A new arXiv paper proposes a way for language models to help reviewers ask a narrower question: do a paper’s methods actually support the novelty claims made in its introduction?
Most automated review tools compare a submission against prior literature. The authors argue that this misses a common human-review concern, where a claimed contribution may be new in wording but not sufficiently realized by the experiments, analysis, or method inside the paper.
Their framework extracts novelty claims, retrieves relevant methodological evidence from the same paper, and generates structured reviewer-style assessments. The evaluation criteria were derived from human peer reviews of 182 ICLR 2025 papers, covering concerns around novelty, methodology, clarity, and related issues.
In tests on accepted and rejected papers, human evaluation found meaningful alignment between the framework’s comments and reviewer concerns, especially on novelty-related issues. The work does not replace peer review, but it shows how LLM tools could help reviewers focus on internal claim support rather than only external similarity.