A new arXiv paper proposes benchmarking AI agents on scientific challenges that span multiple scales.
That framing is important because scientific work rarely looks like a single prompt with a clean answer. It often requires planning, tool use, evidence gathering, and decisions across different levels of abstraction.
As labs push agents into research workflows, benchmarks need to test whether systems can make useful progress on messy scientific problems, not only score well on isolated tasks.