A new arXiv paper examines a failure mode in long-horizon research agents: optimizing an aggregate metric that hides structural mistakes.
The authors show that when validity depends on disaggregated regions, slices, or cohorts, a single headline number can rank the wrong candidate first. The agent may accept a result that looks better overall while breaking the model underneath.
The lesson is broadly useful for AI-assisted science: agentic search needs discipline around metrics, slices, and verification, not just higher aggregate scores.