The study reviewed seven benchmarks and found that agent success rates can rise substantially when token budgets are increased. On software engineering tasks, performance reportedly jumped when the budget was expanded tenfold.

The result is important for safety and evaluation. If benchmarks cap resources too tightly, they may understate what deployed agents can do when given more time, tools, and compute.