A new ICML position paper argues that AI evaluation should pay much more attention to how systems work with people, not only whether they can beat humans acting alone.

The authors say today's dominant benchmark culture often rewards superhuman autonomous performance. In their view, that implicitly steers development toward replacing human work rather than complementing it. They propose shifting more evaluation toward human-AI teams, where the relevant question is whether the combined system produces better outcomes than either side would alone.

The argument matters because benchmarks shape product goals, research funding, and public claims about progress. A model that performs well in isolation may still fail to help a doctor, teacher, analyst, or engineer make better decisions under real constraints. The paper is a position piece rather than a new benchmark suite, but it captures a growing concern: measuring autonomy alone can miss the practical value and risks of AI in actual workplaces.