A new arXiv paper revisits the WorkBench benchmark for workplace agents two years after earlier tests. The authors report a large jump in task completion and a sharp drop in unintended harmful actions for frontier agents.
The result suggests that capability and safety can improve together on some practical agent tasks. Models that completed more work also tended to cause less unintended damage in the benchmark.
The caution is that serious errors have not disappeared. The paper notes that agents can still make basic mistakes with irreversible consequences, which keeps evaluation and human oversight central for workplace deployment.