An Apple-affiliated research paper argues that elaborate multi-agent systems may add little value on current autonomous machine-learning engineering benchmarks. The researchers compared leading open-source harnesses with a minimal coding agent that could directly read files, write code and run shell commands.
Under the same time budget and using the same frontier language model, the more complex systems did not outperform a single minimal-harness session. Systematic ablation tests also suggested that extra orchestration layers, including specialized retrieval agents and multi-agent coordination, became redundant once the underlying coding model was strong enough.
The result points to the model backbone, rather than hand-built orchestration, as the main performance driver in the tested setting. That has practical implications for teams deciding whether to spend engineering time on elaborate agent architectures or first improve the model, tools and execution environment available to one agent.
The finding should not be generalized to every agent workload. It covers autonomous machine-learning engineering tasks and the benchmarks used in the study, where goals and execution environments are relatively structured. Long-lived business processes, collaboration among specialists or tasks requiring distinct permissions could still benefit from additional orchestration. The paper’s narrower claim is that current benchmark gains do not by themselves justify increasingly complicated harnesses.