Teams of AI agents consumed between 1.8 and 5.1 times the cost of single agents in Vals AI testing, but only one of four matched comparisons produced a statistically significant quality gain. The evaluation used GPT-6 Sol and Claude Opus 5.5 on Vibe Code Bench at medium and maximum reasoning settings.

The exception was a GPT-6 Sol team at medium reasoning, which scored 7.3 points above the solo setup. At maximum reasoning, neither model gained a meaningful quality advantage from the team structure, suggesting that parallel workers add less when an individual agent already receives a large compute budget.

Anthropic’s own experiments show a similar plateau. On knowledge-base and theorem-proving tasks, moving from one to ten Opus 5.5 agents improved scores substantially, but increasing the group from ten to 100 added only a few hundredths. ProgramBench tests found that larger groups could finish sooner while consuming more tokens.

Parallelism therefore appears most useful when a task can be divided cleanly and speed matters. OpenAI researcher Noam Brown has said web research and mathematics split more naturally than work such as writing a novel. Teams can buy latency, but coordination and duplicated effort mean they should not be assumed to buy better answers.