Vals AI: Agent teams cost 1.8x–5.1x more than solo agents, with one significant gain in four tests
Vals AI tested GPT-6 Sol and Claude Opus 5.5 on Vibe Code Bench as solo agents and as teams. Teams cost 1.8x to 5.1x more, and only one of four comparisons showed a statistically significant improvement. Anthropic's own scaling data and comments from OpenAI's Noam Brown point the same way: more agents mostly buy speed, not quality.
Agent teams cost 1.8x to 5.1x more than single agents on Vibe Code Bench, and only one of four team-versus-solo comparisons produced a statistically significant quality gain, according to evals company Vals AI.
Vals AI results
Vals AI tested GPT-6 Sol and Claude Opus 5.5 on Vibe Code Bench, each as a solo agent and as a team, at two reasoning levels: medium and maximum.
- The single significant result was GPT-6 Sol at medium reasoning, where the team scored 7.3 points higher than the solo agent.
- At maximum reasoning, the team setup gave neither GPT-6 Sol nor Opus 5.5 a real advantage.
- Absolute Vibe Code Bench scores and per-token pricing were not included in the source material.
Vals AI's chart plots cost per app against score, with arrows from each solo configuration to its team counterpart. The arrows run far to the right (higher cost) but only slightly upward.
Anthropic's scaling tests
Anthropic reported diminishing quality returns as it added agents in two of its own tests with Opus 5.5. These figures are self-reported. Larger teams reached a given performance level faster, but going from 10 to 100 agents only nudged scores up slightly after 24 hours.
| Task (Opus 5.5) | 1 agent | 10 agents | 30 agents | 100 agents |
|---|---|---|---|---|
| Knowledge base | 0.53 | 0.70 | 0.71 | 0.74 |
| Lean theorem proving | 0.39 | 0.66 | 0.66 | 0.68 |
Most of the gain arrives by 10 agents. In separate ProgramBench tests, speed gains came with higher token usage.
A model called Fable 5.1 showed stronger quality gains on Lean theorem proving above 10 agents. It still scored below Opus 5.5 across all tests. On the knowledge base task, its score dipped slightly when scaling from 30 to 100 agents. The source does not identify Fable 5.1's developer.
OpenAI voices: speed, not quality
OpenAI researcher Noam Brown said on the Dwarkesh Podcast that multi-agent systems mainly buy speed. By his account, four agents solved tasks twice as fast at twice the cost, and the pattern held at 16 agents with slightly lower efficiency.
Brown said the effect depends on the task. Web research and math parallelize well, while writing a novel does not. He said scaling to very large agent counts remains largely unexplored because the costs are too high.
OpenAI developer Eric Provencher recently warned against agent swarms, arguing they are most likely wasted money because coordination between agents breaks down. He called this the "coordination tax."
What this means
The three sources agree that parallel agents trade tokens for wall-clock time, not for better answers. If the goal is output quality, a single agent at higher reasoning effort looks like the better buy. Vals AI's maximum-reasoning results suggest team setups add little once a model is already running at full compute.
Teams still make sense when latency matters and the task decomposes cleanly, such as research, math, or large parallel code searches. Builders should measure cost per completed task, not just score, because a 1.8x to 5.1x cost multiplier needs a matching quality gain to pay off, and Vals AI found one in only one of four cases.
Two caveats apply. One benchmark with two models is a narrow sample, and the Anthropic figures come from the company's own tests. The open question is how teams perform on tasks built for parallelism, where Brown suggests the economics look different.
Related Articles
Mercor study: AI models beat 12 licensed CPAs on simplified accounting tasks, but top score on full APEX benchmark is 61
A Mercor study found AI models beat 12 licensed CPAs on simplified tasks from the APEX Accounting Benchmark. On the full 160-task benchmark, the top model, Claude Opus 5.5, meets 61.8% of grading criteria, and no model fully solved almost 60% of tasks.
GitHub launches ReviewBench, an open benchmark for AI code review agents built on real pull requests
GitHub has launched ReviewBench, an open benchmark for evaluating AI code review agents. According to GitHub, it uses representative pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. Dataset size, model scores, and licensing details were not included in the announcement summary.
Aleph Alpha benchmark: Chinese AI models balanced on just 17-41% of 967 sensitive-topic prompts
An Aleph Alpha study of 967 politically sensitive prompts found that only 17 to 41 percent of responses from Alibaba's Qwen, DeepSeek and Moonshot's Kimi were balanced, according to the company's own AI scorer. DeepSeek V4 Pro refused roughly two-thirds of questions. Aleph Alpha sells "sovereign AI" to governments, which gives it a commercial interest in the result.
Microsoft's ThinkingBox: Claude Opus 5.5 passes all 20 runs on just 241 of 507 stateful agent tasks
Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the database state they leave behind rather than their responses. Across 507 workflows run 20 times each, Claude Opus 5.5 leads at 67.16% pass@1 but passes all 20 attempts on only 241 tasks.
Comments
Loading...