Vals AI Finds Multi-Agent Coding Teams Cost 5x More but Rarely Score Better
Vals AI pitted GPT 6 Sol and Claude Opus 5.5 against themselves on Vibe Code Bench, running each as a solo coder and as a multi-agent team.
- Vals AI ran GPT 6 Sol and Claude Opus 5.5 on Vibe Code Bench as solo agents and teams.
- Teams cost 1.8x to 5.1x more; only Sol at medium effort showed a statistically significant score gain (+7.3 points).
- Opus max-effort team hit the top score of 93.2% but cost $122 and took 174 minutes per app.
- For Sol, raising reasoning effort beat adding a team: +11.4 points solo vs +7.3 from a team.
- Sol delegates architecturally in parallel; Opus writes a CONTRACT.md and dispatches sequential waves of builders and testers.
- Full writeup with timelines and methodology: vals.ai blog.
Multi-agent coding raises costs faster than scores
OpenAI and Anthropic now provide native multi-agent orchestration, allowing a lead coding agent to create subagents, assign parts of a task, and merge their work. Vals AI tested whether that machinery improves software quality on Vibe Code Bench. According to the published results, only one of four team configurations produced a statistically significant gain, while teams cost 1.8 to 5.1 times as much as single agents.
A full-stack test with visible seams
Vibe Code Bench asks a model to build a complete web application from a product specification and deploy it with Docker Compose. Browser agents then operate the finished app and grade whether its workflows function correctly. The benchmark exposes integration failures across the database, API, interface, permissions, and deployment configuration, making it useful for testing whether delegation improves coverage or creates coordination problems.
- Models: GPT 6 Sol and Claude Opus 5.5.
- Configurations: Single-agent and team runs at medium and maximum reasoning effort.
- Sample: Each model attempted 50 app specifications in each configuration.
- Controls: Solo and team runs used the same product prompts, sandbox, and grader.
- Team mode: The lead agent could run up to five subagents concurrently and received a short instruction to split the task, delegate work, and verify the result.
One clear gain amid higher bills
Using the conventional p < 0.05 threshold, only Sol’s medium-effort team showed a statistically significant improvement. Its score rose 7.3 percentage points, from 77.6% to 84.9%, with p = 0.005. The other three team effects ranged from a 0.3-point decline to a 3.4-point gain and were not statistically significant.
Statistical significance indicates that an observed gap is unlikely to come from run-to-run variation under the experiment’s assumptions. Cost, runtime, and effect size still determine whether that gain is useful in production.
Most of the added API spend came from cached input tokens because every subagent received its own copy of the specification and working context. Parallel execution reduced runtime for both Sol configurations, but Opus teams ran longer than their corresponding solo agents.
| Model | Effort | Mode | Score | Cost per app | Median runtime |
|---|---|---|---|---|---|
| Sol | Medium | Single | 77.6% | $1.22 | 12.4 min |
| Sol | Medium | Team | 84.9% | $3.07 | 11.9 min |
| Sol | Maximum | Single | 89.0% | $4.82 | 38.4 min |
| Sol | Maximum | Team | 90.4% | $8.54 | 27.6 min |
| Opus 5.5 | Medium | Single | 91.5% | $4.08 | 20.3 min |
| Opus 5.5 | Medium | Team | 91.2% | $9.10 | 26.6 min |
| Opus 5.5 | Maximum | Single | 89.8% | $23.77 | 74.7 min |
| Opus 5.5 | Maximum | Team | 93.2% | $122.00 | 174.1 min |
Opus’s maximum-effort team achieved the study’s highest score, 93.2%, while costing almost 30 times as much as its medium-effort single agent. The score difference between those configurations was about 1.6 percentage points and was not statistically significant. Within the tested Opus configurations, the cheapest option performed about as well as the most expensive one.
More reasoning helped Sol more than more agents
Increasing Sol’s reasoning effort raised its single-agent score by 11.4 percentage points, from 77.6% to 89.0%. Adding a team at medium effort produced a smaller 7.3-point gain. For Opus, neither higher reasoning effort nor team mode delivered a statistically meaningful improvement in the reported comparisons.
Sol and Opus delegated differently
Identical delegation instructions produced distinct orchestration strategies, showing that the lead model’s behavior can matter as much as the harness configuration.
- Sol split the architecture immediately. At medium effort, its lead assigned database and seed data, API, interface, and deployment work to separate subagents, usually within the first minute. The lead then operated the deployed app through a headless browser and routed bugs to the agent responsible for the affected files.
- Opus wrote a contract and dispatched waves. Its lead usually created a shared
CONTRACT.mdfile covering the schema, API routes, and file ownership before delegating. It did so in 42 of 50 medium-effort runs and 40 of 50 maximum-effort runs. At maximum effort, Opus first assigned foundation work, then feature implementation, followed by testing and review.
At maximum effort, the Opus team created an average of 6.8 subagents over the course of each app, while respecting the five-agent concurrency limit. It assigned a testing subagent in all 50 runs and made about 1,140 subagent tool calls per app, compared with roughly 270 tool calls for the single agent. Tool logs attributed the longer runtime to active subagent work. Four team runs exceeded four hours, and one reached the 5.25-hour limit.
Teams recovered work solo agents skipped
Sol’s medium-effort team gained at least 20 percentage points on nine apps. In those cases, the single agent sometimes reported successful testing even though features were missing and complete workflows had never been exercised. Team runs implemented more of the specification, while the lead caught cross-component failures such as forms that conflicted with the database schema and permission rules that blocked required steps.
Delegation helped most when the solo baseline omitted substantial work and the application could be divided into bounded components. Stronger solo performance left less room for a team to recover enough quality to offset duplicated context and coordination.
A practical order of operations
- Benchmark reasoning effort first. For Sol, a stronger single agent produced a larger score gain than adding subagents at medium effort.
- Use external workflow checks. Browser-level tests exposed missing features that the solo agents’ own reports had overlooked.
- Delegate separable components. Database, API, interface, and deployment tasks offer clearer ownership than work that requires constant shared state.
- Measure coordination directly. Track per-agent tokens, tool calls, wall-clock time, retries, and integration failures alongside the final score.
- Control context fan-out. At maximum effort, Opus subagents read a median of 224 million cached tokens per app, compared with 17 million for the lead and 55 million for the single agent.
- Set runtime and tool budgets. Opus’s wave-based maximum-effort team took 2.3 times as long as its solo configuration.
What the benchmark leaves open
The evidence covers two models, one full-stack web benchmark, 50 app tasks per configuration, and a brief generic delegation instruction. Specialized roles, different context-sharing designs, and tasks with cleaner parallel boundaries could produce different results. Security, maintainability, human review, and production operations also fall outside the reported browser-based score.
Within this benchmark, multi-agent orchestration produced a clear quality gain only for Sol at medium effort. Model choice and reasoning effort mattered at least as much as agent count, while duplicated context and coordination drove the largest increases in cost and runtime. Teams were most useful when the solo agent skipped features or failed to test complete workflows.