GitHub's HydraFusion Beats Claude Opus 5 Quality at 67% Lower Cost
GitHub extends its HydraFusion research preview from the Copilot CLI into VS Code and the Copilot app, exposing multi-model routing to more developers.
- GitHub expanded Project HydraFusion from Copilot CLI into VS Code and the Copilot app.
- Appears as a normal entry in the model picker but routes tasks across multiple models.
- Chooses between Single, Cascade, and Critique workflows based on task signals.
- Available to Copilot Pro, Pro+, Business, and Enterprise; billed at each underlying model's standard token rate.
- Offline benchmarks showed up to 67% cost reduction vs. Opus 5 on TerminalBench 2.1, mixed on DeepSWE.
- Preview is tuned for single-prompt, well-scoped tasks; multi-turn improvements are next.
GitHub has expanded Project HydraFusion, a research preview that routes one coding request through one or more models. Previously limited to the Copilot CLI, HydraFusion is now available from the model picker in Visual Studio Code and the GitHub Copilot app. The release addresses the most common request from early users and adds clearer progress updates while the orchestrator runs.
One picker entry coordinates three workflows
HydraFusion appears alongside individual models in the picker, but it acts as an orchestrator. For each task, it evaluates signals related to reasoning, code generation, debugging, and tool use. It then chooses the lowest-cost workflow expected to meet its quality threshold and returns one final response.
The router can choose among three execution patterns:
- Single: One selected model solves the task directly.
- Cascade: An efficient model drafts a solution. A quality gate accepts the result or escalates the task to a stronger model.
- Critique: One model produces a draft, a read-only critic from another model family reviews it, and the original model performs one revision.
The critic runs without tools in an isolated context, allowing it to review the draft without modifying the repository. HydraFusion applies no patch when a workflow is cancelled or fails validation. Each workflow also includes cost tracking, timeouts, and fallback behavior, making the system a controlled agent loop rather than a conventional model router.
Enable the preview
Visual Studio Code users need version 1.140 or later, including compatible Insiders builds. Select HydraFusion from the Copilot Chat model picker. If the option does not appear, enable chat.copilot.hydraFusion.enabled in the editor settings.
GitHub Copilot app users should install the latest version, open Settings, search for HydraFusion, and enable the toggle. Access is available on Copilot Pro, Pro+, Business, and Enterprise plans. Business and Enterprise administrators must allow preview features before organization members can use it.
HydraFusion has no separate subscription fee. GitHub bills the tokens consumed by each selected model at that model’s standard rate. Cascade and critique workflows can invoke multiple models during one turn, so they may consume more tokens than a direct request to a single model.
Routing extends beyond model selection
GitHub’s existing Auto option selects one model for each request. HydraFusion also chooses an execution pattern and can coordinate multiple models within the same turn. That broader scope lets it trade cost, latency, and answer quality through escalation or review.
GitHub’s research report compares HydraFusion with Claude Opus 5 in offline tests. All runs used the same medium reasoning level.
| Benchmark | Cost | Quality |
|---|---|---|
| TerminalBench 2.1 | 67% lower | 4.9 points higher |
| DeepSWE | 36% lower | 1.5 points lower |
| CheckpointBench | 65% lower | 0.1 points lower |
These figures apply to the benchmark revisions, workflow configurations, model pool, and pricing assumptions used in GitHub’s evaluation. Production results will vary with repository size, task complexity, selected tools, and current model pricing. HydraFusion posted its largest quality gain on TerminalBench 2.1. On DeepSWE, which tests repository-level fixes across multiple files, it reduced cost while trailing Opus 5 in raw quality.
Best suited to bounded coding tasks
HydraFusion’s current tuning favors first-turn coding tasks expressed in a single prompt. GitHub lists stronger support for long, iterative conversations as future work. The preview therefore fits self-contained jobs with a clear goal, relevant context, and a result that can be validated after one orchestration cycle.
Intermediate drafts remain hidden while the workflow runs. The new progress indicators expose more granular stages, but developers cannot inspect every tool call or draft as they can in some single-model flows. Complex cascade and critique runs can also take longer because they invoke additional models and validation steps.
Developers who already move difficult tasks from an inexpensive model to a stronger one can use HydraFusion to automate that pattern. Its router selects the workflow, records model costs, and prevents failed or cancelled reviews from leaving partial changes in the working tree.