GitHub Adds Claude Sonnet 5.5 to Copilot, Cutting Coding Costs by 30%
Anthropic's new mid-tier model lands in Copilot's picker, promising the same coding quality as Sonnet 5 with fewer tool calls and lower per-task cost.
- Claude Sonnet 5.5 is generally available in GitHub Copilot across every major IDE and the CLI.
- Matches Sonnet 5 on coding while using fewer steps, tokens, and tool calls, and finishing faster.
- Scores 70.6% on Terminal-Bench 4.0 versus 10.3% for Sonnet 5 and 66.4% for Opus 5.5.
- Pricing is unchanged at $2 per million input and $10 per million output tokens.
- First Sonnet to ship with Opus-tier cyber safeguards and anti-distillation classifiers.
- Available to Copilot Pro, Pro+, Max, Business, and Enterprise plans with gradual rollout.
Claude Sonnet 5.5 arrives in GitHub Copilot
GitHub is gradually rolling out Claude Sonnet 5.5 in GitHub Copilot, adding Anthropic's latest workhorse model to paid plans. GitHub says it can match Claude Sonnet 5 on coding tasks while using fewer steps, tokens, and tool calls, which shortened completion times in early tests.
Each unnecessary tool call adds a network round trip, expands the context the model must process, and consumes billable tokens. Those costs accumulate during agentic sessions, where Copilot reads files, runs tools, checks results, and revises its work with limited developer input.
One model across eight Copilot surfaces
- Editors and IDEs: VS Code, Visual Studio, JetBrains IDEs, Xcode, and Eclipse
- Agent tools: Copilot CLI and the Copilot coding agent
- Mobile: GitHub Mobile
Access covers Copilot Pro, Pro+, Max, Business, and Enterprise plans. The model may take time to appear because GitHub is enabling it gradually. Business and Enterprise administrators control access through the model policy in Copilot settings; organizations using default model enablement receive new models automatically unless an administrator disables them.
Sonnet leads one agentic benchmark
Sonnet 5.5 is the second model in Anthropic's Claude 5.5 family, following Opus 5.5 by several days. Anthropic targets Opus at its most demanding coding and reasoning work and positions Sonnet as the lower-cost option for frequent tasks that still require substantial reasoning.
Terminal-Bench 4.0 measures whether an AI agent can complete practical tasks through a terminal, including choosing commands, inspecting results, and recovering from mistakes. Anthropic reports the following scores:
| Model | Score |
|---|---|
| Claude Sonnet 5.5 | 70.6% |
| Claude Opus 5.5 | 66.4% |
| Claude Sonnet 5 | 10.3% |
Sonnet 5.5 leads Opus 5.5 by 4.2 percentage points on that test and improves on Sonnet 5 by 60.3 points. Its scope remains one agentic benchmark, and results can vary with repository structure, tools, prompts, and task type.
Other reported evaluations
- Humanity's Last Exam: 64.5% with tools on a difficult multidisciplinary question-answering benchmark
- OSWorld 2.1: 80.1% partial credit on tasks involving visual interaction with desktop software
- Chartography: 61.6% without tools on a chart-based evaluation
- GDPval-AA: two points below Opus 5.5 on work modeled on real occupational tasks
Anthropic also says Sonnet 5.5 is the first Sonnet model to finish Pokémon Red using screenshots as its only visual input. The demonstration probes long-horizon planning and visual grounding, although its result is not directly comparable with the scored benchmarks.
Stable rates, lower claimed task cost
Anthropic kept Sonnet 5.5's token rates unchanged from Sonnet 5. Its published list prices are:
| Token type | Price |
|---|---|
| Input | $2.00 |
| Output | $10.00 |
| Cache read | $0.20 |
| Cache write | $2.50 |
Opus 5.5 lists at $4 per million input tokens, $20 per million output tokens, and $5 per million cache-write tokens. Sonnet therefore costs half as much for those token categories.
Anthropic attributes its advertised 30% reduction in cost per task to lower consumption. Early testers reported that Sonnet 5.5 understood codebases faster, batched tool calls, and completed work in fewer steps. Effort settings control how much reasoning budget the model uses; in reported evaluations, Low or Medium effort surpassed Sonnet 5's best result on several tests at roughly one-tenth the cost per task.
GitHub's billing documentation explains how provider pricing applies under usage-based billing. Included allowances, premium-request rules, organization policies, and spending limits can still change the final Copilot charge.
Cyber prompts can trigger a fallback
Anthropic is applying safeguards to Sonnet 5.5 that it previously reserved for its highest-capability models. The company says the model's cybersecurity capabilities are comparable to Opus 5, so higher-risk cyber requests can be visibly routed to Sonnet 5. Security researchers working with red-team or exploit-development prompts are the users most likely to encounter the fallback. Biology safeguards remain unchanged.
Separate anti-distillation classifiers are designed to detect attempts to extract internal reasoning traces for training competing models. These classifiers can block extraction-style prompts, while ordinary software development remains within the model's intended use.
Select it, then measure it
- Open the model picker in a supported Copilot surface and choose Claude Sonnet 5.5.
- If the model is missing, confirm that the rollout has reached the account and that an organization administrator has enabled it.
- Run a representative set of coding tasks and record completion time, token use, tool calls, correctness, and review effort.
- Compare those results with Sonnet 5 or the team's current default before changing organization-wide policies.
Teams moving from Sonnet 5 retain the same per-token rates and need no repository changes to make the switch. Actual savings depend on task complexity and agent behavior, with long-running repository work offering the clearest opportunity to benefit from fewer tool calls and shorter execution times.