GitHub Copilot Adds Claude Sonnet 5 With Near Opus-Level Agentic Power

Anthropic's Claude Sonnet 5 lands in GitHub Copilot with Opus-class agentic performance at Sonnet-class prices, excelling at CLI tasks and multi-step coding workflows.

·
·
  • Claude Sonnet 5 is now GA in GitHub Copilot, available to Pro, Pro+, Max, Business, and Enterprise users via the model picker -- rollout is gradual.
  • Near-Opus performance at Sonnet prices: Sonnet 5 scores 63.2% on agentic coding benchmarks vs Opus 4.8's 69.2%, closing the gap significantly over Sonnet 4.6's 58.1%.
  • Standout for CLI and agentic tasks: GitHub's internal testing flagged particularly strong performance on CLI-style tasks and excellent prompt-cache utilization at lower effort levels.
  • Introductory API pricing at $2/$10 per million tokens through August 31, 2026, then moving to $3/$15 -- roughly 60% cheaper than Opus 4.8 at standard rates.
  • Zero Data Retention (ZDR) applies for Enterprise and Business users; admins enable it via model policy settings in Copilot.
  • New tokenizer uses up to 1.35x more tokens for the same input -- Anthropic set introductory pricing to be cost-neutral, but benchmark your own workloads before assuming direct cost parity with Sonnet 4.6.

Claude Sonnet 5 is now generally available in GitHub Copilot. The rollout is gradual, so you may not see it in the model picker immediately, but it is live today for Copilot Pro, Pro+, Max, Business, and Enterprise subscribers. This is not just a routine model swap: Sonnet 5 is Anthropic's biggest leap in the Sonnet line in terms of agentic capability, and it lands at a moment when the entire coding assistant market is racing to make autonomous, multi-step workflows the default.

Opus-class power, Sonnet-class price

The headline story from Anthropic is the performance gap it closes. In GitHub's internal testing, Claude Sonnet 5 showed strong results across a range of coding scenarios, including particularly strong performance on CLI-style tasks. It also demonstrated excellent prompt-cache utilization and competitive latency at lower effort levels, making it a strong choice for developers who want fast, capable Sonnet-class performance in Copilot.

Anthropic describes it as their most agentic Sonnet yet, with top-tier intelligence for coding and everyday professional work. More concretely, its performance is close to Opus 4.8 and represents a notable improvement over Sonnet 4.6 -- Anthropic specifically notes that it performs much better on tasks involving reasoning, tool use, software coding, and knowledge work.

A key benchmark to know here: on one agentic coding benchmark, Sonnet 5 scores 63.2%, compared to Opus 4.8's 69.2% and Sonnet 4.6's 58.1%. That gap between Sonnet 5 and Opus 4.8 is now small enough that for most everyday tasks, you will not notice the difference -- but you will notice the price.

What makes it different from Sonnet 4.6

The biggest shift is not raw benchmark numbers, it is follow-through. Feedback from early access partners was consistent: Sonnet 5 is much more agentic than its predecessors. Testers described how it finishes complex tasks where previous Sonnet models would stop short, how it checks its own output without explicitly being asked, and how it handles sustained coding, tool use, and debugging well across messy technical contexts.

The model also introduces a tunable effort level. On Claude Sonnet 5, the effort parameter defaults to high on the Claude API and Claude Code -- you can set it explicitly to use a different level. This means you can dial down compute for quick questions and dial it up for complex multi-file refactors, which directly controls both latency and cost.

One practical detail worth knowing: the model's excellent prompt-cache utilization means repeated context -- like a large codebase you keep referencing -- gets reused efficiently, keeping costs down on long sessions.

Where it shines (and where it doesn't)

Based on Anthropic's evaluations and early partner feedback, here is where Sonnet 5 is strongest:

  • CLI and terminal tasks: GitHub's own testing flagged this as a standout area, making it a natural fit for Copilot CLI workflows.
  • Multi-step agentic coding: For Devin, Sonnet 4.5 (the predecessor architecture) increased planning performance by 18% and end-to-end eval scores by 12%, and it excels at testing its own code, enabling longer runs and harder tasks.
  • Brownfield codebases: One early tester noted it is at its best on legacy code -- race conditions, hidden tests -- tracing failures to their actual root cause rather than patching symptoms.
  • Prompt injection resistance: It is better at refusing malicious requests and sidestepping hijack attempts in prompt injection attacks, and hallucinates and engages in sycophantic behavior at a lower rate than Sonnet 4.6.

The areas where it falls short are equally clear. It is not on the same level as Opus 4.8 and Claude Mythos Preview when it comes to misaligned behavior, and evaluations show it has a much lower ability to perform dangerous cybersecurity tasks than current Opus models. For high-stakes security research or the most complex reasoning tasks, Opus 4.8 remains the better choice.

How to access it and what it costs

Claude Sonnet 5 is available to Copilot Pro, Pro+, Max, Business, and Enterprise users, selectable in the model picker across VS Code, Visual Studio, Copilot CLI, the GitHub Copilot cloud agent, GitHub Mobile, JetBrains, Xcode, and Eclipse.

This model is billed at provider list pricing under Usage Based Billing. GitHub Copilot moved to a token-based AI Credits system in June 2026, where when usage exceeds the included allowances for any Copilot plan, additional usage is billed in GitHub AI Credits at per-token rates (1 AI credit = $0.01 USD). Code completions and next edit suggestions are not billed in AI credits -- they remain unlimited for all paid Copilot plans.

On the API side, Claude Sonnet 5 is available at an introductory price of $2 per million input tokens and $10 per million output tokens through August 31, 2026, after which it moves to standard pricing at $3 per million input tokens and $15 per million output tokens. It vaults into a performance tier that overlaps substantially with Anthropic's flagship model, while costing roughly 60% less per token at standard pricing.

One important footnote on pricing: Sonnet 5 ships with an updated tokenizer. The same input can map to roughly 1.0 to 1.35x more tokens depending on content type. Anthropic set the introductory pricing to make the transition roughly cost-neutral, but you should benchmark your own workloads before assuming a straight cost comparison with Sonnet 4.6.

Enterprise and data retention

Copilot Enterprise and Copilot Business plan administrators can enable Claude Sonnet 5 for their organization through the model policy settings in Copilot. Like other Sonnet models in GitHub Copilot, Claude Sonnet 5 operates under Zero Data Retention (ZDR). This is a meaningful distinction from some other models in the Copilot lineup, where Anthropic retains data for safety classifiers.

The bigger picture

The emphasis on agentic capabilities -- the ability to plan, use tools like browsers and terminals, and execute multi-step workflows autonomously -- reflects where the AI industry's center of gravity has shifted in 2026. Enterprises are no longer simply asking chatbots questions; they are deploying AI systems that can navigate complex software environments, execute multi-step coding tasks, and operate with minimal human supervision.

The GitHub Copilot agentic harness already supports 20+ frontier models across the GPT, Claude, Gemini, and MAI families. You can choose the right model for the capability and cost profile of each task, or let Auto model selection choose for you, balancing task intent and model health to optimize token efficiency. Sonnet 5 slots in as the new default Sonnet-class workhorse for that system -- the model you reach for when you want strong agentic performance without paying Opus prices.

For teams running Copilot CLI heavily, building multi-step coding agents, or doing sustained autonomous work across large codebases, Sonnet 5 is worth switching to immediately. For quick inline questions or simple completions, the difference will be marginal. The real test will be in production agentic sessions, where Sonnet 5's ability to finish what it starts -- rather than stalling halfway -- is where it earns its place.

Trending
  • No trending articles

Comments

avatar

Next Reads