GitHub Copilot Adds Moonshot AI's Kimi K2.7, Its First Open-Weight Model

Kimi K2.7 Code becomes the first open-weight model selectable in Copilot's model picker, offering a lower-cost option for agentic coding workflows.

·
·
·
  • First open-weight model in Copilot: Kimi K2.7 Code is the first open-weight model selectable in GitHub Copilot's model picker, completing a five-lab roster.
  • Lower-cost agentic coding: Roughly 5-7x cheaper than GPT-5.5 per API call, positioned for high-volume, long-horizon coding workflows.
  • 30% fewer thinking tokens: K2.7 cuts reasoning-token usage by ~30% vs K2.6, directly reducing cost and latency for agentic tasks.
  • Availability: Rolling out now to Copilot Pro, Pro+, and Max; Business/Enterprise admins must explicitly enable it in org settings.
  • Benchmark caveats: Strong gains over K2.6 are reported, but all three headline benchmarks are proprietary Moonshot evals — independent results pending.
  • Open weights on Hugging Face: Full 1T-parameter weights available under Modified MIT License at Hugging Face for self-hosting.

GitHub Copilot just added Moonshot AI's Kimi K2.7 Code to its model picker, and the milestone is bigger than it sounds. Kimi K2.7 Code is now generally available in GitHub Copilot, and it is the first open-weight model offered as a selectable option in the Copilot model picker. Every other model in the picker , GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro , is proprietary. You cannot download it, audit its weights, or run it yourself. Kimi K2.7 is different.

Kimi K2.7 is MIT-licensed, the full 1T parameter weights are public on Hugging Face, and GitHub is simply running a hosted copy on Azure for Copilot users who prefer not to manage infrastructure. The addition also completes a five-lab Copilot roster, with the model picker now spanning OpenAI, Anthropic, Google, Microsoft, and Moonshot AI , the broadest competitive set that any major coding tool has ever offered through a single subscription.

What K2.7 Code actually is

K2.7 Code is built on Kimi K2.6, with better long-horizon coding task completion and about 30% lower thinking-token usage versus K2.6. Under the hood, it packs 1T total parameters, 32B activated parameters, 384 experts, 8 selected experts per token, a 256K context window, native multimodality, and a MoonViT vision encoder. The MoE (Mixture-of-Experts) architecture is the key to making a trillion-parameter model affordable to serve: only a fraction of those parameters activate on any given token, keeping inference costs manageable.

Reasoning models tend to overthink, spending thousands of tokens deliberating on problems that don't need it. Kimi K2.7 Code significantly reduces this tendency, cutting thinking-token usage by approximately 30% on average compared with K2.6. Across every benchmark Moonshot tested, K2.7 achieves higher scores than K2.6 while consuming fewer tokens , meaning faster responses in interactive coding sessions, lower API costs in production, and agent workflows that complete more work within the same context budget.

One important behavioral detail: K2.7 Code forces thinking and preserve_thinking on, and you cannot turn them off. The model always reasons before answering, and it keeps its full reasoning chain across multi-turn conversations. Moonshot says this "preserve thinking" mode is what boosts performance in coding-agent scenarios where context builds up over many steps.

What it's good at , and where it falls short

Real-world software engineering rarely ends in a single step. Tasks like refactoring a codebase, implementing a feature across multiple files, or debugging over long agent sessions require a model to follow instructions reliably across extended contexts and carry a task through to completion. Kimi K2.7 Code is optimized for these long-horizon scenarios, and compared with K2.6, it follows instructions more reliably in long contexts and achieves higher end-to-end task success rates.

On benchmarks, the picture is nuanced. Moonshot reports strong gains over K2.6:

  • +21.8% on Kimi Code Bench v2 (62.0 vs 50.9), +11.0% on Program Bench (53.6 vs 48.3), and +31.5% on MLS Bench Lite (35.1 vs 26.7)
  • MCP Mark Verified (correct tool invocation): K2.7 scored 81.1, beating Claude Opus 4.8's 76.4
  • SWE-bench Verified (real GitHub bug fixes): K2.7 reached 60.4%, setting a new high-water mark among open-source models

But there is a caveat worth knowing. Moonshot claims gains of 21.8% on Kimi Code Bench v2, 11% on Program Bench, and 31.5% on MLS Bench Lite , and all three are proprietary benchmarks run by Moonshot AI. Practitioners have already challenged Moonshot directly on the benchmark choices, with one developer noting that "every model 'improves' double digits on its own test suite." Independent benchmark submissions are still pending.

The honest comparison against frontier models: K2.7 is not the best model on any single benchmark, but it is the best value model for agentic coding workflows that involve tool use, long-running agents, and cost-sensitive scaling. K2.7 is roughly 5-7x cheaper per API call than GPT-5.5 and is open-source.

How to access it in Copilot

Kimi K2.7 Code is beginning to roll out to Copilot Pro, Pro+, and Max plans, selectable in the model picker in Visual Studio Code. Rollout is gradual, and GitHub will expand to Copilot Business, Enterprise, and additional surfaces over the coming weeks. The full list of supported surfaces includes:

  • VS Code version 1.127.0 or later
  • Visual Studio version 17.14.6 or later
  • Copilot CLI
  • GitHub Copilot cloud agent
  • GitHub Mobile (iOS and Android)
  • JetBrains version 1.9.1-251 or later
  • Xcode and Eclipse

For Business and Enterprise plans, Kimi K2.7 Code is off by default. Plan administrators must enable the Kimi K2.7 Code policy in Copilot settings before anyone in their organization can select it. If the policy is left off, the model stays unavailable to that organization. GitHub's own docs note it is an open-weight model that may be less aligned than other Copilot models, with an elevated risk of producing harmful content, and recommend reviewing the model card and conducting your own evaluations before enabling it.

What it costs

This model is billed at provider list pricing under usage-based billing. Each token is priced based on the model used, and the total is converted into AI credits, where 1 AI credit = $0.01 USD. Kimi K2.7 is billed at provider list pricing within that credit system , roughly GPT-5.4 mini pricing tier. For reference, the official pricing docs have the per-token breakdown. On Moonshot's own API, pricing is $0.19 per 1M cached input tokens, $0.95 per 1M cache-miss input tokens, and $4.00 per 1M output tokens.

The bigger picture

This is not just a model addition. GitHub shipped a coordinated agent stack update: model choice, browser control, automatic routing, and spend boundaries. The combined pattern is clear: Copilot is turning into an agent runtime with a paid model router behind it. Kimi K2.7's arrival is the first signal that the Copilot model picker is open to the open-source ecosystem, not just the major proprietary labs.

K2.7's combination of open weights, aggressive pricing, token efficiency, and strong MCP tool use makes it the pragmatic choice for a specific but growing set of use cases. If you are running high-volume agentic coding pipelines, iterating on large codebases, or just want a lower-cost option that can handle multi-file tasks without burning through your AI credit budget, K2.7 Code is worth testing. The official model page has the full benchmark breakdown, and the weights are on Hugging Face if you want to self-host.

Trending
  • No trending articles

Comments

avatar

Next Reads