GitHub Drops Moonshot AI's Open-Weight Kimi K2.7 Into Copilot's Model Picker

Moonshot AI's Kimi K2.7 Code is the first open-weight model in GitHub Copilot's picker, bringing a 1T-parameter MoE coding model to millions of VS Code users at lower cost.

·
·
·
  • First open-weight model in Copilot: Kimi K2.7 Code from Moonshot AI is now selectable in GitHub Copilot's model picker, a first for open-weight models.
  • 1T-parameter MoE architecture: Only 32B parameters activate per token, making inference significantly cheaper than equivalent dense models.
  • Lower cost, strong agentic performance: ~30% fewer thinking tokens vs K2.6; priced below frontier closed models in Copilot's usage-based billing system.
  • Available now on Pro/Pro+/Max: Accessible in VS Code 1.127.0+, JetBrains, Xcode, Eclipse, CLI, and GitHub Mobile; Business/Enterprise requires admin opt-in.
  • Benchmarks are vendor-reported: K2.7 beats K2.6 by up to 31.5% on Moonshot's own evals but still trails GPT-5.5 and Claude Opus 4.8; no independent third-party results yet.
  • Copilot now spans 5 AI labs: OpenAI, Anthropic, Google, Microsoft, and Moonshot AI — the widest multi-provider roster of any major coding assistant.

GitHub just made a quiet but consequential change to its Copilot model picker: Kimi K2.7 Code, an open-weight model from Beijing-based Moonshot AI, is now generally available as a selectable option. It is the first time a model whose full weights are publicly downloadable has appeared in Copilot's picker alongside Claude, GPT, and Gemini. For most developers, the practical change is simple: a new entry in a dropdown. But the implications run deeper than that.

A new kind of model in a familiar place

Every other model in Copilot's picker is proprietary. You cannot download it, audit it, or run it yourself. Kimi K2.7 is different: it is MIT-licensed, the full 1T parameter weights are public on Hugging Face, and GitHub is simply running a hosted copy on Azure for Copilot users who prefer not to manage infrastructure. GitHub's move is distribution: bringing an open-weight 1T-parameter MoE coding model into the same picker where developers already choose Claude, GPT, and Gemini variants.

The Copilot model picker now spans five independent AI labs: OpenAI, Anthropic, Google, Microsoft, and Moonshot AI, making it the only major coding tool in the market that currently routes across five separate AI providers under a single subscription. That happened fast. Moonshot AI published the weights on Hugging Face just 19 days before GitHub shipped them into general availability. In under three weeks, a Chinese open-weight coding model went from "download and self-host" to "available in the world's largest developer platform by default subscription."

What the model actually is

Kimi K2.7 Code is built on a Mixture-of-Experts (MoE) architecture with 1 trillion total parameters and 32 billion activated parameters per token. The model supports a 256K context length and uses Multi-head Latent Attention (MLA). It also includes MoonViT, a 400M-parameter vision encoder.

MoE is the key architectural idea here. In a standard dense model, every single parameter activates for every token processed. MoE breaks that: it routes each token through only a small subset of specialized "expert" sub-networks. Kimi K2.7 Code has 1 trillion total parameters but only 32 billion active per token. That means you get the knowledge capacity of a massive model at a fraction of the inference cost per call.

Kimi K2.7 Code is purpose-built for coding tasks. For general-purpose work such as writing, analysis, and conversation, Moonshot recommends K2.6, which offers more well-rounded capabilities. One notable constraint: it runs exclusively in thinking mode and does not support temperature adjustment. Moonshot AI has fixed it at 1.0, meaning teams cannot tune output determinism the way they might with other models.

The performance story

On coding benchmarks, K2.7 Code shows substantial gains over K2.6: +21.8% on Kimi Code Bench v2 (62.0 vs 50.9), +11.0% on Program Bench (53.6 vs 48.3), and +31.5% on MLS Bench Lite (35.1 vs 26.7). The efficiency gains are arguably more interesting than the raw scores. K2.7 Code improves reasoning efficiency, reducing thinking-token usage by approximately 30% compared with K2.6. Since thinking tokens bill as output tokens in agentic runs that span hundreds of steps, that 30% compounds quickly into real cost savings.

The key improvement under the hood is how the model generates code. Where K2.6 produced implementations by wrapping existing libraries and routing through established frameworks, K2.7-Code authors implementations directly. Moonshot AI says this produces more reliable generalization across Rust, Go, and Python, and across task types including frontend development, DevOps, and performance optimization.

That said, the benchmark picture has caveats worth knowing:

  • K2.7-Code still trails both GPT-5.5 and Claude Opus 4.8 on most benchmark rows.
  • The one clear win over Opus 4.8 is MCP Mark Verified (81.1 vs 76.4), a tool-use suite. As of the model's release, there were no independent third-party numbers on standard public benchmarks such as SWE-bench Verified or LiveCodeBench.
  • Researcher Elliot Arledge ran K2.7-Code against K2.6 and Claude Fable 5 on KernelBench-Hard and wrote: "K2.7 is more honest but not more capable."
  • K2.7 is not the best model on any single benchmark. It is the best value model for agentic coding workflows that involve tool use, long-running agents, and cost-sensitive scaling.

Where it fits in your workflow

Kimi K2.7 Code is explicitly built for long-horizon software engineering tasks: scenarios where an AI must maintain context over thousands of steps, navigate repositories, invoke tools, edit code across modules, run tests, debug, and iterate until completion. Think multi-file refactors, test generation across a large codebase, or agentic loops that call external tools repeatedly.

Where it is less suited:

  • General-purpose chat or writing tasks (use K2.6 for those)
  • Tasks requiring the absolute frontier on hard reasoning (Claude Opus 4.8, GPT-5.5), a 1M-token context, or a proven third-party benchmark record.
  • Workflows where you need deterministic, low-temperature outputs. The fixed temperature and mandatory thinking mode make it unsuitable for trivial, cheap completions.

How to access it

Kimi K2.7 Code is beginning to roll out to Copilot Pro, Pro+, and Max plans. You can select the model in the model picker in Visual Studio Code. Access is also available via Visual Studio version 17.14.6 or later, JetBrains version 1.9.1-251 or later, Xcode, Eclipse, Copilot CLI, GitHub.com, and GitHub Mobile.

For enterprise users, the story is different. Kimi K2.7 Code is off by default for Copilot Business and Copilot Enterprise. Plan administrators must enable the Kimi K2.7 Code policy in Copilot settings before anyone in their organization can select it. If the policy is left off, the model stays unavailable to that organization. GitHub recommends administrators review open-weight models against their own security, compliance, and data-governance requirements before enabling them.

On billing: billing follows the AI Credit system GitHub introduced, where one credit equals $0.01. Kimi K2.7 Code is billed at provider list pricing, placing it in a lower cost tier than the frontier proprietary models in Copilot's current roster. No new subscription tier is required. Outside of Copilot, Moonshot's direct API pricing is $0.19 per 1M cached input tokens, $0.95 per 1M cache-miss input tokens, and $4.00 per 1M output tokens.

The open-weight angle matters more than it sounds

Unlike fully proprietary models where the weights remain locked away, Kimi K2.7 Code's parameters are publicly available for download. This means organizations can pull the model onto their own infrastructure for fine-tuning, offline use, or security audits. For enterprises that balk at shipping proprietary code to a cloud API, the open-weight approach offers a path to run the model in a self-hosted environment, potentially even on-premises servers.

Self-hosting is technically possible but not casual. The Hugging Face file listing shows a repository size of roughly 595 GB, which is why K2.7 Code is better thought of as an open-weight frontier-style deployment target or hosted API model, not a small local coding assistant. For practical interactive use, you need enterprise GPU hardware (8x H100 80GB minimum for INT4 inference).

For the vast majority of Copilot users, the Azure-hosted version is the right path. Because Kimi K2.7 Code is served through Azure, prompts flow through Microsoft's Azure AI content safety stack. GitHub confirmed that customer code snippets are not stored or used to train the model when accessed via Copilot, adhering to the same data privacy commitments that apply to other models in the picker.

What this signals for the field

GitHub is turning Copilot from a single assistant into a managed marketplace of model behavior, price, provenance, and risk. For developers and administrators, the model picker is becoming less like a preference menu and more like a policy boundary. The addition of an open-weight model from a non-Western lab is part of that shift.

For the industry, this is a signal: frontier open-weight coding models are now default IDE options, not side projects. The community reaction on X captured it well: early community reaction framed it as "the moat keeps getting cheaper" and "open weight in a Microsoft product is like a stray cat becoming a service dog." That framing is half-joking, but it points at something real. The gap between "open-source model you self-host" and "model available in your IDE by default" just closed significantly.

The practical advice for now: if you are on a Pro plan, the model is available in VS Code today. Run it on your actual tasks, particularly multi-file refactors and agentic tool loops, and compare premium request consumption against your current default. K2.7 is roughly 5-7x cheaper per API call than GPT-5.5 and is open-source, which makes the experiment low-risk. The benchmarks are a starting point, not a verdict.

Trending
  • No trending articles

Comments

avatar

Next Reads