GitHub Copilot Routes Tasks Between Local and Cloud Using MAI Code 1.1 Flash

GitHub Copilot will soon route coding tasks between on-device models and the cloud automatically, backed by a new local model and OS-level sandboxing.

·
·
·
  • GitHub Copilot will auto-route tasks between local and cloud models to save AI credits, rolling out by month end.
  • New on-device MAI Code 1.1 Flash: 137B MoE with 6.8B active params, quantized to 53GB.
  • Hits 70.80% on SWE-Bench Verified on-device vs 72.6% for the cloud Bfloat16 version.
  • Targets NVIDIA RTX Spark hardware like Surface Laptop Ultra with 128GB unified memory.
  • Ships with Microsoft Execution Containers, open-source sandboxing using native OS primitives on Windows, macOS, and Linux.
  • Available in Copilot CLI, Copilot app, and VS Code via Auto mode or explicit model selection.

Copilot will route coding tasks between local and cloud models

GitHub plans to let Copilot choose whether a task runs on a laptop or a cloud model. Microsoft says the feature will arrive by the end of the month, although the announcement gives no calendar date. Automatic routing aims to reduce paid cloud usage while preserving access to larger models for demanding work.

The feature extends Project HydraFusion, GitHub’s system for selecting among multiple AI models. The new version adds compute location to that decision. Microsoft is also releasing a quantized coding model for local inference and an open-source sandboxing library that limits an agent’s access to the host system.

A 137-billion-parameter model in 53 GB

MAI Code 1.1 Flash is a coding-focused mixture-of-experts model with 137 billion total parameters and 6.8 billion active parameters. A mixture-of-experts architecture activates a subset of the model for each token, reducing the computation required during inference.

Microsoft fits the model onto a laptop through quantization and speculative decoding. Quantization stores weights at lower numerical precision, while speculative decoding lets a smaller model draft several tokens for the main model to verify in parallel. Microsoft says the resulting build occupies 53 GB, about 80% less space than its Bfloat16 cloud version.

Microsoft reports the following benchmark results. The figures are vendor-provided and have not been independently replicated in the announcement.

Reported coding and terminal benchmark scores
Benchmark Local MAI Code 1.1 Flash Cloud MAI Code 1.1 Flash GPT-OSS-120B
SWE-Bench Verified 70.80% 72.6% 32.0%
Terminal-Bench 2.1 66.29% Not reported 23.6%
Decode throughput for MAI Code 1.1 Flash at several prompt lengths
Microsoft’s reported decode throughput across increasing context lengths.

On a Surface Laptop Ultra, Microsoft measured prompt processing at 923.5 tokens per second with a 64K context and 769.8 tokens per second with a 128K context. Decode throughput declined from about 65 tokens per second at 2K context to roughly 40 at 256K. Peak memory use reached 75.5 GB at 256K context.

The 53 GB model is only the start

Microsoft’s reference system is a Surface Laptop Ultra built around NVIDIA RTX Spark, with as much as 128 GB of unified memory and one petaflop of advertised AI compute. Unified memory gives the processor and GPU access to the same large memory pool, allowing the 53 GB model to remain resident without fitting inside a conventional discrete GPU’s dedicated memory.

Inference also requires memory for the operating system, applications, runtime components, and the key-value cache. That cache stores attention state for previously processed tokens and grows as an agent reads source files, command output, and conversation history. Long sessions can therefore consume substantially more memory than the model weights alone.

Copilot gets automatic and manual routes

GitHub plans to expose local models through the Copilot CLI, Copilot app, and VS Code using two selection modes:

  • Automatic orchestration: Copilot selects local or cloud inference using the task, session context, and available cached state. The router can reconsider that choice during a multi-turn session.
  • Explicit selection: Developers can select MAI Code 1.1 Flash through the Windows ML provider or connect Copilot to an OpenAI-compatible local endpoint. Compatible runtimes include llama.cpp, Ollama, and LM Studio.
Copilot model selection and local inference flow
Copilot can select a local model automatically or use a developer-configured endpoint.

The sandbox defines the real boundary

A local agent retains access to files, shells, and development tools unless the operating system restricts it. Microsoft addresses that risk with Execution Containers (MXC), an open-source library that converts a declarative security policy into each operating system’s native isolation mechanisms.

Sandbox implementation by operating system
Operating system Isolation backend
Windows BaseContainer tier of the ProcessContainer backend
macOS Seatbelt
Linux bubblewrap

MXC uses process-level isolation without requiring a virtual machine or container image. When enabled, shell commands and local Model Context Protocol servers run inside the sandbox. The working directory remains writable, while most other locations are read-only or inaccessible.

Where the boundary stops

  • Built-in file tools: These run inside GitHub Copilot. Copilot’s agent controller checks their requests against the active policy, but the operating system does not isolate them as sandboxed child processes.
  • Remote MCP servers: Remote Model Context Protocol servers operate outside the local process sandbox and require separate trust and network controls.
  • CLI configuration: Developers enable and configure the boundary with the /sandbox command.

Workflows that benefit from local execution

Local inference and process isolation support several practical development workflows:

  1. Offline automation: A scheduled job can inspect repositories, run tests in isolated working copies, and generate an HTML report. Full offline operation requires an explicitly selected local model and locally available dependencies.
  2. Lower cloud usage: Routine searches, boilerplate changes, lint fixes, and small refactors can run locally, while automatic routing sends more complex tasks to cloud models when needed.
  3. Sensitive repositories: Explicit local selection can keep prompts and model inference on the device. Network restrictions, sandbox policies, and careful handling of remote MCP servers remain necessary to control other data paths.

What teams should verify before rollout

  • Memory capacity: The 53 GB model requires additional headroom for context caches, the inference runtime, development tools, and the operating system.
  • Routing behavior: Automatic mode may use cloud models. Workloads with strict data-location requirements should pin an approved local model.
  • Security policy: Sandbox rules should cover writable paths, command execution, local MCP servers, and network access. Remote MCP services require separate controls.
  • Repository-specific quality: Microsoft’s benchmark scores provide a baseline, but latency, accuracy, and tool use should be tested against each team’s languages, build systems, and codebase size.

Microsoft’s reported 70.80% SWE-Bench Verified score places the local build close to its 72.6% cloud counterpart on that benchmark. Actual adoption will depend on whether the reduction in cloud usage offsets the cost and availability of machines with enough unified memory.

Trending
  • No trending articles

Comments

avatar

Next Reads