GitHub Copilot will automatically route coding tasks between local and cloud models by the end of October, expanding Project HydraFusion from model selection to execution-location decisions1. Microsoft has not disclosed how much repository context the Auto routing mode sends to cloud models, whether developers can inspect routing decisions, or whether Auto can be restricted to local-only inference.
The plan was outlined Wednesday in a post co-written by Patrick Nikoletich, a GitHub product manager, and Stuart Schaefer, a Windows platform partner architect. Copilot will weigh task context and cache state when switching between local and cloud inference, including during multi-turn sessions. In Copilot CLI, the Copilot app, and VS Code, developers can use Auto routing or manually select a local model. Local model options include MAI Code 1.1 Flash through the WindowsML provider and OpenAI-compatible local endpoints.
MAI Code 1.1 Flash specs
MAI Code 1.1 Flash is a mixture-of-experts model with 137 billion total parameters and 6.8 billion active. Microsoft used mixed-precision quantization at roughly 3.3 bits per weight to shrink the model to 53 GB, an 80% reduction from the bfloat16 cloud version. Speculative decoding is paired with quantization to speed up local inference. Microsoft measured peak memory use of 75.5 GB at a 256K-token context. The initial rollout targets NVIDIA RTX Spark Windows PCs such as Surface Laptop Ultra, which offers up to 128 GB of unified memory.
On SWE-Bench Verified, Microsoft reports the full-precision version scored 72.6% and the quantized model scored 70.8%. On Terminal-Bench 2.1, a dataset of 89 tasks, the quantized model outperformed the original, scoring 66.29% against 62.9%.
Sandboxing controls and their limits
GitHub made its new sandboxing controls generally available alongside the routing announcement. When sandboxing is enabled, Copilot applies OS-enforced restrictions to shell commands and, where supported, local MCP and language servers, using Microsoft's open source Execution Containers (MXC) library. The OS-level backend varies by platform: BaseContainer on Windows via ProcessContainer, Seatbelt on macOS, and bubblewrap on Linux.
Built-in file tools run inside the agent process, where the harness checks requests against sandbox policy rather than relying on OS-enforced isolation. Remote MCP servers sit outside the local process sandbox; Copilot checks their connection policies when MCP sandbox controls are enabled. GitHub says those restrictions apply regardless of whether a task runs on a local or cloud model.
Nikoletich and Schaefer acknowledge that selecting a local model keeps inference on the device but does not stop the agent from reaching external services or making network requests through its tools. Developers who need a fully local session will have to lock down what those tools can access.
ANALYSIS Anthropic said it could route Claude Sonnet 5.5 requests to Claude Sonnet 5 when it detects higher-risk activity. Microsoft's version operates at a different layer, deciding not which model but which hardware handles a given turn. The unanswered transparency questions, particularly whether developers can audit or override Auto's routing choices, will determine how much trust enterprise teams place in the system.