On October 8, 2026, Microsoft detailed plans for GitHub Copilot to route work between local and cloud models, alongside separate sandbox controls for agent tools. Microsoft set an end-of-October target for Auto routing across Copilot CLI, the Copilot app and Visual Studio Code. This follow-up adds specific routing and sandbox details to NeoTeo’s earlier coverage of Microsoft’s Windows AI strategy.
BYOK and Auto put model choice in different hands
With BYOK, or “bring your own key,” you configure a provider or select a model yourself. Microsoft described Auto as a planned option that would route tasks between local and cloud inference based on task context and cache state across a multi-turn session. They solve different problems: one gives you direct control over the provider, while the other is intended to choose where inference runs.
| Approach | Who controls model choice | Copilot surfaces described |
| BYOK | The user configures a provider or selects a model. | Copilot CLI and VS Code chat and utility tasks. |
| Auto | Copilot is intended to route work between local and cloud inference. | Copilot CLI, the Copilot app and VS Code. |
For Copilot CLI, the documented BYOK provider types are openai, azure and anthropic. The openai type also works with Ollama, vLLM, Foundry Local and other OpenAI Chat Completions-compatible endpoints. A CLI model must support tool calling and streaming; GitHub recommends a context window of at least 128k tokens for best results.
To connect a local Ollama service, set COPILOT_PROVIDER_BASE_URL to http://localhost:11434 and COPILOT_MODEL to the identifier of a model you have installed, then start copilot. An API key is not needed when that local service does not require authentication.
In VS Code, open the chat model picker, choose Manage Language Models, then add a provider or configure a compatible endpoint. VS Code directs Ollama users to its official extension because the built-in Ollama provider is deprecated. Models used by agents need to support tool calling.
Local inference does not make every Copilot task offline
A local model handles inference on your device, but the location of inference and the network behavior of the rest of a Copilot session are separate matters. In Copilot CLI, COPILOT_OFFLINE=true prevents contact with GitHub. If the configured model provider is remote, it still receives prompts and code context.
VS Code supports fully offline BYOK chat with a local model, without a GitHub account or Copilot plan. Some other functions still require a GitHub account, including inline suggestions, semantic search and features that rely on embeddings. So local chat can work offline while those account-dependent features remain separate.
What the sandbox covers—and where its boundary stops
Microsoft describes sandboxing as a control over tool execution, separate from whether a model runs locally or in the cloud. Shell commands and, by default, local MCP (Model Context Protocol) servers and language servers run inside a process boundary.
Built-in file tools follow a different path: the agent harness applies policy checks, rather than isolating them as operating-system-enforced child processes. Remote MCP servers sit outside the local process sandbox. Choosing a local model, by itself, does not change these tool boundaries.
The sandbox backend varies by operating system
Microsoft names the MXC BaseContainer tier of ProcessContainer for Windows, Seatbelt for macOS and bubblewrap for Linux. It says these local backends do not require a separate virtual machine or container image.