On October 7, 2026, Microsoft said its coding model MAI-Code-1.1-Flash was available to download and run locally on Windows, with no per-call inference charge for local use. The trade-off is hardware: Microsoft reports a 53 GB model footprint and 75.5 GB peak memory use at the model’s 256,000-token context. Experimental local-model options in GitHub Copilot were planned separately, for the end of October.
Microsoft announces local execution for MAI-Code-1.1-Flash
Microsoft describes MAI-Code-1.1-Flash as a coding-focused mixture-of-experts model. It has 137 billion total parameters, with 6.8 billion active parameters. In a mixture-of-experts design, different parts of the model handle different inputs; the active-parameter count is not the model’s full on-device footprint.
Microsoft says local calls carry no model-inference fee. That covers inference charges, not the cost of a computer or the electricity it uses. NeoTeo previously covered the October 7 Windows and AI event; Microsoft’s announcement that day put a specific local coding model in focus.
The local model’s memory demands
Microsoft gives three distinct figures for the local version: a 53 GB quantized model footprint, 75.5 GB peak memory use at a 256,000-token context, and a recommendation for devices with more than 120 GB of RAM for best performance.
These numbers describe different things. The footprint is the model’s on-device size; the peak-memory figure is for use at the maximum stated context; and the RAM figure is Microsoft’s broader recommendation. Microsoft also notes that unified memory is shared by the CPU and GPU, while the operating system, applications, inference runtime, and key-value cache use part of that pool. A device’s advertised memory capacity therefore does not all remain available to the model.
Microsoft’s reported benchmark results
Microsoft says it ran the local-model benchmark tests on October 5, 2026, using mixed-precision quantization at about 3.3 bits per weight, DFlash2 sliding-window speculative decoding, and a Windows ARM64 llama.cpp CUDA runtime. The company reported these results for the local quantized and full-precision versions:
| Benchmark and dataset size | On-device quantized MAI-Code-1.1-Flash | Full-precision MAI-Code-1.1-Flash |
| SWE-Bench Verified, 500 tasks | 70.80% | 72.6% |
| Terminal-Bench 2.1, 89 tasks | 66.29% | 62.9% |
The local version scored lower than the full-precision version on SWE-Bench Verified and higher on Terminal-Bench 2.1 in Microsoft’s table. Those figures describe the company’s stated test setup, not a promise of the same results on every device.
The planned GitHub Copilot workflow
Microsoft said experimental local-model options were planned for GitHub Copilot CLI, the Copilot app, and VS Code by the end of October 2026. The plan described Auto routing between local and cloud inference, as well as manual selection of a local model. Microsoft presented these options as forthcoming, rather than as a completed rollout.
Microsoft also described sandbox controls for agent commands. Developers must configure the sandbox separately to set boundaries for file and network activity; local inference alone does not impose those controls.