AI NewsDev toolsAnnouncement

Microsoft adds a 53GB on-device coding model to GitHub Copilot

Microsoft said GitHub Copilot will run coding tasks on the device by the end of October, using a 53GB quantized build of its MAI Code 1.1 Flash model on NVIDIA RTX Spark PCs.

AI News

Editorial3 min read

LinkedInX
Header image for the Microsoft blog post on local models and sandboxed tools in GitHub Copilot

Image: Microsoft

Why it mattersA team with repository-confidentiality rules has to find out before adoption what Auto still sends to the cloud, because the on-device option does not by itself make a session private.

A coding model small enough to run on a laptop still weighs 53GB, and most laptops cannot hold it.

Patrick Nikoletich of GitHub and Stuart Schaefer of the Windows platform team wrote on 7 October that GitHub Copilot will run coding tasks on the device by the end of October, using a quantized version of Microsoft's own MAI Code 1.1 Flash as the on-device model. Copilot will decide, through what Microsoft calls Auto, whether each task runs locally or in the cloud, or a developer can pin a specific local model.

The model and the hardware it needs

MAI Code 1.1 Flash is a mixture-of-experts coding model with 137 billion total parameters and 6.8 billion active parameters. Microsoft says the on-device build applies mixed-precision quantization at about 3.3 bits per weight, plus sliding-window speculative decoding, to bring the file down to 53GB, which it describes as an 80 percent reduction from the Bfloat16 cloud version.

Microsoft says peak memory use was 75.5GB at a 256k-token context on a Surface Laptop Ultra, the first NVIDIA RTX Spark PC. The weights alone leave no room for the operating system, the inference runtime, and the key-value cache, which grows with the agent's context. Microsoft lists the initial target as RTX Spark Windows PCs with up to 128GB of unified memory, which rules out most 16GB and 32GB developer laptops.

The benchmarks

Microsoft's own numbers compare the quantized on-device model with the full cloud version. On SWE-Bench Verified, 500 tasks, the cloud model scored 72.6 percent and the quantized on-device model scored 70.80 percent. On Terminal-Bench 2.1, 89 tasks, the cloud model scored 62.9 percent and the on-device model scored 66.29 percent. On a benchmark that small the gap is three tasks, which Microsoft points out rather than compares further. Prompt-processing throughput was 923.5 tokens per second at a 64k context and 769.8 at 128k. Microsoft dates the test to 5 October and reports it on a Windows ARM64 llama.cpp CUDA runtime.

What Microsoft has not said about Auto

Microsoft writes that "local inference does not make the session offline." Auto decides where each task runs across a multi-turn session, using task context and cache state, and may route work to the cloud. The post does not say how much repository context or conversation history Auto sends to the cloud when it does so, whether a developer can see those routing decisions in the UI, or whether a team can pin a project to local inference only.

The New Stack, which covered the post on 8 October, says the result is that a team with strict data-handling rules still does not know what Copilot sends to the cloud under Auto. Picking a local model keeps model calls on the device, but tools inside the session, including remote MCP servers, can still reach the network.

A team that was waiting for an on-device coding option so it could run Copilot on a confidential repository now has a decision to make before the end of October. The hardware a Windows laptop needs to run the model is one half of that decision. What Auto sends when it routes is the other, and Microsoft has not answered it yet.

Source

Bringing local models and sandboxed tools to Windows and GitHub Copilot, Microsoft. GitHub Copilot is going local, but Microsoft won't say what gets sent to the cloud, The New Stack.

SourceMicrosoft

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX
Start a project