AI NewsOpen sourceAnnouncement

Edge0 open-sourced its iOS, Android, macOS and Windows apps for running 35B models from a phone's SSD

Edge0 open-sourced its iOS, macOS, Android and Windows apps on 30 September, so a 35 billion parameter mixture-of-experts model can run on a phone from its own SSD, with 2.9 GiB of active memory on a Mac mini M4 Pro.

AI News

Editorial2 min read

LinkedInX

Why it mattersA phone that can run a 35 billion parameter model from SSD, in 2.9 GiB of active memory, moves on-device assistants and offline agents out of the lab and into apps a mobile team can ship.

A phone with 4 GiB of free memory can run a 35 billion parameter mixture-of-experts model today, because the model's expert weights stream from the device's SSD and only the active set stays in RAM. Edge0 open-sourced its iOS, macOS, Android and Windows apps on 30 September, with native engines for each platform in the same repository. The project's open-source release went live on 8 September and the repository is at 3,063 stars under Apache-2.0, pushed yesterday.

What the framework does differently

Edge0 is a streaming inference framework for sparse mixture-of-experts models, built around a pattern the authors call "SSD expert offload plus Recover-LoRA plus prerouter routing prediction." Expert weights are stored on disk and loaded on demand, so peak memory is bounded by the active expert set rather than the full parameter count. A trained head predicts routing one step ahead, so the SSD read for the next expert overlaps the current forward pass, which the project's own benchmark reports gives "up to +59%" more decode speed than naive loading. A frozen 4-bit base runs under LoRA adapters trained by distillation, which the authors say recovers most of the quantisation loss.

Two released tiers, with the authors' own benchmark

Edge0 publishes two model tiers on Hugging Face and ModelScope. Edge0-35B-A3B-preview is 4-bit, 40 layers, 256 experts, with the prerouter picking four at a time. Edge0-8B-A1B-preview is 4-bit, 24 layers, 128 experts, prerouter picking eight. On a Mac mini M4 Pro with 24 GB of RAM, the 35B tier decodes at 14.9 to 17.7 tokens per second with 2.9 GiB of peak active memory, and the 8B tier decodes at 23.9 to 25.3 tokens per second with 1.0 GiB of peak active memory, measured in the project's own bench.py on a 3.3 thousand-token prompt.

The quality trade, in the authors' own numbers

The project reports its own OpenCompass results for both tiers against the full-precision base models they are built on. On AIME 2026, HumanEval, GPQA-Diamond, MMLU-Pro and IFBench, averaged, the 35B tier loses 3.9 points against the fp16 Qwen3.6-35B-A3B base, and the 8B tier loses 2.8 points against the fp16 Ling 3.0 tiny base. The 8B tier scores 70.1 on MMLU-Pro against the fp16 base's 65.8, which is the only cell in the published table where the quantised model beats the full-precision one. These are the authors' numbers, run under settings they chose, and no independent evaluation has published yet.

Four platform engines ship as separate code paths in the same repo: a macOS CLI and desktop app in Rust, an on-device iPhone app in Swift with MLX Swift, an Android app and native engine in Kotlin, and a Windows app and engine in C++ with Vulkan. One unified framework that adapts runtime to each device under one access layer is scheduled for Q4 2026 on the project's roadmap. The technical report, "The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction," is on arXiv as paper 2609.18063.

Source

SourceEdge0

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX