Backburner uses an iPhone to help a Mac run a 27B model
Backburner, a new MIT-licensed llama.cpp fork, plugs an iPhone into a Mac over USB-C and runs the last 24 of 64 model layers on the phone, with 29 to 44 percent faster prompt reading at 16k to 48k context on a 24 GB Mac.
Image: GitHub
Why it mattersA Mac with 24 GB of memory can now hold 128,000 tokens of 8-bit context next to a 27B model, which is the size an agent session reaches after a few long file reads.
A Mac with 24 gigabytes of memory runs out of room long before an agent session fills up. Backburner is a new project that borrows the iPhone sitting on the desk and uses it as a second compute device for a local 27 billion parameter model, over the USB-C cable.
The project went up on GitHub on 1 October and reached 129 stars by 3 October. It is an MIT-licensed fork of llama.cpp plus an iPhone app, and the author says it is pre-release. The author posted measured numbers from a MacBook Pro M4 Pro with 24 GB of memory, an iPhone 17 Pro Max, and the Qwen3.8-27B model at 4-bit precision.
How the two halves split the work
The 64-layer model is cut in two. The Mac runs layers 1 to 40 on its own GPU. The iPhone runs layers 41 to 64 on its GPU, using the matrix units in the A19 Pro chip. Each 256-token batch of the prompt is pipelined between the two: while the phone works on one batch, the Mac starts the next. The last batch runs on the Mac so the final output stays on the laptop. Only the state that changes between batches crosses the cable.
Past 64,000 tokens, the Mac runs every layer and the phone changes jobs. It holds the oldest key-value pages of the context and does the attention over them for every new token. That is how a 24 GB Mac fits 128,000 tokens of 8-bit context, which the Mac alone can only reach by dropping to 4-bit.
What was measured
Reading a 2,000-token file into a saved agent session that already holds some context, from the author's turn-bench.py:
- At 16k context: 109 tokens per second on the Mac alone, 157 tokens per second with the iPhone attached. That is 44 percent faster, and the wait falls from 18.8 seconds to 13.1.
- At 32k context: 101 tokens per second alone, 130 with the iPhone, 29 percent faster.
- At 48k context: 87 tokens per second alone, 113 with the iPhone, 30 percent faster.
Starting a new session with a 26,849-token system prompt, project notes and twelve tools fell from 245 seconds on stock llama.cpp to 168 seconds with the fork and the iPhone. Writing speed below 64,000 tokens does not change when the iPhone is plugged in, because the Mac handles decoding on its own below that threshold. The author also reports that greedy output is token-identical with and without the phone, in every length tested.
What it needs, and what it does not fix
An Apple Silicon Mac, an iPhone 15 Pro or newer, and a 10 Gb/s USB-C cable. The cable in the iPhone box is USB 2 and too slow. The code is in a public GitHub fork of llama.cpp. One request at a time, no batching.
For a working developer, this changes the ceiling on a local model used behind Codex, Claude Code or Opencode. A long tool result that would have made the Mac swap at 70,000 tokens now fits, and the first file read of a session is roughly a third faster. Small reads under 512 tokens stay on the Mac, which in a real agent session was 29 of 36 requests. If the phone fails for any reason, Backburner turns it off for 60 seconds and the current batch reruns on the Mac. Saving a session while the phone holds keys past 64,000 tokens is implemented but not yet tested end to end.
Source
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.

