Backburner Turns Your iPhone Into a Mac's AI Co-Processor, 44% Faster
Backburner turns a plugged-in iPhone into a co-processor for a 24GB MacBook, boosting Qwen3-27B prefill up to 44% and extending context past 128k tokens.
- Backburner uses an iPhone over USB-C as a co-processor for llama.cpp on Apple Silicon Macs: repo here.
- Prefill speedup of 29-44% at 16k-48k context on Qwen3.8-27B, measured on an M4 Pro 24GB.
- Extends usable context from 64k to 128k+ at 8-bit on a 24GB Mac without quantizing KV cache.
- Output is token-identical with and without the phone; this is throughput, not approximation.
- Mac runs layers 1-40, iPhone runs 41-64 on its GPU, pipelined per 256-token ubatch.
- Includes custom llama.cpp fork with SME2 kernels, Metal fusions, and DFlash2 speculative decoding.
Backburner turns an iPhone into a local LLM accelerator
Backburner is an open-source project that distributes local large language model inference across an Apple Silicon Mac and an iPhone. Connect the devices with a 10 Gb/s USB-C cable, launch the companion iOS app, and the phone contributes its GPU and Neural Engine to a 27-billion-parameter model running through llama.cpp.
In the project’s reported benchmarks, the added hardware increased long-context prompt processing throughput by as much as 44%. It also let a 24GB Mac extend an 8-bit key-value cache beyond its approximate 64,000-token memory limit, with tests reaching 128,000 tokens.
| Mac | 24GB M4 Pro MacBook |
|---|---|
| Phone | iPhone 17 Pro Max |
| Model | Qwen3.8-27B at IQ4_XS quantization |
| Runtime | Custom llama.cpp fork and companion iOS app |
| Connection | 10 Gb/s USB-C |
One cable, two compute paths
During prompt processing, also called prefill, Backburner divides the model by layer. The Mac processes layers 1 through 40 for each 256-token microbatch, then sends the residual activations, the intermediate values passed between model layers, to the phone. The phone runs layers 41 through 64 on its GPU while the Mac begins the next microbatch.
This pipeline keeps both processors working and limits cable traffic to intermediate activations. The model weights remain split across the devices, with about 5.1GB of upper-layer weights stored on the phone.
Once the context exceeds the Mac’s memory budget, Backburner also moves the oldest key-value cache pages to the phone. A KV cache stores the attention data from earlier tokens so the model does not recompute them for every generated token. Backburner organizes those entries into pages containing 4,096 keys each.
For each attention step, the Mac sends the query tensor to the phone. The phone calculates attention against its cached keys and returns partial output, maximum and sum values, which the Mac merges with its local result. A custom GPU matrix kernel performs this work.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.