Inco AI's Splash 1.3.0 Cuts Mac Prefill Wait by 1.49x for Coding Agents

Splash 1.3.0 cuts time to first token on local Macs by offloading KV cache to SSD and splitting prefill between the GPU and Neural Engine.

·
·
·
Inco AI's Splash 1.3.0 Cuts Mac Prefill Wait by 1.49x for Coding Agents
  • Splash 1.3.0 ships Neural Engine prefill and SSD-tiered KV cache for Apple Silicon via GitHub
  • GPU plus Neural Engine FFN split cuts cold time to first token by 1.14x to 1.49x on Qwen3.8-27B
  • SSD offloading claims first-token drop from 19s to 1s when resuming a project on a 24 GB M6
  • Enabled by default on supported dense models; --disable-ane and --max-cache-disk control the behavior
  • Requires M3 or newer and macOS 26.4+; community fork adds M1/M2 support
  • Open source under Apache-2.0, install via brew install incoai/tap/splash

Splash 1.3.0 adds ANE prefill and SSD-backed KV cache

Inco AI has released Splash 1.3.0, targeting two common sources of latency in local coding agents. The Apple Neural Engine now shares long-prompt prefill work with the GPU, while an SSD tier preserves attention state when users switch between projects.

Splash is a Mac-only inference engine that serves a curated set of models through OpenAI- and Anthropic-compatible APIs. Inco builds model-specific runtimes with fused Metal kernels and a dedicated speculative-decoding draft for each checkpoint. Version 1.3.0 extends that optimization to hardware scheduling and cache management.

Prefill recruits the Neural Engine

Before generating an output token, a model must process the entire input and populate its attention cache. This prefill stage often dominates time to first token for coding agents because repository context can span thousands of tokens. On a 24 GB Mac running a 27B model, the release notes show prefill exceeding 45 seconds.

For supported dense models such as Qwen3.8-27B, Splash now divides feed-forward network layers between the GPU and Apple Neural Engine, or ANE. The engine calibrates the split for each Mac and model, identifies prompt sizes that benefit, and saves the result. The scheduler selects the fastest measured path automatically.

Short prompts, Qwen3.6-35B-A3B’s mixture-of-experts architecture, and token-by-token decoding remain on the GPU. The ANE path uses Hadamard rotation to distribute numerical outliers before W8A8 computation, which represents weights and activations with 8-bit values. It stages weights from the loaded model, making a second full-model int8 copy unnecessary.

Inco AI’s cold-prefill measurements for a 14K-token Qwen3.8-27B prompt
Mac Weights GPU only GPU + ANE Speedup
M5 Max MLX 4-bit 14.09 s 12.36 s 1.14×
M5 Pro MLX 4-bit 28.28 s 21.57 s 1.31×
M6, 24 GB UD-IQ3_XXS GGUF 45.77 s About 31 s 1.46–1.49×
M3 Max, Low Power MLX 4-bit 91.04 s 62.43 s 1.46×

Systems with slower GPU prefill show the largest relative gains in these measurements. Each row serves as its own comparison because the machines use different processors, weight formats, memory capacities, and power modes.

After an ANE evaluation failure, timeout, or non-finite result, Splash reruns the affected chunk on the GPU. Automatic calibration can also choose the GPU path when it performs better. Developers can force that path with splash serve --disable-ane.

KV cache moves to SSD

During inference, a model stores previously computed attention keys and values in a KV cache. Coding agents can reuse that state when prompts share a repository or conversation prefix, avoiding another full prefill. Limited unified memory often evicts the cache when a user changes projects, so Splash 1.3.0 can move it to SSD and restore it later.

The new implementation transfers evicted cache pages directly between shared-memory extents and disk through an I/O worker. Splash removed the earlier GPU staging ring, KV-copy kernel, and copy-only GPU commands, returning the staging ring’s memory to SSD-offload configurations.

Splash also compacts live KV pages into fewer extents, allowing empty allocations to return to macOS. That reduces pressure on the unified memory pool shared by the model, IDE, browser, and other applications.

Inco reports that restoring a cached project on a 24 GB M6 reduced first-token latency from roughly 19 seconds to 1 second. Results depend on the reusable prefix, available cache state, SSD performance, and memory pressure.

Set the disk budget

Homebrew installs the default build, and --max-cache-disk sets the maximum storage available to the on-disk KV cache:

ruby
brew install incoai/tap/splash
splash serve --model unsloth/Qwen3.8-27B-GGUF:UD-IQ3_XXS \
  --language-only --max-cache-disk 16G

# Run in another terminal
splash opencode

In this example, Splash serves the language-only model, permits up to 16 GB of disk-backed cache, and launches its OpenCode integration from a second terminal.

  • Hardware: The default build requires Apple Silicon M3 or newer.
  • Operating system: macOS 26.4 or later is required.
  • Memory: Inco recommends 36 GB of unified memory; 48 GB or more provides additional headroom.
  • Upgrade: Existing users should restart running servers because the native/server protocol changes from version 7 to 8.
  • License: Splash uses Apache-2.0 and is free to run locally.

A separate, unsupported community fork adds the GPU kernels needed for ANE splitting on M1 and M2 Macs. Its first-start calibration reportedly takes 35 to 85 seconds. On an M1 Max, the calibrated scheduler assigns half of the feed-forward work to the ANE and processes long prompts about 1.5× faster.

Runtime telemetry exposes the split

The /status endpoint now includes an ane_ffn object for inspecting scheduler behavior. It reports the current state, the share of work routed to the ANE, minimum chunk size, evaluation counts, timing, and GPU reruns after failures.

  • ANE state: Shows whether the split is calibrating, active, disabled, or unavailable.
  • Routed share: Indicates how much feed-forward work the scheduler assigns to the ANE.
  • Chunk threshold: Identifies the minimum workload that qualifies for splitting.
  • Timing and reruns: Exposes performance measurements and GPU fallback activity.

Warm contexts change the workflow

ANE scheduling and SSD-backed KV storage address separate delays in local agent use. The hardware split shortens prefill for qualifying long prompts, while disk restoration can skip most of that computation when a cached project returns. Both optimizations target workflows that repeatedly load large repository prefixes.

The boundaries remain concrete: short prompts, mixture-of-experts models, and decoding stay on the GPU, while SSD restoration requires a reusable prefix and retained cache state. Within those constraints, Splash 1.3.0 reduces waiting without reserving another full copy of the model in unified memory.

Trending
  • No trending articles

Comments

avatar

Next Reads