Liquid AI's LFM2.5-230M Beats Models Twice Its Size Running on a Raspberry Pi
Liquid AI's 230M-parameter LFM2.5 runs at 213 tok/s on a Galaxy S25 Ultra and beats models twice its size on tool use and data extraction

- New model: Liquid AI releases LFM2.5-230M, their smallest model yet, targeting edge and on-device agentic workloads.
- Speed: 213 tok/s on Galaxy S25 Ultra (CPU) and 42 tok/s on Raspberry Pi 5 — fastest in its class.
- Punches above weight: Beats models 2x+ its size on instruction following (IFEval: 71.71), tool use (BFCLv3: 43.26), and data extraction.
- Architecture: Hybrid of 8 gated short convolution blocks + 6 GQA attention blocks; pre-trained on 19T tokens with 32K context.
- Robot demo: Deployed on a Unitree G1 humanoid running on NVIDIA Jetson Orin, acting as a natural-language skill-selection layer.
- Availability: Free open-weight model on Hugging Face, with support for llama.cpp, MLX, vLLM, SGLang, and ONNX.
Liquid AI just pushed the floor on what a useful language model can look like. LFM2.5-230M is their smallest model yet: 230 million parameters that can run on a Raspberry Pi, a flagship Android phone, or a humanoid robot, all while outperforming models more than twice its size on the tasks that actually matter for agents and pipelines.
A different kind of small model
Most sub-500M models are curiosities. They're fine for demos but fall apart the moment you ask them to follow structured instructions, call tools reliably, or extract data from messy text. LFM2.5-230M competes with and often beats models more than twice as large, spanning instruction following, data extraction, and tool use. That's the headline claim, and the benchmarks back it up in the areas that count for agentic workloads.
The model scores 71.71 on IFEval (instruction following) and 43.26 on BFCLv3 (function calling), beating IBM's Granite 4.0-350M (53.48 and 39.58 respectively) and Gemma 3 1B (63.49 and 16.61) despite being a smaller model. On CaseReportBench, a data extraction benchmark, it scores 22.51 versus Gemma 3 1B's 2.28. These aren't marginal wins.
Where it does fall short is anywhere that requires deep reasoning. Given its compact size, Liquid does not recommend it for reasoning-heavy workloads such as advanced math, code generation, or creative writing. On MMLU-Pro, Qwen3.5-0.8B (a model more than 3x larger) scores 37.42 versus LFM2.5-230M's 20.25. The model knows what it is.
The architecture behind the speed
The speed story starts with the LFM2 architecture, which is neither a standard transformer nor a pure state-space model (SSM) like Mamba. The LFM2 architecture is a hybrid design that combines gated short convolutions for local sequence mixing and a minority of Grouped Query Attention (GQA) blocks for long-range token interaction. For the 230M variant specifically, the model has 14 layers: 8 double-gated LIV convolution blocks and 6 GQA blocks.
The key insight from Liquid's architecture search is counterintuitive. In the on-device regime, most of the benefits attributed to recent hybrid SSM and linear-attention blocks can be captured by their short convolutional submodules plus a small number of global attention layers. In other words, the field has been adding complexity where simplicity works better. Using hardware-in-the-loop architecture search under edge latency and memory constraints, Liquid obtained a compact hybrid backbone that delivers up to 2x faster prefill and decode on CPUs compared to similarly sized models.
The practical result: LFM2.5-230M delivers 213 tok/s decode speed on a Galaxy S25 Ultra (CPU) and 42 tok/s on a Raspberry Pi 5 (CPU), with the highest prefill and decode throughput in its class while keeping the smallest memory footprint.
How it was trained
The model was pre-trained for 19T tokens, including a 32K context extension phase. Post-training follows a three-stage recipe designed to preserve fine-tuning flexibility:
- Supervised fine-tuning with distillation from LFM2.5-350M , the larger sibling acts as a teacher, compressing its behavior into the smaller model
- Direct Preference Optimization (DPO) , aligning outputs with human preferences without full RL overhead
- Multi-domain reinforcement learning , sharpening performance on tool use and structured output tasks across domains
The final checkpoint balances strong out-of-the-box capabilities with adaptability to downstream specialization, while remaining competitive with larger models. The base model (LFM2.5-230M-Base) is also released separately for teams that want to fine-tune from scratch.
A robot demo that's more than a stunt
As an early look at ongoing work, Liquid deployed LFM2.5-230M on a Unitree G1 humanoid robot, running entirely on-device on its onboard NVIDIA Jetson Orin. The model acts as a skill-selection layer: it takes a natural-language command and decomposes it into a sequence of tool calls that invoke pre-trained low-level motor skills from NVIDIA's SONIC framework.
The demo is deliberately simple , timed walking, velocity targets, a one-legged kneel , but the architecture is the point. A 230M model, after a quick fine-tune, becomes the natural-language interface for a humanoid. No cloud round-trip. No GPU rack. Just a Jetson Orin and a model small enough to fit in it.
What you can actually build with it
The two use cases Liquid keeps coming back to are large-scale data extraction pipelines and lightweight on-device agentic workloads. Concretely, that means:
- Structured data extraction at scale , running millions of documents through a pipeline where cost and latency matter more than reasoning depth
- On-device agents on phones , tool-calling assistants that work offline, with no API latency or privacy exposure
- Robotics and automation , natural-language control layers on embedded hardware like Jetson Orin
- Home and network automation , always-on agents on low-power devices that respond to commands locally
- Fine-tuned domain specialists , the base model is a clean starting point for narrow tasks where 230M parameters is more than enough
The broader race this fits into
The edge race is a fight over latency, memory, privacy, battery, and distribution , and LFM2.5 sits directly inside that fight. Gemma, Phi, Llama, Qwen, SmolLM, Granite, Nemotron, Apple Silicon, Copilot+ PCs, Android AICore, Jetson Thor, and DGX Spark are all converging on local inference as a first-class platform. Every major lab has a sub-1B story now.
What makes LFM2.5-230M interesting in that context is the architectural bet. Liquid AI's position is that although SSMs can theoretically handle long dependencies in linear time, short convolutions plus a small number of attention blocks are more practical on edge hardware. The empirical results from their architecture search suggest that the community has been over-engineering the problem. The winning design isn't Mamba, isn't a transformer, and isn't a complex hybrid , it's the simplest thing that actually runs fast on the hardware that exists today.
The uncomfortable truth for frontier labs is that the edge does not need a perfect model , it needs a model good enough to own the first interaction, cheap enough to run constantly, private enough to trust with user context, and portable enough to ship everywhere. At 230M parameters and 42 tok/s on a $80 single-board computer, LFM2.5-230M is a credible answer to that brief.
Getting started
Both LFM2.5-230M and LFM2.5-230M-Base are available now on Hugging Face as open-weight models. The model is free to download, fine-tune, and deploy under Liquid's LFM license. Inference is supported across:
- llama.cpp (GGUF) , for edge and CPU inference
- MLX , optimized for Apple Silicon
- vLLM and SGLang , GPU serving for production deployments
- ONNX , cross-platform, including NPUs
- Transformers , requires
transformers>=5.0.0
The quickest path to running it locally with the Transformers library:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"LiquidAI/LFM2.5-230M",
device_map="auto",
dtype="bfloat16",
)
tokenizer = AutoTokenizer.from_pretrained("LiquidAI/LFM2.5-230M")
input_ids = tokenizer.apply_chat_template(
[{"role": "user", "content": "Extract the company name and revenue from this text: ..."}],
add_generation_prompt=True,
return_tensors="pt",
tokenize=True,
)["input_ids"].to(model.device)
output = model.generate(input_ids, temperature=0.1, top_k=50, max_new_tokens=256)
print(tokenizer.decode(output[0][input_ids.shape[-1]:]))For fine-tuning, Liquid provides Colab notebooks covering SFT with LoRA via Unsloth or TRL, DPO, and GRPO , covering most of the standard post-training workflows out of the box.