Liquid AI has released LFM2.5-230M, a 230-million-parameter open-weight model built for on-device inference. The release includes support across llama.cpp, MLX, vLLM, SGLang and ONNX, broadening where the model can run.

The company positions the model for tool use and data extraction, with reported speeds of 213 tokens per second on a Galaxy S25 Ultra and 42 tokens per second on a Raspberry Pi 5. Those figures make it relevant for applications where latency, privacy or offline operation matter.

Small models remain a competitive category because many production tasks do not need frontier-scale systems. Deployment support may be as important as benchmark performance for adoption.