Liquid AI has released two open models built to make decisions rather than generate text token by token. The 3-billion-parameter d1-3B accepts text and images, while the experimental 600-million-parameter d1-omni handles text paired with either images or audio.

Because the models return a classification in one forward pass, they are designed for low-latency work on local devices. Liquid reports that d1-3B answered a benchmark question in 16 milliseconds on Nvidia’s Jetson AGX Thor, 26 milliseconds on an AGX Orin and 50 milliseconds on an Orin Nano. Across seven public datasets, the company reports a mean score of 82.9 for d1-3B and 78.4 for d1-omni.

The evaluations cover reading comprehension, toxicity, intent classification, medical questions and cross-language understanding. Those are vendor-reported benchmark results, not proof of performance in every deployment. The smaller omni model is also explicitly an early research release. Still, the release offers developers compact, downloadable models for routing, moderation and other tasks where a quick label matters more than a conversational answer.