Liquid AI has announced the release of open d1 models, a new category of multimodal decision models specifically designed for edge computing. According to the Decision Index 0.2.1, the d1-3B model has emerged as the top performer in the sub-10B parameter category, outperforming larger models such as Decider 35B-A3B.
From Generation to Decision
Unlike traditional generative models that produce text token by token, the d1 models operate through a single forward pass. This architecture allows them to provide structured answers almost instantaneously, making them ideal for applications requiring speed and precision without the need for cloud-based processing.
The series includes two primary versions:
- d1-3B: Built on the LFM2.5-VL-3B backbone, supporting text and image inputs.
- d1-omni-600M: An early research release utilizing a bidirectional encoder that can process text, images, and audio.
Performance and Hardware
In collaboration with NVIDIA, Liquid AI evaluated the d1-3B across various hardware platforms. The results demonstrate exceptional speed: the model responds in just 16 ms on the NVIDIA Jetson AGX Thor and 50 ms on the more accessible Jetson Orin Nano. On GPUs like the RTX 4090, response times drop below 10 ms.
The models were benchmarked across seven public datasets, covering fields such as reading comprehension, toxicity detection, and medical QA. In these tests, d1-3B achieved a mean score of 82.9, the highest recorded in its class.