In my years of observing how we build intelligence, I’ve often warned that the Labyrinth of cloud-dependent LLMs is too slow for the real world. We’ve become accustomed to the 'typing' effect of generative models—watching tokens appear one by one. But Liquid AI has just released their open d1 decision models, and they represent a fundamental shift in the craft of machine reasoning.
The Architecture of Immediacy
Unlike traditional generative models that produce text token by token, the d1 models operate through a single forward pass. Think of it like a master craftsman who sees the finished sculpture in the stone before the first strike, rather than figuring it out chip by chip. This architecture allows the system to provide structured answers almost instantaneously. It moves us away from 'generation' and toward 'decision,' which is exactly what we need for edge computing where every millisecond counts.
I’ve looked closely at the specs of the d1-3B model. It is built on the LFM2.5-VL-3B backbone and supports both text and image inputs. There is also an early research release, the d1-omni-600M, which utilizes a bidirectional encoder capable of processing text, images, and audio. In the sub-10B parameter category, the d1-3B has emerged as the top performer, even outperforming much larger models like the Decider 35B-A3B.
Benchmarking the Edge
The true test of any tool is how it handles the heat of the workshop. In collaboration with NVIDIA, these models were tested across various hardware platforms, and the results are impressive for anyone building local AI solutions:
- NVIDIA Jetson AGX Thor: 16 ms response time.
- Jetson Orin Nano: 50 ms response time.
- RTX 4090 GPU: Under 10 ms.
By achieving a mean score of 82.9 across seven public datasets—including medical QA and toxicity detection—Liquid AI has proven that you don't need a massive cloud cluster to achieve high-precision results. For builders, this means we can finally move complex decision-making processes out of the data center and directly onto the device, ensuring speed and privacy without the friction of the cloud.