Multimodal open d1 decision models for the edge
Liquid AI has released two new open decision models, d1-3B and d1-omni-600M (experimental), as part of their d1 decision model family. These models are designed for edge applications and support text, vision, and audio. The d1-omni-600M model shows strong performance across various benchmarks, including SQuAD 2.0 (83.3), Civil Comments (95.8), MASSIVE intent (88.3), PubMedQA (68.3), BoolQ (89.0), XNLI (88.6), and PAWS-X (76.4 and 79.5), achieving a mean score of 82.9.
This release marks the first time Liquid AI has offered multimodal open decision models for edge devices, unlike previous models focused solely on text.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年10月7日 16:54 UTC
收录当时偏移:UTC+02026年10月7日 17:00 UTC
- 发布
- 2026年10月7日 16:54
- 收录
- 2026年10月7日 17:00
- 来源类型
- 官方发布
- 档位
- 当事方
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
Today, we release two open decision models in our d1 decision model family : d1-3B and d1-omni-600M (experimental).
- Best decision model under 10B on the Decision Index 0.2.1: d1-3B scores 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).
- Multimodal: d1-3B supports text and images, while d1-omni-600M supports text and images or text and audio
- Fast: d1-3B answers a question in 16 ms on an NVIDIA Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50ms on a Jetson Orin Nano
How we built decision models for the edge
These open d1 decision models are built on our Liquid Foundation Models (LFMs). Unlike our generative models, decision models don’t produce tokens but answer in a single forward pass.
d1-3B and d1-omni-600M are trained from two very different backbones:
- d1-3B is trained from LFM2.5-VL-3B , our latest VLM, which is decoder-only. It accepts text and images as inputs.
- d1-omni-600M is trained from LFM2.5-Encoder-350M , a bidirectional encoder. It adds vision and audio encoders to handle all three modalities. It accepts either text and image, or text and audio as inputs. This model is currently in an early research release and is undergoing further development.
Benchmark results
We benchmarked d1-3B and d1-omni-600M on seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding. d1-3B achieves a mean score of 82.9, the highest in the table and above Decider 4B. d1-omni-600M scores 78.4, surpassing Decider 2B (77.1) with only a quarter of the parameters.
Benchmark d1-omni-600M d1-3B Decider 2B Decider 4B
SQuAD 2.0 74.0 83.3 67.7 76.0
Civil Comments 95.8 93.3 93.6 92.8
MASSIVE intent 86.1 86.9 81.1 88.3
PubMedQA 61.3 68.3 65.7 63.3
BoolQ 77.7 86.3 87.3 89.0
XNLI 74.7 85.6 85.0 88.6
PAWS-X 79.5 76.4 59.5 69.8
Mean 78.4 82.9 77.1 81.1
We validated that d1-3B retains the vision capabilities of its LFM2.5-VL-3B backbone on standard vision benchmarks, and that d1-omni-600M handles all three modalities. We do not report any vision or audio benchmarks, as the Decision Index v0.3 includes only a private vision split and audio decision benchmarks are currently an open problem.
Speed
In collaboration with NVIDIA, we evaluated d1-3B on the NVIDIA stack across NVIDIA GeForce RTX 4090, NVIDIA Jetson AGX Thor, Jetson AGX Orin 64 GB, and Jetson Orin Nano. Since d1-omni-600M is an early research release, we don’t report any speed numbers for it in this release.
Edge inference. d1-3B answers a single question in under 50 ms on every measured device. Three questions take only 1.3x the time of one, with the AGX Thor going from 16 ms to 20 ms.