边缘多模态开放决策模型 d1 发布
Liquid AI 于 2026 年 10 月 7 日正式推出其 d1 决策模型家族中的两款开源模型:d1-3B 和实验性的 d1-omni-600M。这两款模型专为边缘设备设计,旨在提供快速、结构化的多模态决策能力。

- 性能领先:在 Decision Index v0.2.1 基准测试中(参数规模小于 10B),d1-3B 以 48.57 的得分位居榜首,超越了所有 4B 和 9B 规模的模型,以及参数量更大的 Decider 35B-A3B(得分 47.11)。
- 多模态支持:d1-3B 支持文本和图像输入;d1-omni-600M 则支持文本与图像,或文本与音频的组合输入。
- 极致速度:d1-3B 在 NVIDIA Jetson AGX Thor 上单次问答仅需 16 毫秒,在 Jetson AGX Orin 上为 26 毫秒,在 Jetson Orin Nano 上为 50 毫秒。
技术架构:基于液体基础模型 (LFMs)
与传统的生成式模型不同,d1 系列决策模型不产生 Token,而是通过单次前向传播直接输出决策结果。这些模型构建在 Liquid AI 的液体基础模型 (LFMs) 之上,并采用了两种截然不同的骨干网络进行训练:
- d1-3B:基于最新的仅解码器视觉语言模型 (VLM) 训练而成,能够同时处理文本和图像输入。
- d1-omni-600M:基于 LFM2.5-Encoder-350M 双向编码器训练。该模型集成了视觉和音频编码器,从而支持三种模态(文本+图像 或 文本+音频)。目前该模型仍处于早期研究阶段,功能正在进一步开发中。
基准测试结果
团队在包括阅读理解、毒性检测、意图分类、医疗质量保证及跨语言理解在内的 7 个公共数据集上进行了评估。结果显示,d1-3B 的平均得分为 82.9,高于 Decider 4B;而 d1-omni-600M 以 78.4 的得分超过了仅有四分之一参数量的 Decider 2B(77.1)。
| 基准值 | d1-omni-600M | d1-3B | Decider 2B | Decider 4B |
|---|---|---|---|---|
| SQuAD 2.0 | 74.0 | 83.3 | 67.7 | 76.0 |
| Civil Comments | 95.8 | 93.3 | 93.6 | 92.8 |
| Intent Classification | 86.1 | 86.9 | 81.1 | 88.3 |
| PubMedQA | 61.3 | 68.3 | 65.7 | 63.3 |
| BoolQ | 77.7 | 86.3 | 87.3 | 89.0 |
| XNLI | 74.7 | 85.6 | 85.0 | 88.6 |
| PAWS-X | 79.5 | 76.4 | 59.5 | 69.8 |
| Mean | 78.4 | 82.9 | 77.1 | 81.1 |
验证表明,d1-3B 保留了其骨干模型 LFM2.5-VL-3B 的视觉能力,而 d1-omni-600M 则能处理全部三种模态。由于 Decision Index v0.3 尚未包含公开的视觉分割和音频决策基准,本次未报告相关具体数据。
推理速度评估
通过与 NVIDIA 合作,团队在多种硬件平台上对 d1-3B 进行了延迟测试。由于 d1-omni-600M 尚处于早期版本,本次未公布其速度数据。
边缘设备表现:在所有测试的边缘设备上,d1-3B 回答单个问题的时间均低于 50 毫秒。值得注意的是,批量处理三个问题所需的时间仅为单题时间的约 1.3 倍(例如在 AGX Thor 上从 16 毫秒增至 20 毫秒)。
| 设备 | 一个问题 | 三个问题 | 3.4K-代币状态 | 384px 图像 | 64 个国家/地区打包 |
|---|---|---|---|---|---|
| Apple M5 Pro | 30 ms | 41 ms | 640 ms | 62 ms | 78 / s |
| Jetson AGX Thor | 16 ms | 20 ms | 220 ms | 35 ms | 262 / s |
| Jetson AGX Orin 64GB | 26 ms | 35 ms | 560 ms | 83 ms | 110 / s |
| Jetson Orin Nano | 50 ms | 73 ms | 1,640 ms | 202 ms | 38 / s |
GPU 高性能计算表现:在独立 GPU 上,d1-3B 回答问题耗时不足 10 毫秒,处理 384px 图像耗时不足 18 毫秒。
| 设备 | 一个问题 | 三个问题 | 3.4K-代币状态 | 384px 图像 | 64 个国家/地区打包 |
|---|---|---|---|---|---|
| NVIDIA RTX 4090 | 8 ms | 21 ms | 102 ms | 17 ms | 475 / s |
| AMD MI325X | 9 ms | 14 ms | 44 ms | 18 ms | 1,106 / s |
使用指南与代码示例
当应用场景需要快速的结构化决策且涉及多模态输入时,推荐选用 d1 系列模型。若追求最高决策质量,首选 d1-3B;若需极小的模型足迹,可选用 d1-omni-600M。
安装依赖项(要求 transformers >= 5.14):
pip install "transformers>=5.14" torch torchvision pillow
加载模型时需设置 trust_remote_code=True,因为模型包含自定义代码。以下是使用 d1-3B 进行文本和图像决策的代码示例:
import io
import urllib.request
import torch
from PIL import Image
from transformers import AutoModel
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True,
dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)
questions = {
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
"fraud": "Suspected unauthorised use"}},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["Can wait", "Today", "Blocking the customer now"]},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))
url = "http://images.cocodataset.org/val2017/000000039769.jpg"
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
"criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
images=[photo]))
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))
关于 d1-omni-600M 的具体使用说明,请参考其官方模型卡。
获取方式
这两款开源决策模型现已上线 Hugging Face Hub:
- 下载链接:Hugging Face 上的 d1-3B 和 d1-omni-600M 仓库。
- 在线演示:可在 Hugging Face Spaces 的 System One Arcade 中体验。





