Liquid AI发布d1边缘多模态决策模型

Liquid AI发布d1边缘多模态决策模型。d1-3B在Jetson AGX Thor上单次问答仅需16毫秒,以48.57分登顶小参数基准测试,支持文本图像输入,专为边缘设备提供快速结构化决策能力。

边缘多模态开放决策模型 d1 发布

Liquid AI 于 2026 年 10 月 7 日正式推出其 d1 决策模型家族中的两款开源模型:d1-3B 和实验性的 d1-omni-600M。这两款模型专为边缘设备设计,旨在提供快速、结构化的多模态决策能力。

单次前向传播直接输出决策
  • 性能领先:在 Decision Index v0.2.1 基准测试中(参数规模小于 10B),d1-3B 以 48.57 的得分位居榜首,超越了所有 4B 和 9B 规模的模型,以及参数量更大的 Decider 35B-A3B(得分 47.11)。
  • 多模态支持:d1-3B 支持文本和图像输入;d1-omni-600M 则支持文本与图像,或文本与音频的组合输入。
  • 极致速度:d1-3B 在 NVIDIA Jetson AGX Thor 上单次问答仅需 16 毫秒,在 Jetson AGX Orin 上为 26 毫秒,在 Jetson Orin Nano 上为 50 毫秒。

技术架构:基于液体基础模型 (LFMs)

与传统的生成式模型不同,d1 系列决策模型不产生 Token,而是通过单次前向传播直接输出决策结果。这些模型构建在 Liquid AI 的液体基础模型 (LFMs) 之上,并采用了两种截然不同的骨干网络进行训练:

  • d1-3B:基于最新的仅解码器视觉语言模型 (VLM) 训练而成,能够同时处理文本和图像输入。
  • d1-omni-600M:基于 LFM2.5-Encoder-350M 双向编码器训练。该模型集成了视觉和音频编码器,从而支持三种模态(文本+图像 或 文本+音频)。目前该模型仍处于早期研究阶段,功能正在进一步开发中。

基准测试结果

团队在包括阅读理解、毒性检测、意图分类、医疗质量保证及跨语言理解在内的 7 个公共数据集上进行了评估。结果显示,d1-3B 的平均得分为 82.9,高于 Decider 4B;而 d1-omni-600M 以 78.4 的得分超过了仅有四分之一参数量的 Decider 2B(77.1)。

基准值 d1-omni-600M d1-3B Decider 2B Decider 4B
SQuAD 2.0 74.0 83.3 67.7 76.0
Civil Comments 95.8 93.3 93.6 92.8
Intent Classification 86.1 86.9 81.1 88.3
PubMedQA 61.3 68.3 65.7 63.3
BoolQ 77.7 86.3 87.3 89.0
XNLI 74.7 85.6 85.0 88.6
PAWS-X 79.5 76.4 59.5 69.8
Mean 78.4 82.9 77.1 81.1

验证表明,d1-3B 保留了其骨干模型 LFM2.5-VL-3B 的视觉能力,而 d1-omni-600M 则能处理全部三种模态。由于 Decision Index v0.3 尚未包含公开的视觉分割和音频决策基准,本次未报告相关具体数据。

推理速度评估

通过与 NVIDIA 合作,团队在多种硬件平台上对 d1-3B 进行了延迟测试。由于 d1-omni-600M 尚处于早期版本,本次未公布其速度数据。

边缘设备表现:在所有测试的边缘设备上,d1-3B 回答单个问题的时间均低于 50 毫秒。值得注意的是,批量处理三个问题所需的时间仅为单题时间的约 1.3 倍(例如在 AGX Thor 上从 16 毫秒增至 20 毫秒)。

设备 一个问题 三个问题 3.4K-代币状态 384px 图像 64 个国家/地区打包
Apple M5 Pro 30 ms 41 ms 640 ms 62 ms 78 / s
Jetson AGX Thor 16 ms 20 ms 220 ms 35 ms 262 / s
Jetson AGX Orin 64GB 26 ms 35 ms 560 ms 83 ms 110 / s
Jetson Orin Nano 50 ms 73 ms 1,640 ms 202 ms 38 / s

GPU 高性能计算表现:在独立 GPU 上,d1-3B 回答问题耗时不足 10 毫秒,处理 384px 图像耗时不足 18 毫秒。

设备 一个问题 三个问题 3.4K-代币状态 384px 图像 64 个国家/地区打包
NVIDIA RTX 4090 8 ms 21 ms 102 ms 17 ms 475 / s
AMD MI325X 9 ms 14 ms 44 ms 18 ms 1,106 / s

使用指南与代码示例

当应用场景需要快速的结构化决策且涉及多模态输入时,推荐选用 d1 系列模型。若追求最高决策质量,首选 d1-3B;若需极小的模型足迹,可选用 d1-omni-600M。

安装依赖项(要求 transformers >= 5.14):

pip install "transformers>=5.14" torch torchvision pillow

加载模型时需设置 trust_remote_code=True,因为模型包含自定义代码。以下是使用 d1-3B 进行文本和图像决策的代码示例:

import io
import urllib.request

import torch
from PIL import Image
from transformers import AutoModel

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True,
                                  dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)


questions = {
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
                          "fraud": "Suspected unauthorised use"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["Can wait", "Today", "Blocking the customer now"]},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))


url = "http://images.cocodataset.org/val2017/000000039769.jpg"  
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
                                       "criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
                       images=[photo]))


tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))

关于 d1-omni-600M 的具体使用说明,请参考其官方模型卡。

获取方式

这两款开源决策模型现已上线 Hugging Face Hub:

  • 下载链接:Hugging Face 上的 d1-3B 和 d1-omni-600M 仓库。
  • 在线演示:可在 Hugging Face Spaces 的 System One Arcade 中体验。

评论 0

0/500

评论需审核后展示,请文明发言

💬
还没有评论,来说两句

相关阅读