model releaseLiquid Ai

Liquid AI releases open d1-3B decision model: 16 ms on Jetson AGX Thor, 48.57 on Decision Index

TL;DR

Liquid AI released two open-weight decision models, d1-3B (text and image) and the experimental d1-omni-600M (text with image or audio). Unlike generative models, they answer in a single forward pass, and Liquid AI claims d1-3B scores 48.57 on its Decision Index 0.2.1, ahead of all 4B and 9B models it tested.

3 min read
0

Liquid AI released two open-weight decision models on October 7, 2026: d1-3B, which takes text and images, and d1-omni-600M, an experimental model that takes text plus either an image or audio. Liquid AI claims d1-3B scores 48.57 on its own Decision Index 0.2.1, the best result under 10B parameters, and that it answers a single question in 16 ms on an NVIDIA Jetson AGX Thor.

What a decision model is

Decision models do not generate tokens. They answer structured questions (yes/no, multiple choice, scored criteria) in a single forward pass. Both models are built on Liquid AI's Liquid Foundation Models (LFMs), but from different backbones:

  • d1-3B is trained from LFM2.5-VL-3B, a decoder-only vision-language model. Inputs: text and images.
  • d1-omni-600M is trained from LFM2.5-Encoder-350M, a bidirectional encoder, with added vision and audio encoders. Inputs: text and image, or text and audio. Liquid AI labels it an early research release.

Benchmarks (company-reported)

On the Decision Index 0.2.1, Liquid AI says d1-3B scores 48.57, ahead of every 4B and 9B model it tested and of Decider 35B-A3B (47.11). The company ran the comparisons itself, and the index is its own benchmark.

On seven public datasets, d1-3B has a mean of 82.9 and d1-omni-600M has 78.4. The comparison models are Decider 2B (77.1) and Decider 4B (81.1).

Benchmark d1-omni-600M d1-3B Decider 2B Decider 4B
SQuAD 2.0 74.0 83.3 67.7 76.0
Civil Comments 95.8 93.3 93.6 92.8
MASSIVE intent 86.1 86.9 81.1 88.3
PubMedQA 61.3 68.3 65.7 63.3
BoolQ 77.7 86.3 87.3 89.0
XNLI 74.7 85.6 85.0 88.6
PAWS-X 79.5 76.4 59.5 69.8
Mean 78.4 82.9 77.1 81.1

d1-3B does not win every task. Decider 4B scores higher on MASSIVE intent, BoolQ and XNLI, and d1-omni-600M beats d1-3B on Civil Comments. Liquid AI says d1-omni-600M surpasses Decider 2B with a quarter of the parameters.

Liquid AI reports no vision or audio benchmark scores. It says it validated that d1-3B retains the vision capabilities of its backbone, and that audio decision benchmarks remain an open problem.

Latency

Liquid AI measured d1-3B with NVIDIA on several devices. Speed figures are single-question latency, and no speed numbers were published for d1-omni-600M.

Device One question 384px image 64 states, packed
Jetson AGX Thor 16 ms 35 ms 262/s
Jetson AGX Orin 64 GB 26 ms 83 ms 110/s
Jetson Orin Nano 50 ms 202 ms 38/s
Apple M5 Pro 30 ms 62 ms 78/s
NVIDIA RTX 4090 8 ms 17 ms 475/s
AMD MI325X 9 ms 18 ms 1,106/s

Processing a 3.4K-token state takes 220 ms on AGX Thor and 1,640 ms on Orin Nano. Three questions over one state cost about 1.3x the time of one on edge devices.

Availability

Both models are on Hugging Face as open weights, with a demo in the System One Arcade Space. They require transformers>=5.14 and trust_remote_code=True. The API exposes named questions of type yes/no, choice or score through system_one and system_one_batch. Context window, training cutoff and license terms were not stated in the announcement.

What this means

The design targets classification, routing and triage workloads, where a generative model's token-by-token decoding adds latency without adding value. Sub-50 ms answers on a Jetson Orin Nano make on-device decisions plausible without a cloud call.

The headline claims need outside validation. The Decision Index is Liquid AI's own benchmark, and the vision split of the newer v0.3 is private. Decider 4B still beats d1-3B on three of seven public tasks, so the advantage is not uniform. The absence of vision and audio scores also leaves multimodal quality unquantified. Teams should test on their own decision tasks before replacing existing classifiers.

Related Articles

model release

Musubi Releases Open-Weights PolicyLM-1.7B, a Sub-50ms Content Moderation Model

Musubi released PolicyLM-1.7B on Tuesday, a 1.7-billion-parameter open-weights decision model built for real-time content moderation. The company claims it applies plain-English policies to messages in under 50 milliseconds and needs no retraining when policies change.

model release

Google DeepMind releases EmbeddingGemma 2: 740M-parameter open embedding model spanning text, image, video, audio

Google DeepMind has released EmbeddingGemma 2, an Apache 2.0 open embedding model with 740M total parameters that maps text, images, video and audio into a single 768-dimensional vector space. According to the model card, it improves code retrieval on MTEB (code, v1) from 68.76 to 78.68 over its predecessor while keeping an 8,192-token context window.

model release

Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0

Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says it runs in ~191MB of active RAM for text-only weights and ~567MB for the full multimodal model on a quantized Pixel 11 Pro. Google also released a Mac app, AI Edge Foresight, to demonstrate it.

model release

Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0

Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Google also launched a Mac app, AI Edge Foresight, to demonstrate it.

Comments

Loading...