Liquid AI releases LFM2.5-VL-450M, improved 450M-parameter vision-language model with multilingual support
Liquid AI has released LFM2.5-VL-450M, a refreshed 450M-parameter vision-language model built on an updated LFM2.5-350M backbone. The model features a 32,768-token context window, supports 9 languages, handles native 512×512 pixel images, and adds bounding box prediction and function calling capabilities. Performance improvements span both vision and language benchmarks compared to its predecessor.
Liquid AI Releases LFM2.5-VL-450M Vision-Language Model
Liquid AI has released LFM2.5-VL-450M, a 450M-parameter vision-language model that serves as a refreshed version of LFM2-VL-450M. The model is built on an updated LFM2.5-350M language backbone and optimized for improved real-world performance.
Technical Specifications
LFM2.5-VL-450M features a 32,768-token context window and uses a SigLIP2 NaFlex vision encoder with 86M parameters. The model has a vocabulary size of 65,536 and supports 9 languages: English, Arabic, Chinese, French, German, Japanese, Korean, Portuguese, and Spanish.
The model processes native 512×512 pixel images without upscaling and preserves non-standard aspect ratios without distortion. For larger images, a tiling strategy splits them into non-overlapping 512×512 patches with thumbnail encoding for global context. Maximum image tokens range from 32 to 256, tunable at inference time without retraining.
Liquid AI provides the model in four formats: native Transformers/vLLM checkpoints, GGUF quantization for llama.cpp and CPU inference, ONNX Runtime format for cross-platform deployment, and MLX 4-bit through bf16 variants optimized for Apple Silicon.
Capabilities and Performance
New capabilities in LFM2.5-VL-450M include:
- Enhanced instruction following on vision and language tasks
- Multilingual vision understanding across all 9 supported languages
- Bounding box prediction and object detection for grounded visual understanding (measured on RefCOCO-M)
- Function calling support for text-only input (measured by BFCLv4)
On vision benchmarks, LFM2.5-VL-450M shows consistent improvements over LFM2-VL-450M:
- MMStar: 43.00 vs 40.87
- RealWorldQA: 58.43 vs 52.03
- MMBench (dev en): 60.91 vs 56.27
- MMVet: 41.10 vs 33.85
- RefCOCO-M (bounding box prediction, new capability): 81.28
Language benchmark improvements include:
- GPQA: 25.66 vs 23.13
- MMLU Pro: 19.32 vs 17.22
- IFEval: 61.16 vs 51.75
- BFCLv4 (function calling, new capability): 21.08
Liquid AI recommends the model for general vision-language workloads, captioning, and object detection. The company explicitly states it is not well-suited for knowledge-intensive tasks or fine-grained OCR.
Deployment Options
The model supports inference through Hugging Face Transformers (v5.1+), vLLM, SGLang, and llama.cpp. Liquid AI provides fine-tuning notebooks using both Unsloth and TRL frameworks with LoRA adapters.
A real-time video stream captioning WebGPU demo is available for browser-based testing. The model also offers function calling support for text-only inputs through a ChatML-like template format.
Monthly downloads reached 3,522 as of the release announcement. All vision benchmark scores were obtained using VLMEvalKit, with multilingual scores based on benchmarks translated by GPT-4o-mini.
What This Means
LFM2.5-VL-450M targets the efficient vision-language segment where inference cost and latency matter more than maximum capability. The 450M-parameter size and multiformat deployment options make it viable for edge, mobile, and resource-constrained environments. The addition of object detection and function calling expands use cases beyond pure captioning. Performance gains over the prior version suggest meaningful tuning improvements, though the model remains positioned below larger competitors for knowledge-intensive applications.
Related Articles
Liquid AI releases open d1-3B decision model: 16 ms on Jetson AGX Thor, 48.57 on Decision Index
Liquid AI released two open-weight decision models, d1-3B (text and image) and the experimental d1-omni-600M (text with image or audio). Unlike generative models, they answer in a single forward pass, and Liquid AI claims d1-3B scores 48.57 on its Decision Index 0.2.1, ahead of all 4B and 9B models it tested.
TII releases 1.6B Falcon-ASR, claims 20.92% Arabic WER against best listed 23.17%
The Technology Innovation Institute (TII) released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic and the Emirati dialect. TII claims a 20.92% average word error rate across six Arabic test sets, versus 23.17% for the next-best system on the leaderboard snapshot it used. A demo is live on Hugging Face. Pricing and API availability have not been disclosed.
Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0
Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Google also launched a Mac app, AI Edge Foresight, to demonstrate it.
Google releases Nano Banana 2.1 image model: $1.50/$30 per 1M tokens, 66K context
Google's Nano Banana 2.1 (Gemini Nano Banana 2.1) is an image generation and editing model on the Flash tier, listed on OpenRouter at $1.50 input and $30 output per 1M tokens with a 66K context window. It supports 1K, 2K, and 4K output and succeeds Nano Banana 2 and Nano Banana Pro, according to the listing.
Comments
Loading...