InclusionAI Releases Ling 3.0 Flash VL, Adding Vision to Its 124B MoE Model
InclusionAI has released Ling 3.0 Flash VL, a vision-language extension of its 124B total-parameter, 5.5B active Mixture-of-Experts model. The model adds native image and video understanding, supports a 131K token context window, and is priced at $0.06 per 1M input tokens and $0.18 per 1M output tokens via OpenRouter.
Ling 3.0 Flash VL — Quick Specs
InclusionAI Releases Ling 3.0 Flash VL
InclusionAI has released Ling 3.0 Flash VL, a vision-language variant of its Ling 3.0 Flash Mixture-of-Experts (MoE) model. The new model extends Ling 3.0 Flash's text capabilities with native visual perception, allowing it to process images and video alongside text input.
Architecture and Specs
Ling 3.0 Flash VL inherits the underlying MoE architecture from Ling 3.0 Flash: 124 billion total parameters with 5.5 billion active parameters per forward pass. This sparse activation design keeps inference costs low relative to the model's total parameter count while retaining capacity for complex reasoning and generation tasks.
Key specifications:
- Context window: 131,072 tokens (131K)
- Modality: text + image + video → text
- Pricing: $0.06 per 1M input tokens, $0.18 per 1M output tokens
- Availability: OpenRouter API, listed as
inclusionai/ling-3.0-flash-vl
According to InclusionAI, the model builds directly on Ling 3.0 Flash, adding what the company describes as "advanced visual" understanding on top of the base language model's existing capabilities. The source material does not specify exact benchmark scores for image or video comprehension tasks, so those claims remain unverified pending independent testing.
What's New Versus Ling 3.0 Flash
The base Ling 3.0 Flash model was a text-only MoE system. Ling 3.0 Flash VL's primary addition is multimodal input handling — specifically native visual perception for images and video — while InclusionAI states it also further strengthened the underlying language capabilities. No specific benchmark comparisons between the base and VL versions were disclosed.
Pricing in Context
At $0.06/M input and $0.18/M output tokens, Ling 3.0 Flash VL sits in the low-cost tier of multimodal models currently available on OpenRouter, considerably cheaper than flagship vision-language models from major labs. The combination of a sparse 5.5B active-parameter footprint and a 131K context window positions it as a candidate for high-volume multimodal inference where cost per token matters more than peak benchmark performance.
What This Means
Ling 3.0 Flash VL is InclusionAI's move to bring vision capabilities to its efficient MoE lineup without abandoning the cost advantages of sparse activation. The 5.5B active parameter count relative to 124B total parameters suggests the company is targeting inference efficiency — a relevant consideration for developers running multimodal workloads at scale, where per-token vision model costs can climb quickly.
The lack of disclosed benchmark scores for vision or video tasks means buyers should treat InclusionAI's capability claims as unverified until third-party evaluations appear. The $0.06/$0.18 per-1M-token pricing is aggressive for a multimodal model with a 131K context window, and if the model's actual visual reasoning quality holds up under scrutiny, it could become a viable low-cost alternative for applications like document analysis, video summarization, or image-grounded chat — provided independent testing confirms InclusionAI's positioning.
Related Articles
Alibaba's Qwen Releases Qwen-Drive-1.0-4B, a Unified VLM for Autonomous Driving Perception and Planning
Alibaba's Qwen team has released Qwen-Drive-1.0-4B, a 4B-parameter vision-language model built on Qwen3.5 that unifies 3D perception, driving question answering, and motion planning in one framework. The model reports strong open-loop, pseudo-closed-loop, and closed-loop driving benchmark results while claiming minimal loss of general vision-language ability.
DeepSeek Ships V4.1-Flash With Novel Encoder-Decoder Architecture, Cuts KV Cache to 1/8 of Predecessor
DeepSeek released V4.1-Flash, a 763B-parameter model built on a new causal encoder-decoder architecture that splits 8B active parameters for prefill and 16B for decode. The model adds native vision support, a 1M-token context window, and shrinks KV cache footprint to roughly 1/8 of DeepSeek V4 Flash, while retiring V4 Pro.
Inference.net Launches Schematron V2 Turbo, a 3B-Parameter Model for High-Volume HTML-to-JSON Extraction
Inference.net has released Schematron V2 Turbo, a 3-billion-parameter model built specifically for high-volume HTML-to-JSON extraction. The model supports a 128K context window and is priced at $0.03 per 1M input tokens and $0.15 per 1M output tokens.
Inference.net Releases Schematron V2 Small, a 3B-Parameter Model for HTML-to-JSON Extraction
Inference.net has released Schematron V2 Small, a 3B-parameter model specialized in converting HTML pages into structured JSON output. The model supports a 128K context window and requires extraction schemas to be passed via response_format rather than standard prompts.
Comments
Loading...