Alibaba releases Qwen3.5-27B, a 27B multimodal model with Apache 2.0 license
Alibaba Qwen has released Qwen3.5-27B, a 27-billion parameter model capable of processing both images and text. The model is available under an Apache 2.0 open license and is compatible with standard transformer endpoints.
Qwen3.5-27B — Quick Specs
Alibaba Qwen Releases Qwen3.5-27B Multimodal Model
Alibaba's Qwen team has published Qwen3.5-27B, a 27-billion parameter model designed to handle both image and text inputs. The release marks the latest iteration in Alibaba's open-source model lineup.
Model Specifications
Qwen3.5-27B is a multimodal model with an architecture supporting image-text-to-text tasks. The model carries an Apache 2.0 license, making it freely available for both research and commercial use. It is compatible with standard transformer endpoints and follows the safetensors format for model weights.
The model's parameter count of 27 billion positions it in the mid-range segment—larger than models like Mistral 7B but smaller than many instruction-tuned variants in the 70B range. This size targets deployment scenarios where computational resources are constrained but model capability remains a priority.
Capability Profile
Qwen3.5-27B is tagged for conversational tasks and multimodal understanding, suggesting it can engage in dialogue while processing images alongside text prompts. The image-text-to-text classification indicates the model accepts images and text as combined inputs and generates text responses.
Specific benchmark scores, training data composition, knowledge cutoff date, and maximum context window length have not been disclosed in the initial release metadata.
Availability and Licensing
The model is hosted on Hugging Face and is immediately available for download. The Apache 2.0 license removes legal barriers to commercial deployment, distinguishing this release from many restricted-license models. Support for standard transformer inference frameworks means existing tooling can run the model without custom implementations.
No pricing information or commercial hosting details have been announced.
What This Means
Qwen3.5-27B expands Alibaba's open competition with other mid-range multimodal models like Qwen's own larger variants and offerings from Mistral, Meta, and others. The 27B parameter count targets developers who need multimodal capability without the compute overhead of 70B+ models. The Apache 2.0 license removes deployment friction compared to restricted models. However, without disclosed benchmarks or performance data, comparative positioning against competing 27-30B multimodal models remains unclear. Organizations evaluating this model should establish baselines on their specific use cases before production deployment.
Related Articles
Qwen3.8 Omni Flash: Alibaba's First Agentic Omni-Modal Model Adds Native Audio-Video Understanding, 1M Context
Alibaba's Qwen team has released Qwen3.8 Omni Flash, described as the first Qwen model built around agentic capabilities with native audio-video understanding. It ships with a 1M-token context window and support for two- and four-channel spatial audio.
Xiaomi Releases MiMo-V2.6-Pro-RL, a 1.02T-Parameter Omnimodal Model with 1M-Token Context
Xiaomi's MiMo team has released MiMo-V2.6-Pro-RL, a 1.02-trillion-parameter sparse mixture-of-experts model with 42B active parameters, 1M-token context, and native text/image/video/audio processing. The model was trained via a single mixed reinforcement learning run spanning coding, agentic, visual, and cybersecurity tasks, with benchmark scores that Xiaomi claims approach or match Claude Opus 5 and GPT-5.6 on several agentic and coding tests.
Xiaomi Releases MiMo-V2.6-Flash-RL, a 309B-Parameter MoE Model with 1M-Token Context and Native Omnimodal Support
Xiaomi's MiMo team released MiMo-V2.6-Flash-RL, an efficiency-tier checkpoint in the MiMo-V2.6 series featuring a 309B-parameter (15B active) Mixture-of-Experts architecture, 1M-token context, and native support for text, image, video, and audio. The model uses a single mixed reinforcement learning run across coding, agentic, visual, and cybersecurity tasks rather than domain-specific training.
Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output
Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.
Comments
Loading...