ByteDance releases Lance, 3B-parameter unified multimodal model handling image and video generation, editing, and unders
ByteDance has released Lance, a 3-billion parameter multimodal model that performs image and video generation, editing, and understanding within a single framework. The model was trained entirely from scratch using 128 A100 GPUs and achieves 84.67% on DPG-Bench and 74% on GenEval, competing with larger models despite its compact size.
ByteDance Releases Lance, 3B Unified Multimodal Model
ByteDance has released Lance, a 3-billion parameter model that handles text-to-image generation, text-to-video generation, image editing, video editing, and visual question answering in a single unified framework. The model was trained entirely from scratch using 128 A100 GPUs.
Technical Specifications
Lance operates with 3 billion active parameters and supports video generation up to 121 frames at 768×768 resolution (480p preset). According to ByteDance, the model uses flow matching scheduling with a default timestep shift of 3.5 and 30 denoising steps. The architecture requires at least 40GB VRAM for inference.
The model's training used a "staged multi-task recipe," though ByteDance has not disclosed the training dataset size, training duration, or data cutoff date. Pricing information has not been announced.
Benchmark Performance
On DPG-Bench, a comprehensive image generation evaluation, Lance scores 84.67% overall, with particularly strong performance in relation understanding (93.38%) and entity recognition (91.07%). The model trails larger unified models like TUNA-27B (86.54%) and InternVL-U (85.18%) but outperforms the 7B BAGEL model.
For GenEval, which tests compositional image generation across attributes like object count and spatial positioning, Lance achieves 74% overall. This matches SD3-Medium (2B parameters) but falls behind FLUX.1-dev's 75% (though FLUX.1-dev uses 12B parameters).
ByteDance reports competitive scores on specific categories: 99% for single-object generation, 94% for two-object generation, and 72% for counting accuracy.
Capabilities
The model handles six distinct task types through a unified interface: text-to-image, text-to-video, image editing, video editing, image understanding (visual question answering), and video understanding (video captioning and analysis). ByteDance demonstrates video understanding capabilities including counting actions, spatial reasoning, and temporal analysis.
For generation tasks, Lance uses classifier-free guidance with a default scale of 4.0 for text conditioning. The model supports multi-turn consistency editing, maintaining coherent changes across sequential edit operations.
Availability
Model weights are available on Hugging Face under the bytedance-research organization. ByteDance provides a command-line inference tool and Gradio interface. The system requires Python 3.10+ and CUDA 12.4+.
What This Means
Lance represents ByteDance's entry into unified multimodal AI, directly competing with models like DeepSeek-Janus, Show-o, and OmniGen. At 3B parameters, it's significantly smaller than most unified models while maintaining competitive performance on standard benchmarks. The efficiency suggests progress in model architecture design, though the lack of disclosed training details makes it difficult to assess reproducibility or training costs beyond the stated 128-GPU budget. The model's commercial viability will depend on pricing, which ByteDance has not yet announced.
Related Articles
Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text
Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.
Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter
Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.
China Telecom's Xing4.0-29B-A4B: 29B MoE, 4B Active, 256K Context, Trained Fully on Ascend NPUs
China Telecom AI's Xing4.0-29B-A4B (formerly the TeleChat line) is a mixture-of-experts model with 29B total and 4B active parameters and a native 256K context window, extensible to 512K. The company claims it is the first model of this scale trained entirely on Ascend NPUs with MindSpore. Community GGUF quantizations from Venastine-Research are already available.
Cloudflare releases Clef decision models, claims 39 ms median latency vs. 524 ms for TypeSafe's Jev
Cloudflare has released Clef and Clef-flash, two open-weight decision models that return probabilities over predefined answer options instead of generating text. The company claims median latencies of 39 ms and 209 ms, against just over 524 ms for TypeSafe AI's Jev. Both support text and images and are API-compatible with Jev.
Comments
Loading...