model releaseTencent

Tencent releases OmniWeaving, open-source video generation model with reasoning and multi-modal composition

TL;DR

Tencent's Hunyuan team released OmniWeaving on April 3, 2026, an open-source video generation model designed to compete with proprietary systems like Seedance-2.0. The model combines multimodal composition, reasoning-informed capabilities, and supports eight video generation tasks including text-to-video, image-to-video, video editing, and compositional generation.

3 min read
0

Tencent Releases OmniWeaving, Open-Source Video Generation Model

Tencent's Hunyuan team released OmniWeaving on April 3, 2026, positioning it as an open-source alternative to closed proprietary video generation systems. The model represents a significant step toward unified video generation capabilities, supporting eight distinct task configurations.

Architecture and Technical Foundation

OmniWeaving is built on HunyuanVideo-1.5 as its backbone, integrating an MLLM (Multimodal Large Language Model) + MMDiT (Multimodal Diffusion Transformer) + VAE framework. The architecture incorporates two key improvements:

Thinking Mode: The MLLM activates a reasoning mode that generates intermediate reasoning steps before video generation, translating abstract user intent into semantically precise prompts that condition the diffusion model.

Hidden States DeepStacking: Following mechanisms in Qwen3-VL, the model extracts hidden states from multiple intermediate MLLM layers, capturing semantic information across fine-grained details to high-level abstractions. These multi-level features are injected into the first three layers of the MMDiT conditioning branch.

Supported Tasks

OmniWeaving supports eight video generation configurations:

  • Text-to-Video: Generate videos from text prompts
  • First-Frame-to-Video: Animate static images with text guidance
  • Key-Frames-to-Video: Interpolate videos between start and end frames
  • Video-to-Video Editing: Instruction-based manipulation and stylization
  • Reference-to-Video: Single-subject reference-driven generation
  • Compositional Multi-Image-to-Video: Multi-subject generation from 2–4 images
  • Text-Image-Video-to-Video: Generation conditioned on combined text, image, and video inputs
  • Reasoning-Augmented Generation: Reasoning over user intent before video generation

The reasoning and composition tasks can be optionally enabled via a --think flag during inference.

Benchmarking

Tencent introduced IntelligentVBench, described as the first comprehensive benchmark for assessing unified video generation with reasoning capabilities. According to the team, OmniWeaving achieves state-of-the-art performance among open-source unified video generation models, though specific benchmark scores were not disclosed in the release announcement.

Availability and Deployment

Code and model weights were released on April 3, 2026. The model requires installation of attention libraries for optimized inference:

  • Flash Attention for faster inference and reduced GPU memory
  • Flex-Block-Attention for sparse attention optimization
  • SageAttention as an alternative optimization layer

The inference pipeline requires 8 GPUs by default but can be adapted for limited GPU memory environments through configuration adjustments and memory expansion settings. The codebase is available on GitHub with detailed checkpoint download instructions.

Research Background

OmniWeaving is the result of collaboration between Tencent Hunyuan, Zhejiang University, and Nanyang Technological University. The research was authored by Kaihang Pan, Qi Tian, and others, with the paper published on arXiv on March 26, 2026. The team trained the model on massive-scale pretraining datasets encompassing diverse compositional and reasoning-augmented scenarios.

What this means

OmniWeaving addresses a significant gap in open-source video generation by offering reasoning-aware composition capabilities previously limited to proprietary systems. The explicit integration of intermediate reasoning steps and multi-level semantic conditioning represents a technical approach to bridging user intent and pixel-level generation. For practitioners, this means access to a production-ready model supporting complex video generation workflows without closed-source dependencies. The IntelligentVBench benchmark provides a standardized evaluation framework for next-generation video models, though adoption depends on broader community adoption and reproducibility of claimed performance gains.

Related Articles

model release

Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta

Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

model release

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

model release

China Telecom's Xing4.0-29B-A4B: 29B MoE, 4B Active, 256K Context, Trained Fully on Ascend NPUs

China Telecom AI's Xing4.0-29B-A4B (formerly the TeleChat line) is a mixture-of-experts model with 29B total and 4B active parameters and a native 256K context window, extensible to 512K. The company claims it is the first model of this scale trained entirely on Ascend NPUs with MindSpore. Community GGUF quantizations from Venastine-Research are already available.

Comments

Loading...