MiniMax H3 Becomes First Open Video Model to Top an AI Video Ranking
MiniMax has released open weights for H3, a 33-billion-parameter video model that ranks first in Video Editing and second in Text-to-Video on Artificial Analysis — the first time an open model has topped a video generation category. The model accepts text, images, video, and audio in a single prompt, though its highest-resolution module remains closed.
MiniMax has released open weights for H3, a 33-billion-parameter video generation model, making it the first open model to top a category on Artificial Analysis's video generation ranking. The model places first in Video Editing, second in Text-to-Video, and third in Image-to-Video, according to the benchmark site.
H3 is multimodal by design. It processes text, images, video, and audio together in a single request, generating clips between four and 15 seconds long with stereo sound. According to MiniMax's model card, a single prompt can combine up to nine reference images, three video clips, and three audio clips — a level of multi-input flexibility not typically available in open video models.
What's actually open — and what isn't
MiniMax did not release the full H3 stack. Two components remain closed: the 2K resolution module and H3-Context-IR, an internal system that translates prompts and reference material into a structured intermediate format before generation. Without H3-Context-IR, users running the model locally have to prepare that context manually, following prompting guides MiniMax has published separately.
The resolution gap is also notable in practice. Running H3 locally through ComfyUI currently tops out at 768p, well below the 2K output MiniMax's closed pipeline can produce. Users who want the model's full visual quality still depend on MiniMax's hosted service rather than the open weights.
What the open release does enable is fine-tuning. Developers can adapt H3 to custom footage, specific characters, or a particular visual style — something closed-model APIs generally don't allow. The license does carry a revenue cap, though: commercial use is only permitted for companies with less than $20 million in annual revenue, putting the model out of reach for larger commercial deployments without a separate agreement with MiniMax.
Pricing for MiniMax's hosted H3 service, including the closed 2K module, was not disclosed in the model card or accompanying materials.
Competing launch from ByteDance
The H3 release landed the same day ByteDance shipped Seedance 2.5, a closed video model that generates clips up to 30 seconds long with built-in audio — roughly double H3's maximum clip length. Seedance 2.5 is not available as open weights, and pricing details were not included in the source material.
What this means
H3's ranking result is the first concrete evidence that open-weight video models can compete with closed frontier systems on quality metrics, not just cost. But the release is only partially open: the highest-resolution module and the context-formatting system that likely contribute to H3's ranking performance stay behind MiniMax's hosted API. That split lets MiniMax claim an open-source milestone while preserving a commercial moat around the parts of the pipeline that matter most for output quality. For teams evaluating the model, the practical ceiling running it yourself is 768p — a meaningful gap from the 2K results driving its ranking placement. The $20 million revenue cap on commercial use further narrows who can actually build products on H3 without negotiating separately with MiniMax. Whether this becomes a template — rank-topping performance paired with partial openness — or a one-off will depend on how competitors like ByteDance and other Chinese labs respond in the coming months.
Related Articles
Anonymous Provider Launches Union Alpha, a Free 262K-Context Multimodal Model on OpenRouter
A third-party provider using the alias 'Stealth' has released Union Alpha on OpenRouter, a multimodal model with a 262K context window, currently free to use during its preview period. The model's developer remains anonymous, and OpenRouter states it is not the model's owner or operator.
InclusionAI Releases Ling 3.0 Flash VL, Adding Vision to Its 124B MoE Model
InclusionAI has released Ling 3.0 Flash VL, a vision-language extension of its 124B total-parameter, 5.5B active Mixture-of-Experts model. The model adds native image and video understanding, supports a 131K token context window, and is priced at $0.06 per 1M input tokens and $0.18 per 1M output tokens via OpenRouter.
OpenAI's GPT-6 Astra Beats Pokémon in 18 Hours, Scores 62.7% on ARC-AGI-3
GPT-6 Astra completed Pokémon FireRed in 18 hours 12 minutes, five times faster than its predecessor, and scored 62.7% on ARC-AGI-3 versus 7.78% for GPT-5.6 Sol. The model also ran a 141-hour Minecraft session and finished Fallout 3 in roughly 59 hours, according to independent testers.
Ex-OpenAI Researcher Launches Jev, an AI Model That Scores Options Instead of Generating Text
Startup TypeSafe AI has released Jev, a model built to score predefined answer options rather than generate text, claiming response times of 70 to 500 milliseconds. Co-founder Diogo Almeida, a former OpenAI researcher and InstructGPT co-author, says the model targets background classification tasks like sorting customer requests.
Comments
Loading...