Fal Launches H3 Max Live, a Post-Trained Minimax H3 Variant That Generates Video Faster Than Real Time
Fal released H3 Max Live, a post-trained and inference-optimized version of Minimax's H3 video model that Fal claims runs up to 35x faster than the official endpoint. The model generates video faster than it can be watched, enabling an infinite, chat-directed live video stream.
Fal has released H3 Max Live, a modified version of Minimax's H3 video generation model that the company claims generates video faster than it takes to watch it. The claim, if accurate, marks a threshold in generative video: sustained, real-time, interactive video generation rather than clip-by-clip rendering.
What Fal built
According to Fal, the company took Minimax's H3 model — released last month — and post-trained it for both cost and quality, then optimized it for Fal's proprietary inference engine. Fal claims this results in roughly 35x the speed of the official H3 endpoint. Separately, developer @levelsio cited a 50x speedup figure in a widely shared post, and described prior generation times of 2–5 minutes for 15 seconds of video being cut dramatically. Neither figure has been independently verified, and Fal has not published formal latency or throughput benchmarks alongside the release.
The product interface lets users type a prompt (!prompt) into a chat window, with the corresponding scene appearing on screen within seconds — described by Fal as an "infinite broadcast" where every frame is generated on the fly and every scene is directed by chat input.
Researcher Ethan Mollick was among the first to test the system independently, reporting that H3 Max produced "reasonably high quality" video in less time than it takes to watch it, using only the web interface, with generation time including prompt enhancement.
From demo to livestream
Fal engineer Rehan Sheikh connected the model to a Twitch livestream, branding it an "infinite interdimensional cable" channel. The stream drew over 5.7 million views before Twitch removed it. Indie developer @levelsio built a similar public tool, "Infinite Slop," allowing users to steer an ongoing AI-generated stream via chat. After being removed from Twitch and YouTube, Fal built its own hosting service to keep the live-generation stream running.
Quality caveats
Critics who watched the streams describe the output as low-plot, low-coherence "slop" — content generated frame-to-frame with no narrative structure and visibly RL-tuned imagery artifacts. No independent benchmark scores (FVD, human preference win rates, or standard video-quality metrics) have been published for H3 Max against the base Minimax H3 model or competing systems.
Pricing and access
Fal has not disclosed pricing per generation, per minute, or per token for H3 Max Live. Context window, parameter count, and training cutoff for the underlying Minimax H3 checkpoint used in this variant have not been published by either company.
What this means
The technical achievement here is inference speed, not necessarily a leap in visual fidelity or coherence — by most public accounts, output quality remains rough. But the throughput threshold matters: once generation outpaces playback, video stops being a batch-rendered asset and becomes a live, steerable medium, similar to how real-time text generation enabled conversational chat products. Expect competitors — including Minimax itself — to respond with their own low-latency inference stacks, and platform operators like Twitch and YouTube to face renewed pressure over how to moderate infinite, prompt-driven generative streams. The near-term risk is content quality and moderation, not capability; the underlying claim, that faster-than-real-time generation is achievable at all, is unlikely to be walked back regardless of how rough today's output looks.
Related Articles
AWS Details Reference Architecture for Multi-Tenant Document Chat on Amazon Bedrock Knowledge Bases
AWS has published a reference architecture showing how to build multi-tenant agentic document chat applications using Amazon Bedrock Managed Knowledge Base. The design handles per-user document isolation, asynchronous ingestion up to 50 MB, and agentic multi-hop retrieval with citations, offloading infrastructure work from development teams.
Perplexity Brings Agentic 'Personal Computer' Feature to Windows
Perplexity has expanded its agentic Personal Computer feature from Mac to Windows 10 and 11, letting subscribers on paid plans automate multi-step tasks across local files, native apps, and cloud services like OneDrive and Outlook.
OpenAI Quietly Rolls Out Outcome-Based Pricing, Charging Some Customers Only When Tasks Succeed
OpenAI has quietly begun offering some large customers a pay-per-outcome model, charging only when its AI successfully completes tasks such as customer support requests, according to The Information. The shift joins a broader industry move away from flat subscriptions toward usage- and results-based billing, led by startups like Sierra, Fin, and Cognition.
OpenAI to Cut Off Cursor's API Access After SpaceXAI Acquisition, Effective November 12, 2026
OpenAI announced it will stop providing its models to AI coding assistant Cursor on November 12, 2026, following Cursor's acquisition by Elon Musk's SpaceXAI. The company cited a lack of confidence that SpaceXAI would honor its terms of service, pointing to xAI's admitted use of OpenAI outputs to train competing models.
Comments
Loading...