product update

Fal Launches H3 Max Live, a Post-Trained Minimax H3 Variant That Generates Video Faster Than Real Time

TL;DR

Fal released H3 Max Live, a post-trained and inference-optimized version of Minimax's H3 video model that Fal claims runs up to 35x faster than the official endpoint. The model generates video faster than it can be watched, enabling an infinite, chat-directed live video stream.

3 min read
0

Fal has released H3 Max Live, a modified version of Minimax's H3 video generation model that the company claims generates video faster than it takes to watch it. The claim, if accurate, marks a threshold in generative video: sustained, real-time, interactive video generation rather than clip-by-clip rendering.

What Fal built

According to Fal, the company took Minimax's H3 model — released last month — and post-trained it for both cost and quality, then optimized it for Fal's proprietary inference engine. Fal claims this results in roughly 35x the speed of the official H3 endpoint. Separately, developer @levelsio cited a 50x speedup figure in a widely shared post, and described prior generation times of 2–5 minutes for 15 seconds of video being cut dramatically. Neither figure has been independently verified, and Fal has not published formal latency or throughput benchmarks alongside the release.

The product interface lets users type a prompt (!prompt) into a chat window, with the corresponding scene appearing on screen within seconds — described by Fal as an "infinite broadcast" where every frame is generated on the fly and every scene is directed by chat input.

Researcher Ethan Mollick was among the first to test the system independently, reporting that H3 Max produced "reasonably high quality" video in less time than it takes to watch it, using only the web interface, with generation time including prompt enhancement.

From demo to livestream

Fal engineer Rehan Sheikh connected the model to a Twitch livestream, branding it an "infinite interdimensional cable" channel. The stream drew over 5.7 million views before Twitch removed it. Indie developer @levelsio built a similar public tool, "Infinite Slop," allowing users to steer an ongoing AI-generated stream via chat. After being removed from Twitch and YouTube, Fal built its own hosting service to keep the live-generation stream running.

Quality caveats

Critics who watched the streams describe the output as low-plot, low-coherence "slop" — content generated frame-to-frame with no narrative structure and visibly RL-tuned imagery artifacts. No independent benchmark scores (FVD, human preference win rates, or standard video-quality metrics) have been published for H3 Max against the base Minimax H3 model or competing systems.

Pricing and access

Fal has not disclosed pricing per generation, per minute, or per token for H3 Max Live. Context window, parameter count, and training cutoff for the underlying Minimax H3 checkpoint used in this variant have not been published by either company.

What this means

The technical achievement here is inference speed, not necessarily a leap in visual fidelity or coherence — by most public accounts, output quality remains rough. But the throughput threshold matters: once generation outpaces playback, video stops being a batch-rendered asset and becomes a live, steerable medium, similar to how real-time text generation enabled conversational chat products. Expect competitors — including Minimax itself — to respond with their own low-latency inference stacks, and platform operators like Twitch and YouTube to face renewed pressure over how to moderate infinite, prompt-driven generative streams. The near-term risk is content quality and moderation, not capability; the underlying claim, that faster-than-real-time generation is achievable at all, is unlikely to be walked back regardless of how rough today's output looks.

Related Articles

product update

ElevenLabs Launches Music v2.5, Adds API Access and Free Tier for AI-Generated Songs

ElevenLabs has released Music v2.5, an updated version of its ElevenMusic generator, now available through both the app and API. The company says blind testing with nearly 48,000 comparison pairs showed listeners preferred v2.5 over the prior version, particularly for R&B, Hip-Hop, and orchestral genres.

product update

Perplexity Says It Runs End-to-End Engineering Systems on OpenAI's GPT-6 Astra

Perplexity says it has shifted core engineering workflows, including code changes and production monitoring, onto OpenAI's GPT-6 Astra model. The claim comes from an OpenAI-published case study with no independent benchmark data released.

product update

Perplexity Deploys OpenAI's Astra Model for Autonomous Code and Systems Management

Perplexity is using an OpenAI model referred to as Astra to handle software changes, communications, and production monitoring with less frequent human check-ins. OpenAI published the case study; specific model specs and benchmarks have not been disclosed.

product update

Augment Code Claims 4.5x Developer Output Increase From Internal 'Software Factory' of AI Agents

Augment Code says its internal 'software factory'—a network of specialized agents built on its Cosmos platform—drove a 4.5x increase in size-adjusted developer output and cut median PR merge time from 11.2 to 3.1 hours over nine months. The company frames this as evidence that once AI writes nearly all new code, the bottleneck shifts to review, verification, and incident response.

Comments

Loading...