
AI Video Can Beat Playback Speed Under the Right Conditions. The Price Tag Is the Story.
An accelerated MiniMax H3 variant called H3 Max can generate some clips faster than they play. Under favorable loads, buffered interactive AI livestreams are already running. Here is how the trick works, what happened on Twitch and Kick, and what a nonstop stream actually costs.
AI-assisted draft. Reviewed and edited by the Phosphene team before publication.
On August 29, 2026, Pieter Levels shared a short demo on X that most people in AI video immediately understood. A model called H3 Max, built by fal.ai on top of MiniMax's H3, produced 15 seconds of video in 9 seconds. Watch time slower than generation time. That single crossover is what makes infinite AI livestreams possible, and within days there were several of them running.
The concept sounds magical and the reality is a lot more constrained. Both halves are worth understanding, because real-time generation is going to change how video tools are built, not just how streams are run.
What H3 Max actually is
H3 Max is not a new model from scratch. It is MiniMax H3, a large open video model from the Chinese company MiniMax, given extra fine-tuning and, more importantly, heavy inference optimization on fal.ai's own servers. The claim attached to it is 50 times faster than the original at the same quality, which is the kind of number you should read as marketing until you look at the concrete measurements.
The concrete numbers: five seconds of 768p video in under three seconds, and a 15-second clip in around 15 seconds under typical conditions. Speed varies with settings and server load, so the headline "faster than real time" only holds for targeted demo loads. What they have actually crossed is a different threshold. When generating a short clip takes less time than watching it, a pipeline can start producing clip number two while the audience is still looking at clip number one. The stream becomes a chain of segments instead of a queue of finished files.
An audience that writes the script
On fal.live, fal.ai's own platform, that chain is already public. Several channels run continuously with no preset storyline. Viewers type directions into a chat, the model generates a short clip and tries to stitch it onto the previous one, and a voting system lets the audience pick which proposal gets generated next. The result is collective improvisation: a scene can hold a consistent set for a few segments, then suddenly switch characters or style because the chat decided to.
The seams are visible. Requests arrive with a lag, because several seconds of video have to be generated and stitched before anything new appears. Faces change between segments, sets dissolve mid-scene, and the continuity is best described as generous. That is not a bug report on fal.ai specifically. It is the current state of the format: you are watching short-generation cells glued together live, and the glue is the part nobody has solved yet.
The Twitch and Kick episode
The idea did not start on fal.live. Rehan Sheikh, an engineer at fal.ai, first connected H3 Max to a Twitch stream inspired by the interdimensional television from Rick and Morty, with viewers deciding what came next from the chat. Twitch banned it. The experiment moved to Kick and was banned there too. No official reason was given, but the likeliest culprit is straightforward: licensed characters on a monetized stream is the kind of thing platforms and rights holders actively police.
fal.ai responded by shipping its own platform, which has the neat property of not depending on Twitch or Kick at all. Pieter Levels, after sharing the demo, launched Infinite Slop on the same day with fal.ai as the sponsor, and claims 37,000 people watched it during the first day.
That sponsorship detail does not appear in the marketing. It matters a lot.
The math that nobody posts
A developer named Pascal Lindenau ran the numbers on continuous generation. At fal.ai's promotional API pricing, keeping a 768p stream going costs roughly 144 dollars per hour. Leave a stream running 24/7 for a year and you pass 1.2 million dollars. The standard, non-promotional pricing would roughly double that estimate.
Pieter Levels is not paying that bill because fal.ai is covering part of his compute. The platform streams can therefore continue indefinitely in a technical sense, but they are running on sponsorship, which means the format currently exists in a sweet spot between two economic facts: generating video continuously is far too expensive for an individual, and watching an AI-generated novelty stream is worth enough attention to a model vendor that funding one or two flagship channels is a reasonable marketing expense.
What this means for creators
First, the interactive format itself is new creative territory. An audience that votes on the next scene turns a stream into a game, and the early channels are crude versions of something that will get much better at holding continuity. Watch the stitching problem: the moment segments remember prior context well enough to keep characters and sets stable, interactive AI video stops being a curiosity and starts being a real production format.
Second, the platform lesson is already written. If your stream uses characters people recognize, expect bans regardless of how transformative you think the format is. The safer direction is original IP and audience-controlled storytelling, which is exactly where the format's novelty actually lives.
Third, treat every speed claim as a test case, not a spec. The 15-seconds-in-9-seconds demo and the 5-seconds-in-3-seconds claim are both real measurements of the same system under different loads. Real-time AI video is a matter of thresholds: generation that beats watching time enables streaming, and the gap between the demo and the sustained load is where the infrastructure bill lives.
The infinite AI livestream is here, and it is powered by sponsored compute, stitched clips, and an audience that cannot stop voting. That is not a criticism. It is the most honest description of a format that internet culture was obviously going to invent the first week the math allowed it.