Episodes from AI Engineer about Video Generation.

SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind
Aug 30, 2026 · 56:59
Google DeepMind's Dumitru Erhan, Shane Gu and Nicole Brichtova tell swyx that the new Nano Banana 2 Lite and Gemini Omni Flash APIs are steps toward world models, not just prettier videos. They see video models as zero-shot learners that should mature like language models, with understanding and generation unified once cost allows. Language is a lossy intermediary for audio, taste, smell and skin tone; a wine taster Shane consulted borrowed dating vocabulary to describe flavor. Dumitru says people preferred AI versions of real videos because they are sharper and more saturated, an 'Instagram filter', and warns of reward hacking like models adding wedding rings. Evaluation stays manual: Nicole describes ten-person side-by-side video comparisons and asks for real task data and FDE feedback.

The Next Medium: Why Real-Time Interactive Video Changes Everything — Ahmed Ahres, Reactor
Aug 18, 2026 · 17:30
Ahmed Ahres, head of go-to-market at Reactor, argues that world models are real-time interactive video, a change of medium rather than a speedup, because video becomes programmable like software. GPS enabling Uber and viewfinders enabling Instagram and TikTok show how real time unlocks new applications. Reactor's platform serves infinite interactive video (Helios from ByteDance), controllable worlds (Lingbot from Alibaba, LongLive 2 from Nvidia), and live avatars he admits are still uncracked. Users build interactive live streams, medical and cooking simulations, and video-to-video editing. Real-time infrastructure means streaming pixels, live sessions with memory, and sub-100-millisecond latency; he offers promo code AIE2026 for $75 credits.

Generative Video at the Speed of Light — Keegan McCallum, uRun
Aug 18, 2026 · 8:43
Keegan McCallum, founder of inference provider uRun, argues that generative video's interesting axis is no longer quality but efficiency and long-horizon generation. He shows Helios, a distillation of Wan 2.1 14b, generating clips in real time for ~one-hundredth the cost of a slower frontier-quality clip, and notes at least 40 real-time/long-horizon models released this year. Ten dollars buys three hours of continuous generative video, and fifty buys fifteen; this unlocks magic-mirror webcam transformations, visual mediums for people who don't think in text, and content creation where you steer a generation in under a second. The hard part is serving: global GPUs, WebRTC with ICE/TURN, and synchronized streaming pipelines; uRun is building a React component, Python runtime, and MCP/CLI.

Voice agents with Realtime Video — Sidney Primas, LemonSlice
Aug 18, 2026 · 26:36
LemonSlice CTO Sidney Primas explains how his startup builds real-time video avatars by pointing world models at humans. A Microsoft partnership put a Teddy Roosevelt avatar in a replica Oval Office, generating continuously for eight hours with no reset. He argues audio embeddings drive emotion, and because avatars only look backward, errors compound, so LemonSlice trains with an attention mask and collapses roughly 30 denoising steps to one. He says serving video costs about the same as a voice model, and that the model harness, which orchestrates GPU/CPU threads and queues to avoid stutters, holds much durable value. He predicts an emotion engine for better reactions and an end-to-end EQ layer within two or three years, taking user video and audio and outputting avatar video and audio.

Building an Agentic Video Editor for Mass Consumer — Ekaterina Deyneka, Reelful
Aug 18, 2026 · 12:45
Ekaterina Deyneka, founder and CEO of Reelful, presents her agentic video editor as structurally identical to an agentic app builder, with editing real footage harder than generation. At Reelful, users drop in media plus a prompt, and the agent understands the media, proposes a creative plan for approval, then spins up a remote sandbox where agent skills encode cut rules, font pairings, and b-roll generation. Video is composed as React code via Remotion, a verification layer catches composition errors and sends the agent back to reiterate, and the result renders to a polished clip. To hide this complexity from consumers, Reelful is mobile-first with directional templates and a building editor for tweaks. Deyneka demonstrated clips made purely by the agent and announced Reelful's funding from a16z Speedrun.

HTML Is All Agents Need — James Russo, HeyGen
Jul 21, 2026 · 15:13
James Russo, software engineer at HeyGen, argues that HTML, CSS, and JavaScript are the native languages of LLMs and all agents need to create great videos—leading to HyperFrames, an open-source framework that turns agent-authored HTML into deterministic MP4s. By letting the small Gemini 3 Flash model author code first, they ensured the thinnest wrapper won, with only a few data attributes for timing. Rendering works by freezing the browser clock and seeking frame by frame, ensuring everything loads before capturing each frame—enabling deterministic video from any web technology (Three.js, WebGL, Lottie). HyperFrames skills focus on taste and video craft rather than teaching frameworks, raising the floor for single-shot output. The framework has already rendered over 1.3 million videos in 90 days from 267,000 creators, with 15,000 daily renders and 32,000 GitHub stars—proving the approach at scale.
Powered by PodHood