Guest on AI Engineer.

SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind
Aug 30, 2026 · 56:59
Google DeepMind's Dumitru Erhan, Shane Gu and Nicole Brichtova tell swyx that the new Nano Banana 2 Lite and Gemini Omni Flash APIs are steps toward world models, not just prettier videos. They see video models as zero-shot learners that should mature like language models, with understanding and generation unified once cost allows. Language is a lossy intermediary for audio, taste, smell and skin tone; a wine taster Shane consulted borrowed dating vocabulary to describe flavor. Dumitru says people preferred AI versions of real videos because they are sharper and more saturated, an 'Instagram filter', and warns of reward hacking like models adding wedding rings. Evaluation stays manual: Nicole describes ten-person side-by-side video comparisons and asks for real task data and FDE feedback.

Why ChatGPT Keeps Interrupting You — Dr. Tom Shapland, LiveKit
Jul 31, 2025 · 27:03
Dr. Tom Shapland of LiveKit explains why voice AI agents like ChatGPT's Advanced Voice Mode keep interrupting users: they rely on a simple VAD that triggers after a silence threshold, unlike humans who predict turn endings using semantics, syntax, and prosody. To solve this, LiveKit developed a semantic end-of-utterance model that considers the last four conversation turns, extending the VAD's silence window when the user is not done. A side-by-side demo shows dramatic reduction in interruptions. Shapland contrasts this with full-duplex models like Moshi and Meta's SyncLLM, which process input and generate speech simultaneously but lack instruction-following. He predicts commercial systems will improve via smarter VAD augmentations and faster cascade pipelines, not full-duplex. The talk also covers handling backchannels (e.g., 'mm-hmm'), the challenge of benchmarking turn-taking, and why OpenAI doesn't yet use LiveKit's model.
Powered by PodHood