# SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind

AI Engineer · 2026-08-30

<https://aiengineer.podhood.com/7b6cec67-6b9b-463f-8549-11cee1d334a1>

Google DeepMind's Dumitru Erhan, Shane Gu and Nicole Brichtova tell swyx that the new Nano Banana 2 Lite and Gemini Omni Flash APIs are steps toward world models, not just prettier videos. They see video models as zero-shot learners that should mature like language models, with understanding and generation unified once cost allows. Language is a lossy intermediary for audio, taste, smell and skin tone; a wine taster Shane consulted borrowed dating vocabulary to describe flavor. Dumitru says people preferred AI versions of real videos because they are sharper and more saturated, an 'Instagram filter', and warns of reward hacking like models adding wedding rings. Evaluation stays manual: Nicole describes ten-person side-by-side video comparisons and asks for real task data and FDE feedback.

## Questions this episode answers

### What did Google DeepMind launch this week for generative media?

Nicole Brichtova said they launched Nano Banana 2 Lite, their fastest and cheapest image model, which beats the original Nano Banana, and the Gemini Omni Flash APIs for video generation and editing, priced the same as Vio31 Fast. She noted three-second latency unlocks faster ideation and iteration.

[2:04](https://aiengineer.podhood.com/7b6cec67-6b9b-463f-8549-11cee1d334a1?t=124000)

### Why do AI-generated videos look better than real videos?

swyx described an experiment where they captioned real videos and generated equivalent versions with Omni; humans largely preferred the AI-generated videos. He explained the output is not more realistic but sharper, more HDR, and with better skin tone, called human preference an unreliable barometer, and named the default aesthetic the Instagram filter; Nicole Brichtova agreed.

[33:55](https://aiengineer.podhood.com/7b6cec67-6b9b-463f-8549-11cee1d334a1?t=2035000)

### Why do generated images add wedding rings to hands?

swyx said their image model tends to put wedding rings on hands, a pattern he had not noticed until someone else pointed it out. Shane Gu called it very common reward hacking, saying preference-based training can produce such artifacts in many weird ways.

[40:58](https://aiengineer.podhood.com/7b6cec67-6b9b-463f-8549-11cee1d334a1?t=2458000)

### Is language a good enough intermediate representation for video generation?

Shane Gu called language an extremely helpful representation because it conditions causal information and lets you draw on pretraining, but he also argued audio, taste, smell, and skin color are poorly verbalized because they sit close to survival. Nicole Brichtova agreed language is limiting for sensory experience, so video and language should be complementary foundation models.

[14:05](https://aiengineer.podhood.com/7b6cec67-6b9b-463f-8549-11cee1d334a1?t=845000)

## Key moments

- **[0:00] Intro**
- **[1:57] Launches**
  - [2:04] Nano Banana 2 Lite is Google's fastest and cheapest image model, and it beats the original Nano Banana, says Nicole Brichtova.
  - [3:11] Google finally launched Gemini Omni Flash APIs for developers, priced the same as Veo 3.1 Fast, says Nicole Brichtova.
- **[4:40] Use cases**
  - [4:59] Omni can take any input — storyboard images, an audio voice reference — and output a video, says Nicole Brichtova.
  - [7:00] Shane Gu used Nano Banana to translate a gadget manual into Romanian for his parents, keeping the diagrams identical.
- **[7:46] Video agents**
  - [8:25] Shane Gu says video models and symbolic language models should work together, with language conditioning causal information.
  - [9:58] Shane Gu's paper 'Video Models: Zero-Shot Learners and Reasoners' argues video models are foundational space-time models with physical intuitions.
- **[11:12] One model**
  - [12:16] Shane Gu predicts a single unified model in five years, but says multiple specialized models will remain for six months.
- **[14:05] Representations**
  - [14:20] Q: Is captioning the right intermediate representation for video — or should it be code or raw bits? Shane Gu says text is how humans communicate.
  - [17:15] Shane Gu vs RL maximalists: reasoning is not just additional compute — language's chain-of-thought taps pre-training intelligence.
  - [17:58] Shane Gu defines a world model as the model in model-based RL, citing Jitendra Malik, Schmidhuber and Fei-Fei Li.
- **[19:46] Career paths**
  - [22:27] Shane Gu compares today's video models to early GPT-2: creative demos that will become reliable world models with better instruction following.
- **[24:20] Audio**
  - [25:13] Shane Gu says Veo 3 generated audio and video jointly from one causal latent, solving lip-sync that hacked pipelines failed at.
  - [26:41] Shane Gu: language lacks words for smell, taste, skin tone and audio nuance — a professional wine taster uses dating language to describe wine.
  - [31:11] Swyx: AI video sounds studio-recorded because studio recordings are its training data, so it lacks real room sound and distance.
- **[33:55] Better than real**
  - [34:06] In a human eval, viewers preferred AI-generated versions of real videos — sharper, more HDR, better skin tone, says swyx.
  - [35:10] Shane Gu: a manga artist finds AI images creepy because the eye gaze is subtly wrong and unnatural.
  - [37:22] Nicole Brichtova says average viewers prefer AI's overly smooth, saturated content — swyx dubs it 'the Instagram filter.'
  - [38:31] Nicole Brichtova: Nano Banana Pro's default aesthetic made cluttered infographics go viral, showing model defaults shape output.
  - [41:03] Swyx: Google's image model began adding wedding rings to every hand; Shane Gu calls it reward hacking.
- **[41:33] Evals**
  - [43:04] Nicole Brichtova: when two video models are close, Google picks by putting 10 people in a room comparing clips side-by-side.
- **[46:39] Data**
  - [48:13] Shane Gu says Omni wants professional video data, not random YouTube clips — high-quality footage is key.
  - [48:38] Nicole Brichtova wants real-world task trajectories, like product photo to full ad campaign, to train Omni.
  - [51:21] Shane Gu: '99% of information is inside people. You can only extract it through active dialogue and befriending them.'
- **[52:51] FDEs**
  - [53:21] Shane Gu says FDEs are post-training, not sales — they feed customer insights upstream to improve frontier models.
  - [54:38] Nicole Brichtova wants feedback from niche users like interior designers whose custom rug sizes break image models.
  - [55:22] Shane Gu: a brand language like IKEA's is not just blue and yellow — models must capture the exact shade of blue.

## Speakers

- **swyx** (host)
- **Dumitru Erhan** (guest)
- **Nicole Brichtova** (guest)
- **Shane** (guest)

## Topics

Video Generation, Multimodal Models, Vision Language Models

## Mentioned

Amazon (company), Google DeepMind (company), Replicate (company), YouTube (company), Gemini (product), Nano Banana (product), Omni (product), Vio (product)

## Transcript

### Intro

**swyx** [0:13]
And welcome back, for those on the stream and those in person. We tend to basically take these longer sessions between all the main stage keynotes to reflect on things that are particularly important but don't have a significant launch moment.

Today we're very lucky to have people working on Omni, and Vio, and Nano Banana, the world's best generative models, here with us. Dumitru, I first saw you when you were posting about your office.

I think you're probably Google's number one office influencer, at least in San Francisco. I think you like to bike as well. You like to take photos of things.

**Dumitru Erhan** [0:57]
I bike here.

**swyx** [0:58]
Yeah. But also, you work on video models.

**Dumitru Erhan** [1:02]
That'sright.

**swyx** [1:03]
Shane, I met you, I think, at a dinner.

**Dumitru Erhan** [1:06]
Yeah.

**swyx** [1:10]
And I remember you were trying to get me invested in one of the companies. I forget which one.

**Dumitru Erhan** [1:15]
Forget about that.

**swyx** [1:18]
But

now you're working on Omni Thinking and just a bunch of other things.

**Dumitru Erhan** [1:25]
Gemini, Gemini RL, things like that.

**swyx** [1:26]
Yeah, yeah. And Nicole, also the rest of the gen media models, Nano Banana, and everything you just launched, actually, even this week.

**Nicole Brichtova** [1:36]
Yeah, we launched some APIs.

**swyx** [1:37]
Yeah, yeah, yeah.

**Nicole Brichtova** [1:39]
And I haven't tried to convince you to invest in anything, but maybe I should.

**swyx** [1:42]
I mean, so I try not to be an investor. People just convince me anyway. I'm like, just, OK, well, I'm not that rich, but you can't not try to invest in some of these things. And for those of us who are not working at a frontier lab, this is the best, this is the closest they will ever get.

### Launches

**swyx** [1:57]
So actually, let's kind of recap, since you're closest to it and we just did it. What was launched this week? What should people go try out?

**Nicole Brichtova** [2:04]
Yeah. So yesterday, we had two launch moments. One of them, we launched Nano Banana 2 Lite, which is our fastest, cheapest image model in the Nano Banana model family. And it's better than the original Nano Banana. So really, for most people, that model replaces what you used and loved the original Nano Banana for across generation and editing.

And it gets really close to the frontier quality of the mainland bigger models. So that's really exciting. I think if you look at some of the demos or things that people have been trying, getting kind of that three-second latency just unlocks a whole bunch of things that you can do with ideation and iteration.

And it's just really fun. And the models get until a point where the quality is really good, where you can use it for iteration, but you can also use some of those outputs as just kind of ready production outputs.

So that's really exciting. And then second launch, we finally launched the Gemini Omni Flash APIs that we pre-announced at I/O. So thank you for waiting. And that is the first time that we're making the APIs available for developers.

And it's basically really exciting kind of video generation and editing. And we're pricing it the same as Vio31 Fast. So we're getting you kind of really, really good quality for a really awesome price, hopefully.

**swyx** [3:22]
Yeah, I mean, that's incredible. I actually really so when you guys launched Omni for the first time, you also did a podcast with Logan, who couldn't be here today. And you edited Gusloff, and Ramen, and all these things.

I actually really want to do that to our videos. I just didn't have an API for it, because obviously, I have to automate the whole thing. So thank you for the API.

**Nicole Brichtova** [3:42]
That is my favorite use case. Everybody should do that. I got a cat, which is probably the most boring of the animals. If you don't know what we're talking about, you should look it up. It's very funny. Fofer, who's on the team, did that.

**swyx** [3:55]
Fofer is the number one guy you should follow.

**Nicole Brichtova** [3:57]
You should follow.

**swyx** [3:58]
Get ideas on, OK, what can this thing do?

**Nicole Brichtova** [4:00]
Yes.

**swyx** [4:01]
Right?

**Nicole Brichtova** [4:01]
Yes. He's amazing at that.

**swyx** [4:03]
I've tried to get him for the last two years to come to AIE. He hasn't made it yet. He's actually come in person. He just didn't want to speak, because he's anonymous.

**Nicole Brichtova** [4:11]
I know.

**swyx** [4:12]
I want to say his real name, but I can't say his real name.

**Nicole Brichtova** [4:14]
No, we won't do that to him. But you should really follow him. He's amazing. He did all that work.

**Dumitru Erhan** [4:19]
I actually met him in the office when we did the podcast, I think. And I didn't realize it was him, because his badge doesn't say Fofer.

**swyx** [4:28]
Yeah, I know. So he used to be part of Replicate. And Replicate had this joke where everyone was DeepFakes. DeepFakes is this kind of mysterious character in Replicate. Replicate is a very cool company, and Fofer was part of it.

So, OK, one thing I wanted to get on there before I go into the Omni proper is we added cats. We added sloughs. Very cool, very cute, very fun. What are the inspire people as to what are the more workhorse use cases that maybe are not just demos?

### Use cases

**Nicole Brichtova** [4:59]
Yeah, so obviously, the hero capability of the model or maybe there's two. One is the ability to kind of take in anything as input and then get video on the other side. Obviously, in the future, and we've kind of talked about this as a pre-announced, we want to get the other output modalities out as well.

But basically, what that means is you can take a set of images that you have as maybe a storyboard. You can take an audio track as a reference of a voice that you want a character to speak. And then you can get a video on the other side.

So that just unlocks a whole bunch of things that you can do in short film production or shorts we've launched on YouTube as well to help creators kind of create content more easily. And then the other one is obviously video editing.

That's another thing that we're really excited about, that we're just making easier, because now you can use natural language to take a video, add something, remove something. Slough is obviously a fun example. But there's obviously kind of consumer use cases that we kind of had in mind, where you could take your beach vacation video that was too noisy, and you want to clean up that noise.

Maybe in the past, you wouldn't have, because you didn't have the tools or you didn't know what the tools were that you needed to go to. So that's one use case that you can go to. We've seen a lot of folks use it for kind of marketing, ad campaign creation.

And I'm excited to see more of those use cases as we launch the APIs, because obviously, we don't see all of it in the first-party products. But I'm really excited for people to start to explore that in the API.

So those are just some of the kind of high-level things that have come up. People also use it to create education materials.

**swyx** [6:31]
Yes.

**Nicole Brichtova** [6:32]
And that's really exciting. I think we've all kind of talked about being excited about the future of education, where everything can be kind of customized to you and personalized to your knowledge level and the style that you prefer.

And so this is kind of just a step in that direction.

**Dumitru Erhan** [6:48]
Yeah, I sort of actually used just Nano Banana yesterday. My parents are visiting, and there was a very fun use case. I bought some gadget off Amazon that they wanted. And the instructions to use it were only in English.

And there was plenty of diagrams or whatever. And I took a picture of it and said, translate this into Romanian.

**swyx** [7:05]
Yes.

**Dumitru Erhan** [7:05]
And keep everything else the same. So it was amazing. It was just like, yeah, it looks identical. And it's perfectly translated, more or less. But it's using Gemini under the hood, obviously, to kind of do the translation. So you can see this use case for video as well.

The power of text rendering in Omni is quite next level. And you could think about plenty of use cases of both text rendering, translation, internalization, all sorts of things that would be actually genuinely useful to a lot of different people and sort of broader access to either you could re-dub a video or whatever it is that you wanted to do.

There's plenty of different things that you could think about doing.

**swyx** [7:46]
Yeah. One of the most enlightening conversations I have on my podcast is with just researchers at the frontier of these things. I had one with Ethan from the XAI video team, the Grok video team, who was basically saying the next trend is actually not just a single model.

### Video agents

**swyx** [8:04]
It's more like video agents. And I don't know if that terminology resonates. Obviously, very relevant for RL. But it was basically kind of like giving up on trying to do everything in effectively one pass. Do you feel that same way?

Or is it still an open research question, which way the trends are going?

**Shane** [8:25]
Yeah, so what kind of excites me most is really when the symbolic kind of foundational models and this kind of video foundational model can actually kind of really work together. And in a way, if you look at the beginning of the generative image generation, video generation, a lot of it kind of started when the language model got good enough to provide a very detailed captioning, like from Stable Diffusion days or kind of DALL·E 2 days.

So basically, language is an extremely helpful representation. One is that it's kind of universal. But the other kind of more technical thing, my hypothesis is one really difficult thing about machine learning is this sort of experience correlation. So you don't know if this kind of feature that's kind of predictive is actually a causal factor or not.

There are two ways. One is we can have really diverse data, training data, like from every intervention of the causal graph. The other is you condition the causal information. And conditioning the language is kind of like conditioning a causal information of the kind of world.

**swyx** [9:29]
Which is a prompt or a concept?

**Shane** [9:31]
Yeah. Yeah, exactly. So if you look at how you're going to describe this video, how you're going to describe this kind of image, is actually very close to how we discover this kind of causality behind this, like how this is kind of generated.

So one is that can really allow for very rich generalization and then very kind of just like a good model. The other is, so eight months ago, we put an evaluation paper called "Video Models: Zero-Shot Learners and Reasoners."

**swyx** [9:58]
Yes.

**Shane** [9:59]
So that was kind of it's a kind of fun paper. And then later on, actually, the Nano Banana team fought up with the Vision Banana paper that basically used Nano Banana to do. But essentially, the idea is a video model is an extremely good sort of foundational model for space and time kind of information.

So classic computer vision tasks, a lot of it could be kind of zero-shotted. And when you're, say, feeding someone a visual quiz, there's definitely a lot to improve. It can kind of solve. And it can like robotics kind of seeing.

It has really good kind of physical intuitions, like word model. And I think that the key is really the kind of mix of the visual kind of reasoning and then the text kind of reasoning, kind of all tied together.

Obviously, whether doing it as kind of a unified model versus this kind of agent kind of constraint, I think that's more like it's going to be more kind of incremental. I imagine everything is going to go into a single model eventually.

**swyx** [10:59]
Yeah.

**Shane** [10:59]
Butright now, there's a lot you can do if you basically take very good video understanding, image understanding Gemini agentically with this kind of Omni. And that's actually kind of, yeah, our team is exploring a lot.

### One model

**swyx** [11:12]
Yeah, OK, there's a lot in there. I think one question I am increasingly starting to wonder is, does it all trend towards one product for you guys? Now you have multiple models out. The naming of Omni does imply that eventually, everything will go away and it just goes into Omnies.

Is that the plan?

**Dumitru Erhan** [11:35]
Is it? I don't know. I think

maybe. I mean, I think eventually. I think there's sort of different trade-offs, engineering, research, product trade-offs. It's for the same reason that the sorry, how is it called? Nano Banana Lite? I don't know what the product name is.

**Nicole Brichtova** [11:53]
Nano Banana 2 Lite.

**Dumitru Erhan** [11:54]
Nano Banana 2 Lite, yeah. It serves a particular niche. And it probably doesn't necessarily fit immediately in the same model literally checkpoint as something that can do 4K 30-second videos. They're probably not trainable in the same quite way.

So I don't know. It depends on how far into the future you look like. Sure, in five years from now, will they all be the same model? Probably. But in six months from now, we'll probably still have multiple different models doing different things, because kind of pragmatically, the trade-offs are such that we should have multiple different kinds of models.

**swyx** [12:35]
Yeah.

**Nicole Brichtova** [12:36]
I think that'sright. And just on that note, I mean, we did call it Gemini Omni because we wanted to hint at the future where Gemini just becomes fully multimodal in and out. And so it's definitely a move in that direction.

I think we'll probably see a move in the direction where Omni also generates images and edits images and all those kinds of things. But do me sorry that I think on the way there, there is a bunch of really, really useful applications of some of these more specialized models.

And so we will probably continue to work on those as well, because that serves a certain need at this point in time that may not exist a year from now.

**Dumitru Erhan** [13:10]
There's also a research question about just how much transfer there is between different kinds of modalities. I think you may believe that there's some transfer between coding and video generation. And I think most people don't necessarily believe that.

But you could try to think that there is something there. Or it could be a waste to put them together, to try to learn both tasks at the same time. So I think it's an interesting sort of question, to which extent image and video, obviously, kind of there's some transfer, kind of not that different.

There's value in learning to output video and audio at the same time, because joint audio-visual, that's how it is. And then there's other kind of intersections of modalities that are not super obvious, like 3D representation, coding, I don't know, maybe.

Things like that. So I think it's worth sort of exploring the different corners there. And we are actively doing that with a focus towards what people actually want to do with these models.

### Representations

**swyx** [14:05]
Yeah. One thing I feel like I'm surprised by, but also I feel like it's insufficiently answered, is what is the correct intermediate representation? So captioning,right? XAI does captioning. Omni does captioning. And I understand how captioning works for images.

And I understand that you can extend it into video and sort of guide it across time. It just feels very inefficient. I feel like there should be something better. Maybe it's code. And obviously, I think a lot of FFmpeg and Matplot what's the three blue, one brown one?

Manin? A lot of video is generated through code. And maybe that's the optimal representation. Any hypothesis as to is it better, or is it just English all you need?

**Shane** [15:01]
Well, so I'm into Gemini. And we do a lot of RL agents and, of course, kind of coding. So yeah, we're definitely exploring the coding representations as kind of better kind of way to represent.

**swyx** [15:13]
But

what's your probability estimate on if we just output binaries? It's just ones and zeros.

**Shane** [15:25]
I guess maybe a kind of similar discussion was

basically, is the language theright representation?

**swyx** [15:33]
Yes.

**Shane** [15:34]
So one kind of question, for example, professor, someone asked is, why does the chain of thought need to be in the natural language?

**swyx** [15:42]
Yes.

**Shane** [15:43]
Can it just be kind of any kind of continuous tokens, just any amount of additional computations? So one is, obviously, the adaptive compute is going to give better results. So it's that. But what really kind of made chain of thought so four years ago, I wrote the large model zero-shot reasoner and then self-improvement.

So I kind of know from very early days. But the reason it works really well is,right now, the recipe that works is the pre-training that scales a lot. And then that basically learns other intelligence. There are a lot of scaling RL.

But those are still extremely kind of compute-intensive to extract the information. And you really want to rely the intelligence on that. So basically, by tying the sort of reasoning in the natural language, you basically directly use the intelligence of the pre-training to it.

Why should you remove that kind of constraint staying in and out? And these days, I feel a lot of advancements in the text, but also in this kind of multimodal space, is really driven by this kind of text as a kind of great sort of representation.

**swyx** [16:55]
Yeah. It's a good backbone.

**Shane** [16:56]
Yeah.

**Dumitru Erhan** [16:57]
I think to me, it's even simpler than that. It's text is how we communicate. So I think fundamentally, if you're building kind of products that humans will be interfacing with, we will be using text somehow if it's a text interface, not for everything.

So I think it's natural to default to that.

**Shane** [17:15]
Yeah. Obviously, there's a kind of few discussion. Some RL maximalists, it's like, oh, we don't care about kind of channel that's kind of like stuff. It's just additional compute.

**Dumitru Erhan** [17:25]
Sure.

**Shane** [17:25]
But I personally, yeah.

**swyx** [17:27]
RL maximalists. I wonder who qualifies in that description.

**Dumitru Erhan** [17:32]
David Silver.

**swyx** [17:33]
Ah, OK. Yeah, I mean, they've just left to start their thing. Interesting. OK, so I think I'm very interested in just better representations, because I think that's one of our themes that we're curating today at the World's Fair is world models.

You mentioned the word world models. But it's not something that's super well-defined. I think everyone's kind of converging on some version of it that it's the ideal.

**Dumitru Erhan** [17:58]
Sure. Everything is a world model now. It's sort of.

**swyx** [18:01]
Yeah, it's not that useful,right?

**Shane** [18:03]
So I just gave a keynote at the ICAERS world model workshop. And then, yeah, essentially, I definitely encourage to check out the definition by Jitendra Madik. He's like the OG computer vision professor at UC Berkeley. He has a bit of work to say about world model.

But also kind of Schmidt-Herbert, how he defined the world model from 2019, not like 1990, sort of when it was just basically just that kind of model base. For me, the world model is basically just a model in the model-based RL.

And I feel that has sufficient to describe. But obviously, there are a lot of Fei-Fei had a kind of nice blog post about.

**swyx** [18:39]
Spatial intelligence.

**Shane** [18:40]
Yeah, just kind of broken down. But yeah.

**swyx** [18:44]
Yeah, I mean, so

I'll end this part of the conversation. But I do think that language, to me, relying on language as the narrow pipe through which everything goes through still is like a lossy compression.

**Shane** [18:58]
No, no, no. But we've not seen that. We're basically seeing the video model and the language together. So I think the language alone is not sufficient. That's why we feel like the video is a very complementary to the model.

Right now, the kind of VO Omni, many people kind of feel as generating kind of pretty videos. But I think our vision, it's much more than that. It's a missing foundational model that's absolutely required if you want to make the AGI that matches the humans, not just the jagged one.

**swyx** [19:25]
Yeah. OK, so one other thing. You mentioned on the vision side. And I'm kind of curious how sort of parallel, in terms of your research careers, this development is. I think basically, a lot of vision people have crossed over into world model people.

A lot of vision people also become generative video and image people. And is it just as simple as reversing image to text and then now it's text to image?

### Career paths

**swyx** [20:00]
I mean, that effectively was the diffusion process.

I just see the career paths of the people that I talk to. And I see this overall trend of research directions. And I just wanted you guys to sort of reflect on that.

**Dumitru Erhan** [20:16]
I mean, I certainly went that way. I started a long time ago doing computer vision, sort of object detection, recognition, things like that. I think that's just a simpler problem. Just generation is just harder. It's a different kind of mapping.

The inverse mapping is not as simple as just inverting the kind of procedures. It's more ambiguous to go from cat to image of a cat.

**swyx** [20:41]
And in some ways, it's also a loop, because your vision work creates the synthetic labels that then continues.

**Dumitru Erhan** [20:47]
I mean.

**swyx** [20:47]
Sure.

**Dumitru Erhan** [20:49]
I don't know.

**swyx** [20:49]
I don't know. I'm trying to validate my sort of theories about how fields develop, how careers progress through this.

**Nicole Brichtova** [20:55]
I mean, the better the understanding side gets, we have seen that the generation side also gets better.

**swyx** [21:02]
It's completely bootstrapping. Yeah.

**Nicole Brichtova** [21:06]
And so there's definitely a layer there to that thesis. And I think, yeah, I think a lot of people have kind of I definitely worked with a lot of image understanding people who became image generation people. And then some of them have moved on to video, because it's kind of like the next thing where you have so many more dimensions to work with.

So yeah, I'm curious about you, because you're.

**Shane** [21:25]
So I definitely recommend to start with understanding, recognition, because that's basically discriminator. And then that's going to lead to better generation. And that's what the bridge is basically reinforcement learning. So my kind of journey is, I initially kind of worked on the algorithmic research in the Gemini model, against some like MNIST kind of generation.

And then I worked on RL and robotics. And then six years ago, I was leading a moon shot on the dexterity. It was pretty early. But I see now everyone's kind of doing it. Four years ago, I basically kind of figured out that this like symbolic AGI is going to accelerate much faster than the kind of physical AGI kind of counterpart.

So I decided to kind of like language models and then those things. And then recently, kind of work with Doomie and then Omni team. I quite enjoy kind of collaboration there. What I quite enjoy, what I recommend definitely to the researcher is to definitely kind of explore or at least get exposure to what the top people in each of the community are looking at, how they kind of think about problems.

So when I look at the video model, to me, it kind of reminds me pretty early on sort of like language model, where very early language model was a kind of creative sort of demo. You kind of try to write a story, like a novel, and then you write in GPT-2.

And then those kind of days, like LSTM kind of days. And then instruction tuning, you actually kind of make a user as a chatbot. But then at the chatbot stage, it still had so much hallucinations. And the instruction tuning wasn't good enough.

So it couldn't use for reasoning. And when it got good enough in pre-training and post-training for reasoning, then this kind of test time scaling and RL really took off, rather than many of the kind of best performing models.

Andright now, I think the video model is, as we mentioned, it is a complementary foundational model. And I can imagine it's going to follow a similar path. It's going to be very it's going to improve a lot the instruction following a lot of this.

It's going to improve a lot in reducing hallucinations. To the extent that you become a very reliable world model, so you can kind of intermix the video space-time simulation with a text simulation to solve arbitrary AGI problems.

**Dumitru Erhan** [23:34]
Also, I think the difference still is between sort of text models and image video models is that we haven't quite unified understanding and generation in multimedia, I'd say, yet. I mean, I think without going into the details, of course, it depends on at which level you're thinking about this.

But generally, there's not that many, as far as I know, models, SOTA kind of frontier models that are genuinely kind of good at both understanding and generation of, let's say, videos. It's an interesting challenge. I'm not saying that we should do this.

But I think it kind of stands to reason that understanding and generation are two sides of the same coin. So they kind of should be in the same model in some ways. But we don't necessarily always do that.

**swyx** [24:20]
Yeah. You mentioned audio as well,right?

### Audio

**Dumitru Erhan** [24:23]
Yeah.

**swyx** [24:24]
Is that as hard as video? Or qualitatively different? If so, in what way? What are the interesting directions? Three years ago, was people using, I guess, diffusion to do audio, as in the sort of refusion approach. I don't know if you guys saw that.

And I just think it's very interesting if a modality that we perceive, which is audio is different than video, actually to machines is exactly the same. They see no difference.

**Dumitru Erhan** [24:59]
I mean, I think on a technical level, there are some differences. But I think they're relatively minor. I think from my perspective, audio came into my life when we shipped VO3, which was, I believe, the first model that did like a joint.

**swyx** [25:13]
With the slicing of the.

**Dumitru Erhan** [25:14]
Yeah, yeah, slicing of gold bars or whatever. It was the first model that did sort of joint audio-visual generation.

**swyx** [25:20]
Yes.

**Dumitru Erhan** [25:22]
I mean, there were other models that did kind of agentic hacking under the hood. But this one was truly sort of generating everything at once. And the reason we did that is because we felt, and I think it was theright choice, we felt that it only makes sense to generate them at the same time, because there's sort of kind of like from a machine learning perspective, there's one latent kind of causal kind of generative process.

There's something that generates you speaking. It's not the pixels and then the audio are somehow generated by some other process. The lips have to move in sync with the audio. So I think that solved a lot of the issues that previous models had, or the way that people did video generation before, where it was like, OK, we generate the pixels.

And then we're going to hack something on top of it that moves the lips with the audio that we generate. And that was very bad. And so I think that was, to me, that's the I mean, after VO3, people were like, what do you mean there's no audio in your model?

That makes no sense. Once it's there, you have to have it. So I think that was theright choice. And doing it to one single generative model, I think, was theright choice.

**Shane** [26:29]
One thing I kind of want to also kind of ask you guys kind of opinion as well. One difference I find the audio and then against the image and video is like the audio information is less verbalized. I mean, of course, the TTS and stuff is trivial.

But when you get around to say, how to describe music, how do you describe this person's tone kind of pitch? I feel the sort of verbalization is insufficient. And the interesting thing is that you kind of see that in two other things, like taste, taste sense, and also, say, smell.

And then another interesting thing is the skin color. So skin color, the language is pretty limited to describe the skin color. And the reason is that we're extremely sensitive to the small difference, perturbations on the skin color, because that basically shows us, is this person going to kill me?

Or can I befriend this person? Kind of those kind of information. And then I feel the smell, taste, skin color, and sound kind of stuff is very, very tied into primitive survival kind of stuff. And so our sort of sensory system is so sensitive that it's intractable to.

So for example, I asked one of the wine sort of tasters and the professional. And then he basically said he kind of used language from like a dating, describing a partner as a way to describe the taste, because there's no sufficient vocab to describe.

So I'm kind of curious. Yeah, do you guys feel that?

**Nicole Brichtova** [27:59]
I think, well, to some extent, I think the same is true for visual information. When you think about a certain style or a certain aesthetic, there are some people who just have a much more kind of developed, whether it's palate or kind of visual taste and aesthetic.

I think language just tends to be a bit of a limiting factor when you are trying to describe any of these things that we experience with sensory information. And to your point earlier, I think that is kind of the reason why we are investing in world models and why we are pushing on kind of the perception and the generation side of things, because it is such a large part of how we as humans navigate the world.

It's a large part of how embodied AI navigates the world. And I do think language does have a lot of it's gotten us very far. And it can probably get us really far. But it feels limiting in a lot of these kind of areas.

And yeah, I don't really know how to describe sense and taste. But yeah, I'm curious, Doomie.

**Dumitru Erhan** [29:05]
Yeah, I don't know that I have thought that deeply about this. So yeah, I mean, yeah, I don't have a good answer about audio. I mean, I don't know, because I'm thinking about, well, what is Omni bad at in terms of audio?

But they're all solvable problems, I find, with more data or better data or whatever it is. So I don't know that we have pushed the frontier so much that we have hit some sort of limits that are rooted in evolutionary kind of limits imposed by humans.

I don't know.

**swyx** [29:39]
He's feeling the limits of captioning, which is the thing I was.

**Shane** [29:42]
Yeah, exactly. There's a lot of information in the world. And it connects to basically why we do world modeling, as Dumitru mentioned.

**swyx** [29:48]
You just need SREFs, SREF 15476. And then that's your.

**Dumitru Erhan** [29:51]
But I guess.

**swyx** [29:52]
What my journey does,right?

**Dumitru Erhan** [29:53]
I guess maybe I can't describe this vibe.

**Nicole Brichtova** [29:55]
Well, I think that's kind of the point of providing some of these references, because even just describing how someone talks and their tone and prosody and all of these things, I think some of these terms even, I didn't used to know what they mean.

**swyx** [30:11]
Prosody, yes. Disfluencies.

**Nicole Brichtova** [30:13]
Exactly. There's kind of an entire vocabulary that even if you are not kind of steeped in a domain, which is true for actually most human domains, that you don't even know what it means. And sometimes it's also a question of, if we haven't focused on those things, with the large language models, that they may also have gaps in those areas.

And then we fill them on the other side with generation, because we're fundamentally relying on the language model's understanding of the world to then be able to represent it. So yeah, it all kind of goes back to your question about the language as an intermediary.

But yeah, I think to Doomie's point, some of these might just be focus areas and things that we haven't necessarily pushed on as much as we can. And as we will, we will discover what the actual ceiling is.

**swyx** [30:56]
Yeah. As a podcaster, I think a lot about sound. And I would just offer a couple of things for discussion in case it triggers anything with you guys. I have three domains of rough audio, which is music, voice, SFX.

Is that rough? OK, covers everything. And then also, even within voice, let's just focus on voice. Forget the other two. Room sound, the echoiness of big room, small room, in person, in a car, over a phone, all these are labelable.

But we experience them very differently. And I often think one that tells of an AI video is that it is studio quality, because it was recorded in a studio, because that's your training data. And to me, that's one thing.

Actually, the more interesting thing is just when I tell this is how I convince people who are kind of skeptical about the need for world models, because you need it even for audio, about, well, I'm further away from you.

So I should sound a little bit softer or more diffuse. And the video models need to pick that up, because if they're going to do immersive video and audio, you need that.

**Shane** [32:02]
I love that example of basically studio quality or not. In a way, we don't have enough language to really describe this kind of echoing or some kind of noise kind of happening. We just don't have precise enough. And basically, the reason that I think is quite important to have relatively information-rich kind of captioning is that we kind of rely on the natural language as a representation.

But if you basically don't have enough representation, that basically means that condition on the language, the generation is very multimodal. And if you anything can learn from the VAE, kind of like a very old VAE kind of research, the idea is really we want to capture most of the stochasticity in the lateral representation.

And then the X given the Z should be kind of deterministic.

**swyx** [32:47]
Yeah, yeah. Well, I hope there's more progress there. And I'm sure you guys are doing.

**Nicole Brichtova** [32:52]
I even actually do more like facial expressions. And maybe this gets to your point about things that we're very sensitive to. I think you can tell a lot of AI content also just from people's facial expressions.

**Shane** [33:03]
Yes, slope.

**Nicole Brichtova** [33:06]
Yes. And we try not to contribute to it. And/or skin textures, things that kind of make things look real in real life. I can tell from the way you're nodding or from the way your micro expressions are kind of changing of how you're reacting to what I'm saying.

We haven't quite crossed that chasm, I think. We're so much better than we were a year ago.

**swyx** [33:29]
Yeah.

**Nicole Brichtova** [33:29]
But there's so much more headroom kind of in a lot of those things that we as humans are super sensitive to. And I think image arguably probably is there, because there's a lot of kind of images that I will see that really do look indistinguishable from reality.

And I can't tell if they're generated or not.

**swyx** [33:46]
They're better than reality.

**Nicole Brichtova** [33:48]
Or. Well, that's a different.

**swyx** [33:50]
No, I think that one of the better.

**Nicole Brichtova** [33:52]
Better than what I would take on my vacation as a photo, yes.

### Better than real

**swyx** [33:55]
One of the fun experiments that we did a while ago on the team is like, can we generate videos that are better than real videos? So you just take the same caption from some video. And then.

**Shane** [34:07]
Then recycle it.

**swyx** [34:07]
Just try to describe a real video, and then generate the equivalent version with Omni. And then do a human eval.

**Shane** [34:14]
How does it do?

**swyx** [34:15]
And then humans largely prefer AI-generated videos.

**Shane** [34:18]
Oh, really?

**Nicole Brichtova** [34:20]
Yeah.

**swyx** [34:20]
But.

**Shane** [34:21]
Because it's the RL process.

**swyx** [34:23]
That's the RL process working.

**Shane** [34:24]
It's however you want to rationalize it. It's not necessarily the RL process. It's just like, I think it's just.

**swyx** [34:28]
I'm not saying this is a good result. I'm just saying we have optimized in a way that kind of potentially sort of triggers something in the human brain that like, oh, it looks a lot of the AI videos just look better.

On inspection, on deeper inspection, they would not actually be more useful or whatever. But if you just say side by side, random YouTube video versus generated version of it, you will just have a it will just look better, because it's sharper, more HDR.

The skin tone is better. Again, it's not more realistic. It doesn't solve your problem necessarily. But it looks better.

**Shane** [35:10]
I think it also depends on the sensitivity of the people. I was born and raised in Japan. And I think one thing I kind of know is they're extremely, extremely sensitive about that's why architecture, food, and stuff like that.

So I talked to a manga artist there. And he's kind of disgusted by the image generation AI. And one kind of thing he mentioned is the eye gaze. Eye gaze, that slight difference makes him kind of feel creepy about unnatural.

**swyx** [35:39]
Like if you're looking a little bit off?

**Shane** [35:41]
Yeah, it's just yeah, it looks too fake. Yeah. So I think it does depend on the sensitivity.

**swyx** [35:48]
Yeah, yeah, yeah. All I'm saying is human preferences are not particularly reliable barometer of what you should be optimizing for. If you just ask people, do you like this or not, you'll not necessarily get what you wanted.

**Shane** [36:01]
Yeah, let me just kind of add one thing. But four years ago, there was a debate that if the prompt engineering is going to disappear. And

some very powerful people say it's going to disappear. But I basically said it shouldn't, because the prompt engineering sort of specifying that is the only way you can sort of control the output. So when you have sort of control over the AI, and what allows you to prompt engineer is really that sensitivity.

So sure, mayberight now the AI can do a lot of auto prompting in that. And it can generate something that's sufficient. But if it's like that, never be satisfied. Never be satisfied with the AI's generated content. Always fine-tune your sensitivity.

And always kind of keep prompting for the differences.

**Nicole Brichtova** [36:48]
I think to that end, there's also a big difference between the average human untrained eye, which I would put myself in that bucket. I have some aesthetic sensibilities. And I've done this long enough that I have a preference.

But your example of a manga artist, that's somebody who has honed a craft over possibly many decades. And anybody who does that, whether it's design, architecture, you just have a very different level of expertise. And you see things that the average human will not see.

But Doomie'sright. When we look at if you were to just poll 10 people on the street, they would probably prefer the overly smooth, very saturated kind of content.

**swyx** [37:34]
Yeah, it's called the Instagram filter.

**Nicole Brichtova** [37:36]
It is, yeah. And so there's also a little bit of a question of, what does your default aesthetic look like if you don't specify? But then to Shane's point, one of the things we always try to get these models better at is instruction follow, so that when you want to get them to a different outcome, you should be able to, whether that's through language or whether that's through your references, because language is sometimes too limiting.

And so these models continue to get better at it. But they so much adwarm.

**swyx** [38:04]
Do you feel pressure as a product director to set the default for the world? I mean, kind of.

**Nicole Brichtova** [38:11]
Maybe I should. I don't know. I haven't thought about this.

**swyx** [38:14]
But it's like someone has to have a default. The default has to exist.

**Nicole Brichtova** [38:19]
Actually, I will say, we have thought about this. And I think one of the so for example, actually, if you look at Nano Banana generations, we had an explosion of Nano Banana infographics when Nano Banana Pro came out.

**swyx** [38:31]
I tried it, yeah.

**Nicole Brichtova** [38:32]
Yeah, yeah, yeah. I think NeurIPS papers were all so many had infographics generated.

**swyx** [38:39]
Can you run your watermarking on it and see how many?

**Nicole Brichtova** [38:42]
We probably could. We haven't done that. But I saw my Twitter was maybe this is just also the bias of my algorithm. But they were everywhere. And it was actually very painful, because I think our default aesthetic was a little bit too it was too cluttered.

I think the model was like a bit of an overeager student that just learned it was like, oh, I know all these I know all this information about this concept. Let me shove it into the same image.

**swyx** [39:10]
Japanese infographics 5x that.

**Nicole Brichtova** [39:14]
Or maybe it was but it just.

**swyx** [39:17]
Wait, wait. So same prompt, same content. If it's in Japanese, it's more.

**Shane** [39:21]
Density, density, yeah.

**swyx** [39:22]
Oh, wow. Because that's the style in Japan.

**Shane** [39:25]
Yeah, some very bureaucrat. There's a famous word for it.

**swyx** [39:30]
No, but we do go through this process with Omni. We did it together. We had a bunch of at the very end, OK, we did some tuning. And OK, what kind of style do we prefer? Is it more muted, more saturated?

**Nicole Brichtova** [39:44]
We had a lot of saturations.

**swyx** [39:45]
Yeah, I think Nicole just has PTSD, so has forgotten about it. But she was very much involved in this of like, OK, which kind of color palette do we basically prefer? And it's not something that you have to make a trade-off there.

**Nicole Brichtova** [40:00]
And it's not, because it ends up being us. Actually, it is true. It ends up being the modeling teams. And you could ask the question legitimately of like, are we the best people to do that? Or should we actually work with someone who has a really creative point of view and is more of an art director?

And we kind of go back and forth on this.

**swyx** [40:20]
We have the trusted testers.

**Nicole Brichtova** [40:22]
We have trusted testers who give us a lot of feedback. And we take that seriously.

**swyx** [40:25]
Very well organized, by the way. They have these weekly calls and stuff. It's amazing.

**Nicole Brichtova** [40:30]
Logan's team does a lot of that.

**swyx** [40:32]
Kudos to Logan.

**Nicole Brichtova** [40:33]
Kudos to Logan, who couldn't be here today. And we have a lot of people actually internally at Google, like Fulfer, who give us a ton of no, no, no, truly, who give us a ton of feedback on when we release new checkpoints.

And sometimes it will be stuff that we don't see. We would be like, oh, yeah, this optimization seems OK. And then they would come back and say, what have you done? You completely ruined my grass, because now the detail is all blurry.

**swyx** [40:58]
I think we just noticed not a super secret at this point, but that our model tends to put rings, wedding rings on hands.

**Nicole Brichtova** [41:04]
That's, yeah.

**swyx** [41:05]
Very strange. I had never noticed that. But he's like, I just saw it. There's a Fulfer channel, basically. He posts that. I was like, why is there a wedding ring in every hand? I'm like, that's so strange.

**Shane** [41:14]
That sounds like very common reward hacking.

**swyx** [41:16]
Yeah, yeah, yeah. But it's something that we would not have noticed necessarily while developing this.

**Shane** [41:21]
Is it an RL artifact?

**swyx** [41:23]
I don't know.

**Shane** [41:24]
You do have a lot of preference base. And then you may prefer that. Spirits correlation, reward hacking, it can happen in many weird ways.

### Evals

**swyx** [41:33]
It does, it does. This is related to another topic that, again, I try to use these main stage things as introductions or ties in. We have an evals track. We have character AI and YouTube talking about how they evaluate videos.

How do you evaluate videos? Apart from paying Fulfer. Not everyone has a Fulfer. But also, I think there needs to be something more quantitative.

**Shane** [41:58]
Well, I mean, you improve Gemini to improve the evaluation for video.

**Nicole Brichtova** [42:05]
No, no, that's definitely one way. It's actually very hard.

**swyx** [42:08]
It's very hard.

**Nicole Brichtova** [42:09]
It's very hard to get operators to evaluate things in a video, including especially things like aesthetics. There are some things that are a little bit more objective, especially when we talk let's say we talk about images. And we look at infographics text rendering.

That's actually fine, because you can kind of OCR things out. And then you can look at like, OK, this letter is messed up. And then the whole thing is actually useless, because literally, if a letter is off in the rendered text, you just can't use that asset.

So those things are a little bit more auto-ratable from what we found. We do rely a lot on humans looking at things. And so we do do a lot of human evals. We do a lot of human evals.

**swyx** [42:51]
Do a lot of human evals.

**Nicole Brichtova** [42:52]
And every time Shane is like, and every time we have a new model, we want to do more things. And we want to jam in more capabilities. And so then we have more evals that we have to run.

And then at some point, you do get two models that are kind of close to each other. And then we literally make decisions based on looking at output side by side, sometimes in a room. I've been in rooms where there's like 10 of us.

And we're just looking at videos side by side. And we're like, do you prefer this? Or do you prefer that?

**swyx** [43:25]
Oh, wow. I mean, but it is genuinely very complicated. The more capabilities you add, even just the one capability, but it's like almost AGI complete capabilities, like video editing. Think about video editing as a and editing with audio.

**Shane** [43:39]
My editor would be very happy to hear this.

**swyx** [43:40]
It's the hardest problem in Gen Media. I mean, I don't know if it's the hardest, but it's definitely there. In terms of complexity of evaluation, free-form video editing, you can do anything.

**Shane** [43:57]
I spent a lot of money on that. And it's very hard. Please help me.

**swyx** [44:01]
We don't have add a sloth eval.

**Nicole Brichtova** [44:05]
Well, now we should.

**swyx** [44:05]
Now we should, yeah, yeah, yeah. But things like that, it's not that easy to track.

**Shane** [44:10]
I think I'm just surprised at the sample size that you have. To test the entire surface of your models, you still rely on audio magnitude of hundreds.

**Nicole Brichtova** [44:20]
No, no, no, no. Well, we do a ton of human evals on thousands of things. I think there's also an element of we can talk about things like live experiments, which is also where you get signal on some of these more minute differences at a much larger scale.

Then there's auto-raters, which is definitely kind of a more it's a very well-defined space, I think, for LLMs, much more nascent for media models. And then sometimes you still do rely on human judgment. And we do rely on things like feedback from people who just have a very honed aesthetic and people who just use these models in their workflows day to day.

Because we could also you could have a model that is really well on some slice of human evals. But then it really breaks a workflow for somebody. And so this is why we do early access programs. And we try to get feedback.

And then we try to incorporate it before we release something more broadly. I feel like Shane had a hot take based on his expression when we were talking about this.

**Shane** [45:22]
Every kind of human work should be gradually kind of amortized. And then the interesting thing is the video understanding, especially against AI generated video, detecting AI stuff is an extremely interesting vision task. And then some of it is kind of aesthetics, so this kind of visual quality.

But for some of the kind of cases, semantically it doesn't make sense. For example, you're taking some famous scene from a movie and try to sort of construct that. And then if you kind of generate it, it can generate something there.

But at some point, some of the semantic information doesn't make sense. It's actually inconsistent. So can the AI actually detect that? So when I evaluate the AI video, it's like, oh, I feel I am so smart. It's like AI is still kind of behind.

But we should make a lot of effort. I think the video understanding is an extremely important intelligence task beyond just the pure aesthetics or the preference. And yeah, we should always try to amortize the human label.

**swyx** [46:25]
Yeah. What data do you need? A lot of people I talked to wanted to get in front of you, actually. I mean, they want to be nice about it. They have a lot of video data. They have gaming data.

They have real-world video data. They have images. They have labelers. What do you want?

### Data

**Nicole Brichtova** [46:45]
Are you offering?

**swyx** [46:46]
I'm just like, this is your request for like, OK, OK, I'm sure you get a lot of pitches. You got a lot of people who want to talk to you. I think, actually, it's probably the sorting out signal from noise is the main problem.

So creating a nice API of like, OK, if you actually do A, B, and C, we are interested in that.

**Shane** [47:11]
Loaded question there. So I don't know that there's an easy I think we do already have a lot of data. I think it's hard to talk about this.

**swyx** [47:22]
You want to talk about it in public. I don't want to give you trouble.

**Shane** [47:25]
No, no, I just want to say it's hard to talk about this without trying to I have to think about what I am revealing about our project and where we're going. Generally, high-quality data, I think. Maybe let's just put it this way.

It's not the secret.

**swyx** [47:41]
Embodied?

**Shane** [47:42]
Sorry?

**swyx** [47:42]
Embodied data?

**Shane** [47:44]
I mean, yeah, sure.

**swyx** [47:47]
We have sort of announced, I think, publicly that we have some sort of robotics collaboration. So I think it's like

because we have a robotics team at GDM. So they're always interested in things like that. I mean, for Omni specifically, I think we're just quite interested in just high-quality data. It's not necessarily like, oh, a random YouTube video, but a more professional shot, things like that.

Those are things that we're always on the lookout for.

**Nicole Brichtova** [48:19]
And I think maybe this is easier to some extent to answer for some of the agentic work as well. Actual kind of what are the tasks that people are trying to do? These things are actually kind of difficult to manufacture if you're doing it yourself or if you're doing it with a vendor.

What is the actual if you're creating a marketing campaign, what does that look like? Do you start from, here's a picture of my new product. And then I want to turn that into a video ad. And I want to turn that into a bunch of assets that fit all these different ad formats that I need to push onto the various platforms to promote.

And then so you kind of go from this to that. And what is that kind of trajectory of tasks that you're experiencing along the way? That is really useful. And that is actually kind of difficult to get, because we don't always have theright first-party surface where people are actually doing some of these things.

Or you might work with someone who's a vendor, but they also don't have that product surface. A lot of this kind of information lives in the places where people are doing these tasks. And so that's kind of difficult to get.

If anyone's figured that out, you should reach out to us.

**Shane** [49:31]
Every channel thought.

**swyx** [49:33]
Every channel thought.

**Shane** [49:34]
Every channel thought. And maybe the data the Chinese lab is using.

**swyx** [49:39]
Yes, yeah, yeah.

As a media person myself, there's so many podcasters and people in marketing departments and all these. They would be happy to be your data. Just put a BCI on my head.

**Nicole Brichtova** [49:53]
And talk to us.

**swyx** [49:54]
Watch me think. Because there's just an endless amount of work to do. There's so much work. And this is all these needs to somewhat be a commodity. Obviously, you can be an artisan. You can be Hollywood for the really high-quality stuff.

But actually, a lot of work is commodity. And it should be modelable. And we want you to do it.

**Nicole Brichtova** [50:14]
And we want the high-quality to do me this point. We do what we want, the high quality.

**swyx** [50:19]
We want commodity. Yeah, yes. You want to be on both sides.

**Nicole Brichtova** [50:25]
Thank you for the solicitation.

**swyx** [50:30]
I also added a data quality track. I think that people want to understand

how to raise the bar.

And a lot of it is just educating the market and educating researchers and engineers and founders on, this is where we're going. A lot of this is slop. Stop doing that. Do this instead. And people will listen. Yeah, I don't know.

To that extent.

**Nicole Brichtova** [50:56]
But I think to that point, there's a lot of, again, just craft that goes into this. And there's a lot of process. Even to the marketing campaign example, you don't create that in five minutes. You go through a process.

And you iterate. And you pick something over something else because you liked it for whatever reason. Maybe the eye gaze was correct. We don't know these things, because none of us are marketing directors. And the models don't know these things.

**Shane** [51:21]
I even kind of say this for the natural language as well. I always kind of say 99% of information is inside people. You can only extract it through active dialogue and befriending them. So most of the stuff on the internet is sort of the outcome, the output of that.

But what are all the trajectories? How did this person have this inspiration to write this paper? What is the starting point? What is the inspiration? What are the dialogue that sparked it? Those kind of stuff is kind of inside people.

So even those kind of even the language space, there's kind of that. I think the creative is kind of similar as well. There's a lot of dark knowledge.

**Nicole Brichtova** [51:54]
Yeah, it's like when you write a novel. A novel speaks to you because usually, there's some sort of a personal connection that you feel to the story or the trajectory or the characters. If you read most of the stuff that's written by LLMs today,

it falls into these default patterns. And the language starts to feel really similar. And all the descriptions sound really similar. You can kind of quickly read it as like, oh, this is not that interesting, because I can't connect to it.

And again, that's kind of like a human expertise.

**Shane** [52:26]
One nice thing recently is that Google Cloud and Google DeepMind are kind of starting to invest a lot more in the FDEs for the product engineers. And I also kind of saw some recruiting for the creative Gen Media kind of space as well.

So I think those are kind of really the effort, because we kind of feel that what we can kind of do with a lot of public data, there's limits. But really partnering with that, we can provide kind of better models and products and really kind of feedback.

**swyx** [52:51]
We have an FDE track here for the first time. Every lab is announcing it. It's crazy. One thing I'm actually very keen on doing and I push for this at Cognition as well is to turn the FDEs not just into sales and solutions, but also to evals workers.

### FDEs

**Shane** [53:08]
FDE is not the sales. FDE is way bigger than that.

**swyx** [53:12]
How do you frame FDEs then? Besides, I do think about it as sales. The more you customize the solution.

**Shane** [53:21]
So I define post-training as anything between the pre-training and the final user experience, anything. Anything is a post-training. And to me, when I first sort of learned a lot about I mean, FDE kind of, I guess, originally came from Palantir and then that.

So I guess the kind of history is different. But yeah, I think the key is really that. The key is not only to kind of work with them and ensure that they kind of know how to use, but also to sort of code-derive kind of insights that can basically kind of help both parties.

They can put a lot of harness, how they use the model. We can improve very upstream. So how to get the customer feedback to the modeling, I feel, is kind of more the role I kind of want for the FDEs.

**swyx** [54:04]
Yeah.

**Nicole Brichtova** [54:04]
And even if I was sorry, just on that, if you want to talk to us, or at least me, I'm not going to offer up your time. But it's really helpful for us to actually talk to people who are using our models and understand where they're struggling.

Because again, it's the real-world task that you're actually trying to use them for. I will talk to people who do kind of interior design with some of our image models. And they will say, hey, I really want to take this pattern, but then I want to scale it across 10 different rug sizes.

And sometimes I have a very custom rug size. And then the model fails at replicating the pattern the same way. Or I want to do a try-on for these earrings. And then the earrings have a certain size. And then my head has a certain size.

It has to make sense if you're actually trying to try things on. And the models kind of fail at a bunch of these things that actually happen in the real world. And so that's useful for us, because for some of these things, we don't think about it, because we don't use the models for those tasks.

**Shane** [55:07]
Or I think to your point about ad campaigns or whatever, people have notions of brand languages or whatever, which is a bunch of images or PDFs saying things. It's a pretty kind of ambiguous question as well. What is the IKEA brand language?

Is it blue and yellow? That's not a very.

**Nicole Brichtova** [55:26]
But what shade of blue?

**Shane** [55:28]
Yeah, yeah, yeah. And the brands are pretty specific. They do care about the shade of blue. It shouldn't just be a random blue and a random yellow. That's not going to be IKEA. I'm just thinking about an example.

But this is the kind of stuff that it's not necessarily part of our developing frontier models, kind of necessarily mandate. But it's something that we do want to. We do want to fundamentally build products that people will use to solve concrete tasks, not just research artifacts.

So I think it's useful to understand what people do care about.

**swyx** [55:57]
Well, I'm sure a lot of people are very grateful for your work. And there's a lot more to do. You've made so much progress over the last even just couple of years of Nano Banana and Veo and Omni.

And I don't know what else you got cooking, but we're very excited. This is one of those things where I was very disappointed when Sora shut down. And I think there needs to be more general exploration of generative models and not just coding.

I think that is.

**Nicole Brichtova** [56:27]
We obviously like this.

**swyx** [56:28]
We love coding. We love coding. And yes, but thank you so much for your time. It's been a real pleasure. And I can't wait to see what this looks like next time.

**Nicole Brichtova** [56:35]
Thank you for having us. Great questions.

**swyx** [56:37]
Thank you, everyone.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
