# Evaling Video Slop — Maor Bril, Character.ai

AI Engineer · 2026-07-25

<https://aiengineer.podhood.com/6ef30cf3-8a6c-445e-a77f-29a07b1c466b>

Maor Bril of Character.ai argues that evaluating AI-generated video quality requires pairwise comparison rather than absolute scoring, because CLIP score misses temporal incoherence and LLM-as-a-judge is too slow and expensive. His team trained a small Qwen3-VL judge using Bradley-Terry loss on pairs of real and deliberately broken footage, catching drift early by running the judge as a regression gate in CI—every AgentX release clears an eval wall calibrated against human scores. The judge scores a 15-second video in three seconds, and they avoid becoming an AI detector by ensuring consistent encoding and annotation across real and AI footage. They also evaluate sound using Atmos and correlation with key frames, but admit lip syncing remains unsolved. The fix: score the axes you care about (story, pacing, physics) and put eval inside the generation loop.

## Questions this episode answers

### How did Character.ai train a video quality judge to avoid the 'vibe vs. substance' problem?

Maor Bril's team abandoned absolute scoring and instead used pairwise comparison: 'Is B a better story than A?' They trained a small Qwen VLM with Bradley-Terry loss on pairs of real footage and deliberately broken clips. To prevent the model from becoming a simple AI detector, they ensured encoding was consistent across real and generated videos and annotated both sides identically. The resulting judge runs inside the generation loop, scoring explicit axes like storytelling, physics, and pacing rather than superficial gloss.

[15:15](https://aiengineer.podhood.com/6ef30cf3-8a6c-445e-a77f-29a07b1c466b?t=915000)

### Why do AI video quality judges sometimes give high scores to clearly flawed clips?

Bril found that initial judges learned to score the 'vibe' or visual gloss rather than the actual quality axes. His team's first model gave a cinematography score of 9.2 to a four-second still frame where the camera never moved, and praised the physics in videos of hovering ghosts. This happened because the training data inadvertently rewarded coherent-looking videos rather than checking whether the clip told the intended story with correct motion, physics, and shot continuity.

[17:11](https://aiengineer.podhood.com/6ef30cf3-8a6c-445e-a77f-29a07b1c466b?t=1031000)

### How can you catch and fix drift in long AI-generated videos efficiently?

Bril explains that drift is cheapest to catch early in the generation process. By placing a fast VLM judge inside an agentic workflow, drift between starting frames of consecutive shots can be spotted before full video assembly. For long-form content built from many short generations, detecting a six-second clip that drifted allows it to be regenerated immediately, avoiding expensive rework later and keeping the final output coherent.

[10:10](https://aiengineer.podhood.com/6ef30cf3-8a6c-445e-a77f-29a07b1c466b?t=610000)

## Key moments

- **[0:00] Intro**
  - [0:14] 'The generation of video is basically free, but most generated video is not that good' — Maor Bril, Character.ai
  - [1:19] 'The hard part was never making video; the hard part is judging if it is good enough' — Maor Bril
- **[1:24] Challenges**
  - [1:55] Human judgment remains essential for high-quality video because current evaluation tools miss storytelling coherence
  - [3:05] Video is a storytelling medium, and CLIP score fails to check whether the video tells the intended story across frames
- **[3:14] Story Axes**
- **[3:59] LLM Judge**
  - [4:43] Character.ai's first eval harness combined frame metrics, an LLM judge, and human annotation for calibration but was too slow and costly
  - [5:57] Catching character drift between starting frames is significantly cheaper than fixing it after generating the full video
  - [7:01] Story and sound failure modes only manifest over time, so evaluation must be embedded early in the generation pipeline
  - [8:28] A distilled small vision-language model scores a 15-second video in 3 seconds, trading a slight accuracy drop for dramatic speed gains
- **[9:20] Pairwise Prefs**
  - [9:20] Don't score, compare: pairwise preference (video A vs B) yields more reliable and generalizable video quality evaluation
  - [10:31] 'It was so wrong, but it was wrong in a very confident way' — Maor Bril on the first video judge scoring a frozen 4-second frame 9.2 for camera work
- **[10:47] Vibe Trap**
- **[11:53] Real vs AI**
  - [11:53] To fix overconfident vibe scoring, Maor paired real footage with AI footage, ensuring consistent encoding to avoid building an AI detector
- **[13:27] Agentic Eval**
  - [13:27] Switching from rigid pipelines to an agentic workflow with self-verification tools allowed the system to adapt to diverse, user-generated stories
- **[14:01] Takeaways**
  - [14:01] Maor's three takeaways: go relative, score real axes like story and physics, and put evaluation inside the generation loop
- **[15:05] Q&A**
  - [15:16] Q: How does Character.ai evaluate sound matching? Uses Atmos for quality and correlates key frame timestamps with audio spikes
  - [17:00] Q: How does Character.ai handle lip syncing? Lip syncing is still an unsolved problem, especially for talking animations without clear lip movement
  - [17:54] Q: How do you align subjective human taste in video evaluation? Periodic team annotation sessions on random axes calibrate the AI judges
  - [19:40] Q: Why did Character.ai choose Qwen's small vision-language model? Prior positive post-training experiences and sufficient performance drove the choice
  - [20:09] Q: At what scale does the custom video eval model become cost-effective? It depends on unit economics; you can scale instances horizontally to handle any volume
  - [21:58] Q: Will Character.ai export auto telemetry from the open-source eval harness? Maor accepted the feature request for future addition

## Speakers

- **Maor Bril** (guest)

## Topics

LLM-as-Judge

## Mentioned

Character (company), Atmos (product), CLIP (product), IPS (product), Qwen (product)

## Transcript

### Intro

**Maor Bril** [0:14]
So, hi. I'm Maor Bril. I've been with Character for a bit over 2 years, and we'll talk about AI slop,right? I think that, you know, when we look at video generation as a whole,right, we have like two kind of parallel tracks.

One is the video generation, which became insanely good from models like Kling and Seedance and Veo and Sora. We still remember Sora. But the part that got left behind is how we evaluate the quality of the video that was generated,right?

So on the one hand, we still kind of squint at it and decide whether or not it's good, but on the other hand, we know that the generation has gotten a lot better. And when we look at X or whatever social you're consuming your content on, there are a lot of guides on how to create amazing videos with this model or that.

So the hard part was never how to make video. The hard part was how do we generate

### Challenges

**Maor Bril** [1:24]
good enough video, and how do we judge if the video is good enough? So now we've gotten to a world where the generation of video is basically free,right? Free, especially when you compare it to how much studios would charge.

But the problem is the most, the grand majority of video that is generated is not that good,right? We have like a lot of hallucinations, like a third limb, opening and closing the door at the same time, hovering, physics, etc.

So unfortunately, in order to get high-quality content, we need a human to judge. And I don't know, when was the last time you've seen how someone is creating these long-form

generated video? It's usually a lot of shorter generations and a lot of editing. The problem is because we're using a lot of the tools that we built for the text era, for the image era, for videos,right? We're using things like CLIP score, which is great to judge a single frame.

Things like

IPS will help us kind of detect the drift between frames, but we don't have, I mean, but the problem is when you kind of combine all these together, all these tools, they're good at watching the individual frames. They're good at checking, does this specific frame, does it match the prompt that generated it,right?

It will check consistency between frames, and it will check whether or not it matched the prompt that drove it. But what it won't do, it doesn't tell you if

you told the story that you meant to tell,right? If you think about what is video, video is a storytelling medium. Video is just another form on how we tell a story,right, for any type of story. So one of the things we have to look at, does it tell the actual story?

### Story Axes

**Maor Bril** [3:21]
Does the physics make sense? Like, for example, if we want a video of a character walking downstairs, does it actually walk or hover?

Does the character stay the same character across multiple shots? Does the pacing make sense? Like, you know, for example, people take time going from one place to another. We need to make sure that the pacing makes sense as well.

And especially when we add audio, we want to make sure that the audio is kind of synced with the imagery. Like, for example, if someone is slamming a door, we want that sound of the door being slammed to be exactly when the door is actually being slammed.

### LLM Judge

**Maor Bril** [3:59]
Now, the next iteration we all went to a while ago, we started using LLM as a judge for everything. And we have amazing foundational models that we just throw videos at them. The problem with them is that, A, they're slow; B, they're only as good as your prompts.

And multiple people will prompt multiple ways, and the same model may respond in a very, very different way. And sometimes the prompt we use, like, is it consistent? Does this match the prompt? But then the question we really care about, is it good?

And the answer varies. So, oops, sorry about that. So our first iteration is like, let's take all these things and build a repeatable benchmark on how we test video that we can rerun over and over and over again.

So that combines both metrics, as I said earlier, that knows how to view individual frames, but also a consistent LLM as a judge,right, where we also use human annotation to calibrate the LLM as a judge. So for every report that we generate with that harness, we're able to have humans annotate it and basically feed that feedback back into

the LLM as a judge prompt to make sure that it's aligned with what I think or what the annotator thought is good. And we use it to score the videos. The problem with this approach, it's very slow, it's very expensive, and especially when we want to bring it for our users to be able to generate a lot of video, because creation is a very hard process.

And so the problem, as I said, the problem is, so this is a slow process, and we need to bring it as close to the users as possible and also earlier into the process. The reason for that is if we take a look at all the metrics and there are mistakes that we can find earlier

than later, then it's a lot cheaper to correct that particular mistake. So for example,right on the left, we have

two starting frames of different shots,right? But it's easy to correct, to view it at this point and see, did the character drift between frame one and frame two? Because those frames will be used as starting frames to generate videos.

So if you can catch the drift at this point and correct it, then it's much cheaper to generate the video as a whole because we can correct it at a much cheaper cost. And the same thing applies when we look at longer form video,right?

When we see all these three, four, five-minute long videos, they're usually a collection of a lot of shorter videos. And being able to catch a six-second generation that drifted and regenerate that before we combine the whole video will end up being a better result as a whole.

And now the other problem we're trying to solve is some of these axes,right, only exist across time,right? So for example, when we look at the,right, we mentioned the story,right? So does the story that we're trying to tell with that video, does it hold in that video?

Does the video tell the exact story? Does the pacing make sense,right? And we mentioned the sound. So as I said,right, the underlying goal is to bring that evals closer to the online generation because the sooner we're able to catch those mistakes, the sooner we're able to catch that drift,right, then it's much easier, much cheaper to fix.

Now, so now the problem is that, as I said, this is a very slow process. So the solution is actually to take all these committee of experts and distill it into one small model that is also very, very fast, but is able to give us a response that is not whether or not this video is slop or not, but why is it slop,right?

Why is that video scored low versus the other? Because, for example, it added an extra limb, because it didn't obey physics, because the audio was out of sync. So the goal was, A, build it on top of a small VLM.

And why is it a VLM? VLM because we needed the model to be able to see the image. But also we needed to work fast,right, because we brought it closer to the generation where, in fact, it takes about, with the model we have trained, it takes about three seconds to score a 15-second video.

Now, we also tested a bigger model and the results were better, but it was significantly slower. And the decision was to go with the smaller model because the added value from the bigger model didn't justify

the slowness. The other very interesting realization we came to is don't score, compare. What does that mean? For example, if I'll ask any person in this room to look at a particular video and rank it from one to ten on storytelling,right?

### Pairwise Prefs

**Maor Bril** [9:33]
I'm pretty sure that, you know, what will be a six for you will be a five for you, will be a four for you, and an eight for you,right? But if I'll show you two videos and I'll ask you which one of them is telling a better story, the grand majority will probably agree that B is telling a better story than A,right?

And if you do it enough times, then it's easy to generalize the model towards detecting what's better versus not. So we trained on pairs,right? A versus B as opposed to one through ten. Now, we manufactured badness. So luckily, the internet is full of very high-quality videos and it's very, very easy to get good videos.

And it was very fun to create bad videos, A, by either corrupting good videos or by, you know, just generating random slop.

Now, we shipped v1 and it was so wrong. It was wrong, but it was wrong in a very confident way. So for example, the frame you see here is from a video that the model scored 9.2 on the camera work and the camera didn't move.

### Vibe Trap

**Maor Bril** [10:49]
For four seconds, it was like a still image of the same character, but the model was very, very happy with the cinematography. So the physics in some other videos, which I'm not showing because of time limitations, it says that the physics look great, but it said it on ghosts hovering and people flying, etc.

So, I mean, so then the question is like, why was it wrong? The reason it was wrong is because how we generated that data,right? It scored the vibe as opposed to the axes. So it learned how to detect coherent videos and it learned how to detect the artificial artifacts, basically the gloss of the video as opposed to whether or not the video actually told the, sorry, the videos actually told the story.

### Real vs AI

**Maor Bril** [11:54]
And so the solution was to fix the dataset. And so the way we fixed the dataset, we actually, I started pairing real footage versus AI footage. Now, the risk with that, and that's the reason why I avoided doing it at first, is because I didn't want to create an AI detector,right?

Because if you start creating pairs of good as human-generated video and bad as AI video, then there's a very big chance of the model overfitting and becoming an AI detector as opposed to a video quality detector. So there are two things I did in order to avoid that.

A, I made sure that the encoding is consistent across both sides of the equation. So there's no artificial artifacts for video A versus video B. And I used the exact same method of annotating both videos. So both the axes, so all the axes in those videos were annotated in the same way.

And surprise, it turned out pretty awesome. And so now what we're able to do, especially when you're looking at videos, A, we changed from a very complex pipeline,right, to an agentic workflow. The reason behind this is, A, the pipelines work great if you have a very, very unique use case.

### Agentic Eval

**Maor Bril** [13:32]
But once you put it in front of users, they'll have a very, very distinct story that they want to tell with their own characters, with their own images, and their own voice. So that's when it starts to drift.

But by providing the agents with tools to validate the quality of the output, it's able to adapt to changes better, but it's also able to verify its own work and fix things as they go along. So if you're going to steal

from these, from the stock of you things, one, go relative, not absolute,right? As I explained earlier, the value of comparing video A versus video B will always give you a better result going forward. B, score the real axes that you care about.

### Takeaways

**Maor Bril** [14:20]
So if you care about storytelling, if you care about pacing, if you care about physics, score those axes. Don't expect them to miraculously appear. And put eval inside the generation loop,right? Especially if your goal is to have a higher quality of generation, get the eval as close to the generation loop as possible.

Eventually, evaluate it as a story. Videos are stories. Videos are just another way for us to tell stories to others. And

thank you very much.

### Q&A

**Host** [15:06]
Allright, any questions? Okay, down here. Awesome. Allright, I got two down here. Here you go.

**Guest** [15:16]
Sure. Hi, how do you eval sound? Sound and video matching.

**Maor Bril** [15:22]
I'm sorry, can you repeat?

**Guest** [15:23]
How do you eval sound? Sound as effects and matching with the video.

**Maor Bril** [15:27]
Oh, yeah, that's a fantastic question. So

sound is actually a combination of a few things. One, I'm using Atmos

to make sure that

the sound quality is high enough and is understandable. B, the model will learn to identify key frames,right? And especially because when I feed something into the model, it can be just the video or it can be the video plus the prompt that generated that video.

So for example, if the prompt will say the door slammed,right, it will look for a door being slammed and will match the sound at that same frame.

Did that answer your question?

**Guest** [16:21]
How does the model recognize sound?

**Maor Bril** [16:23]
So it's both by using

Atmos and also to correlate. So for example, when it's looking at the frames,right, it's making sure that, for example, the door being slammed at frame six. Frame six has a specific timestamp. So it's looking for that

spike in the sound at that timestamp. It doesn't know that it is that sound, but it's looking for a specific spike of sound at that timestamp.

**Guest** [16:57]
What about lip syncing?

**Maor Bril** [17:00]
Lip syncing is an unsolved problem yet.

**Guest** [17:04]
One more question.

**Maor Bril** [17:05]
We're trying though.

**Guest** [17:06]
Can you repeat the question?

**Maor Bril** [17:07]
Yeah.

**Guest** [17:08]
No, go for it.

**Maor Bril** [17:10]
Yeah, so the question was, what about lip syncing?

**Host** [17:15]
Oh, that wasn't me, but.

**Maor Bril** [17:17]
Oh, I'm sorry.

**Host** [17:17]
I guess the lip syncing answer would be interesting before I ask my question.

**Maor Bril** [17:22]
Yeah, as I said, it is an unsolved problem still. We're still working through it. Especially for us, you know, some of the characters that we're trying to do are talking head, are humans,right? Which, you know, we can look into different techniques to try to identify the lips, but some of them are just talking, you know, talking animations that have no real correlation between, you know, the movement of the mouth and speech.

So unfortunately, I don't have a solution for that yet.

**Host** [17:54]
So I'm curious about, for example, if you wanted to further enrich the dataset with human evaluation.

**Maor Bril** [18:02]
Yes.

**Host** [18:04]
The question of taste and what is good, because I think there's a big question mark about, is that going to remain the domain of humans? But I've also seen people say that, well, most humans are really, they have terrible taste anyway in videos and games and books.

**Maor Bril** [18:20]
Fair.

**Host** [18:21]
So how would you construct and align sort of like any human judges?

**Maor Bril** [18:28]
Yeah, so this is actually solved at first at the judge duty part, where every report it'll generate a human can go and annotate it. And we actually, we do that. We periodically have sessions where everyone spends 10 to 15 minutes just annotating videos.

And that usually happens on multiple axes. I won't ask everyone to annotate the same video on 10 different things. It'll be random. And I use the data to calibrate the AI judges. And the results from that is actually being served as a dataset for training for the next version of that model.

So it's a process that does take a little bit of time and hopefully, and it does evolve over time, but it's not immediate. You know, because also taste is very subjective and things that are great for me, you know, that I think are fantastic, some people that come and say, "Ah, are you sure they're great?"

Because, you know, so yeah, it's a process and I use the human feedback to calibrate the models all the time.

**Host** [19:40]
How did you land on the Qwen small VLM? Did you try any others?

**Maor Bril** [19:45]
I did. So the intent I had was to, A, you know, find a small enough model. The reason I went with Qwen is because we also had a very good experience with post-training Qwen on other use cases. So I mean, yes, I could have, I did try a few others, but it just, you know, everything was just there and it was good enough.

**Guest** [20:09]
So my question is about scale. So obviously Character.ai produces thousands, millions, a bajillion videos.

**Maor Bril** [20:16]
Yeah.

**Guest** [20:16]
At what scale does this become reasonable for my domain that is not Character.ai? So my domain has hundreds, maybe a thousand videos.

**Maor Bril** [20:23]
Sure. So if you're happy with the cohort of experts and you don't need,right, so I'll rephrase that. The scale is both for speed,right, as well as capacity, because I can serve this model as one instance on one GPU or I can serve it as, you know, a hundred instances,right?

So that determines my scale. The reason I chose to go towards a model is because I wanted to speed up the creation process,right? It would have worked just as well if I didn't have this particular model. I would have used like the cohort of experts,right, from metrics that are available both on CPU and GPU, as well as frontier models,right?

So it was a balance out as, you know, A, how long did it take me to train this model and to curate the dataset and get it to a working set,right? And how much does it cost to serve it versus how much it would have cost me to do this, A, slower.

Now, potentially it is better,right? I mean, like I assumed that if you're going to use the FABO, which came back today,right, it will probably give you a better result. But at what cost,right? If you do it for one or two, it's probably fine.

If you do it for thousands or tens of thousands per day, it adds up. So it's a matter of your unit economics.

**Guest** [21:55]
Cool. Right over here.

**Maor Bril** [21:56]
Yeah.

**Guest** [21:58]
On yourright. There you go. Last question.

**Maor Bril** [22:01]
It's very bright. I'm sorry.

**Guest** [22:02]
No worries, no worries. My question is, I looked a bit at the repo. You guys don't export auto traces of the LLMs' judges yet.

**Maor Bril** [22:12]
Correct.

**Guest** [22:12]
Is that something, are you open to that? So you can connect to other platforms?

**Maor Bril** [22:18]
Sure. So the repo itself, it's a harness and you can connect any

agents or any LLMs you want. We actually have an internal version of this, which is running it as a service,right, with an agentic harness on top of it that has all

the metrics we care about. But I do accept your feature request and I'll be adding auto telemetry to the harness.

**Guest** [22:49]
Awesome. Thank you very much. A warm welcome or round of applause for Maor. Thank you.

**Maor Bril** [22:55]
Thank you all.

**Guest** [22:56]
Thanks.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
