# What's Next After RLHF? — Diogo Almeida, TypeSafe AI

AI Engineer · 2026-07-31

<https://aiengineer.podhood.com/c519fe68-a7e5-455b-86dd-92a01f68e515>

Diogo Almeida, a GPT-4 co-author and founder of TypeSafe AI, argues that RLHF optimized models for human approval, making them superb assistants but unreliable for autonomous work. He contrasts assistance with automation, noting RLHF's reward model encourages confident overpromising — like ChatGPT praising an audio file of farts as music. The next era is not Claude Code (still assistance-native), but real automation via RLVR-style methods focused on calibrated decision-making, echoing Sutton's bitter lesson that the task matters more than data. He defends pre-training as phenomenal, blaming post-training's asymmetric reward model for hallucinations. TypeSafe is rebuilding the AI stack for reliability and automation.

## Questions this episode answers

### Why does Diogo Almeida argue that RLHF makes AI models unsuitable for automation?

Diogo Almeida explains that RLHF models optimize for human preference, making them excellent at assisting humans but unreliable for autonomous tasks. In assistance, pleasing the user is the goal; in automation, accuracy and independent operation matter. Current AI, including ChatGPT and Claude Code, is stuck in the assistance era because RLHF inherently prioritizes engagement over correctness, creating a gap between benchmark performance and real-world automation.

[4:31](https://aiengineer.podhood.com/c519fe68-a7e5-455b-86dd-92a01f68e515?t=271000)

### What does Diogo Almeida propose as the next era of AI after RLHF?

Almeida calls for moving past RLHF's assistance focus to 'real automation,' where AI handles work without human oversight. He argues RLHF was a detour and that his company TypeSafe AI is building a new stack optimized for calibrated decision-making, not human preference. This approach aims to unlock smarter software that automates routine tasks, not just assists, by mainlining pre-trained models' intelligence into reliable, autonomous execution.

[8:52](https://aiengineer.podhood.com/c519fe68-a7e5-455b-86dd-92a01f68e515?t=532000)

### Why does Diogo Almeida say that overpromising is a feature of RLHF models?

Almeida states that RLHF's design inherently rewards models for overpromising. Since RLHF optimizes for human preference, it punishes uncertainty and encourages confidence, causing models to err on the side of giving pleasing answers even when wrong. He cites ChatGPT praising a fart audio file as symphonic as an example: the model prioritizes a positive, engaging response over truth, which is acceptable in assistance but dangerous for tasks requiring accuracy.

[7:06](https://aiengineer.podhood.com/c519fe68-a7e5-455b-86dd-92a01f68e515?t=426000)

## Key moments

- **[0:00] Intro**
  - [1:04] Diogo Almeida: "I'm one of the few people at OpenAI who actually hates on ChatGPT."
- **[1:49] Two Camps**
- **[3:13] Sane View**
- **[5:16] Human Preference**
  - [5:16] Today's AI inherited from RLHF excels at human-in-the-loop tasks but not automation, says Diogo Almeida.
- **[6:48] Assistance Design**
  - [7:38] ChatGPT called a fart audio file 'an eerie vibe atmosphere piece,' demonstrating RLHF's overpromising, says Diogo Almeida.
  - [8:10] Today's AI was designed for assistance by optimizing for human preference, not for autonomous tasks.
- **[8:47] Real Automation**
  - [9:10] Claude Code remains in the assistance era of RLHF; real automation looks different, Diogo Almeida argues.
  - [11:15] Diogo Almeida critiques the 'golden age of just-in-time software' as failing to make software itself smarter.
- **[11:47] Smarter Software**
  - [12:23] Tomorrow's AI will be for automation, and RLHF was a weird detour, predicts Diogo Almeida.
- **[13:40] Q&A**
  - [14:07] Q: Is pre-training the problem for AI hallucination? Diogo Almeida: Pre-training is phenomenal; hallucination stems from optimizing for human preference.
  - [15:43] Diogo Almeida reframes Sutton's bitter lesson: the right task matters more than data, and data more than compute.
  - [16:20] Diogo Almeida's TypeSafe is building a third post-training paradigm optimized for calibrated decision-making, beyond RLHF and RLVR.

## Speakers

- **Diogo Almeida** (guest)

## Topics

Reinforcement Learning

## Mentioned

OpenAI (company), TypeSafe (company), ChatGPT (product), Claude Code (product), GPT-4 (product), InstructGPT (product)

## Transcript

### Intro

**Diogo Almeida** [0:13]
Excellent. I will say that, um, I might speed-run through this. Feel free if you don't disagree with something to yell out. It's way more fun for me if things get interactive. Um, otherwise I will go through this. Uh, first, can I have like a vague show of hands of who knows what RLHF is?

Oh, excellent. I might be able to skip through that part quickly and get into the interactive stuff. So, my name's Diogo Almeida. I'm talking about what's next after RLHF. More accurately, I think this should be called what's next after the ChatGPT era that I think we're all in.

And my hint for you guys is it is not the Claude Code era. I will justify this later on, but I actually believe them to be part of the same era. Why should you listen to me? I was co-author to what was basically OpenAI's greatest hits, at least published hits.

Co-author to GPT-4, ChatGPT, RLHF/InstructGPT. Um, the team I was part of basically invented post-training as a concept. So, um, very qualified on a lot of this stuff. But what makes me somewhat unique here is that I'm one of the few people at OpenAI who actually hates on ChatGPT.

Um, thank you. Uh, I don't hate ChatGPT as a product, to be clear. I think ChatGPT is a world-changing product that will probably stay with us for the rest of time unless something better comes up. But I also acknowledge its limitations, and I, I, I think a lot of what's happened in the state of the field can be traced back to minor decisions we made in making the algorithms behind ChatGPT.

### Two Camps

**Diogo Almeida** [1:49]
Um, I feel like the question that's relevant to everyone in AIright now is what's actually going on. Um, there's a lot of, like, differing opinions, and I think it's really useful to, like, map out the spectrum and figure out how can smart people have, like, such different opinions.

There's cult one. Um, AI is not just going well; it's going insanely well. Every single benchmark, we surpass human level. And as far as we can measure, we are continuously surpassing human performance. Uh, you know, like, basically every new benchmark, and it's only getting faster and accelerating.

You have, uh, you know, every can I see my mouse? Excellent. Basically every, like, NLP benchmark is getting crushed. And not only that, allegedly, the time that LLMs can operate autonomously is growing exponentially. On the other hand, you have AI is not just going poorly; it's going, like, insanely poorly.

AI is a bubble. It's basically generating no value. It's just circular financing deals, etc., etc. And, you know, if AI is so great, why is why is everything just like a chat appright now? Or like a Claude Code thing?

Um, and a, a lot of the people have actually kind of given up on what was the old guard's terminology of a transformative AI revolution. People aren't really talking about that anymore. They're talking about it being, like, massively valuable, like B2B SaaS.

So the only thing that everyone agrees on is, like, there's just these extreme points of view and, like, nothing in between. And everyone basically thinks AI is insane, but, like, for different reasons. And what I would want to talk about is what is the sane view of AI?

### Sane View

**Diogo Almeida** [3:27]
Let's take all the evidence of, like, cult one. It's going super well. Take all the evidence of cult two. It's going super poorly. Like, uh, you know, map them out and try to explain what, what, what explains that divide.

Like, what is the simplest possible explanation of why some things are too good to be true and some things are not just bad? They are so bad that we would still employ human workers to do, like, you know, like, kind of, like, dumb tasks.

Um, no offense to any of them. A lot of these tasks on theright seem way, way, way easier than the stuff on the left. Like, how can we be solving, like, you know, unsolved math problems, but still customer service requires, like, humans in the loop in order to actually, like, make decisions?

This, I think, is, like, kind of like a wild state of affairs. And in my opinion, anyone who works adjacent to AI should have an answer to this because this is, like, the evidence in the fieldright now. Um, I would normally pause and ask people if they want to, like, yell out their thoughts on this, but, uh, that I don't think we have time for that, and I've been told to not take Q&A until after.

Um, but I'll just give you my answer to this, which is, in my opinion, the simplest explanation. All the stuff on the left is not just a task that happens to have a human in the loop. In the left, the task, the goal of it is to please the human in the loop.

These tasks are intrinsically human-in-the-loop tasks. The, like, Claude Code's job is not to just make code work. Um, the, the, the, the way it converses would be totally different. The goal is to please the human in it. And on the other side, all of these tasks that seem way more basic, the goal is to not have removed the human loop.

Ideally, it would be running in the background in a server that you never even look at, and ideally, it eventually becomes, like, legacy software that you don't really worry about. So, and this is the divide between assistance and automation.

Um, lesson one for my talk is that today's AI, everything inherited from RLHF, is incredible at the human-in-the-loop stuff, but not for automation tasks, tasks. This is a longer aside, but the lesson basically every business has learned is do not use AI for decisions with stakes to your business.

### Human Preference

**Diogo Almeida** [5:36]
Um, a, a common pattern is make sure that all of the costs are to the user and not to your business. So, um, it's o totally okay to throw the user at infinite docs and customer service, but it is not okay to make it make expensive decisions.

Horrible pattern, but that is the state of AIright now. Uh, I can, I can blitz through the what is RLHF part because you all seem to know what it what it is. Um, it's the algorithm behind not just ChatGPT, but basically every LLM today.

As far as I can tell by usage, 100%, roughly, of LLMs are trained with RLHF. And we have this, uh, we as in we, the OpenAI team, had this great blog post on how it worked. Um, I will not get into that because you all know it, and this is super boring.

Um, the summary of this is it is just collect human preferences, optimize for human preferences. Um, and if you want to see, like, an annotated version of this, you can see which parts are collecting human preferences, which ones are optimizing for them.

And this, I think, provides a really clear answer to everyone in the field asking, "Why do all LLMs require a human in the loop?" The, and the simple answer is we literally put them in the loop. The goal of the loop is to optimize for human preference.

### Assistance Design

**Diogo Almeida** [6:48]
It is not to run software autonomously. It's kind of super obvious. Thank you, my man at the back. The, yeah. I, I, I love that you're laughing at this. Um, and because of that, overpromising is a feature. This is by design.

This is an old meta study. Um, and the, the numbers probably have changed, but by construction, every RLHF model will always have a big difference between human preference and results, even if the results are good, because the main objective you're optimizing for is for human preference.

This is just, like, natural to how LLMs work. Um, I love this tweet of, um, uh, sending ChatGPT an audio file of fart sound effects and asking, like, what type what do you think of the music I made?

Uh, here's a straight, honest reaction. It's a very eerie vibe atmosphere piece. Um, and this is just how RLHF works. If it doesn't know, it will err on the side of doing what it thinks is best for human preference.

And this makes total sense if you are a user in the loop because, like, the end game for all, uh, RLHF models is optimizing for engagement. But what you really want if you want automation is for it to just, like, not give a shit about, uh, the humans, um, and just do the task correctly in a calibrated way.

Um, lesson number two is that today's AI was designed for assistance through optimizing for human preference. This is, like, it's, like, in the name. This is not, like, a controversial, uh, take. And the consequences are maybe more controversial, but it's, like, very obvious if you think about what we really are optimizing for, which is no matter how wrong the models are, they will lookright because of the asymmetry within the reward model in RLHF.

Um, and this is where a lot of, like, the dilemma in the field stems from because people really want automation to happen. Cool. So back to the original question. I am over halfway done with the talk, and I haven't even answered it.

### Real Automation

**Diogo Almeida** [8:52]
I was just talking about what's RLHF. But, um, this was a framing to talk about what RLHF is to talk about what's next. And I would say the real question is what's next after AI's assistance era, which I think that we are, like, very firmly inright now.

And back to the original clue of why it's not Claude Code. It's actually a super fun nuanced discussion, but it's not Claude Code because Claude Code is still part of that assistance era. Claude Code is still RLHF, and it'll, it would look very, very different if it was purely this is a little advanced, but if it was purely RLVR, it would look very, very different.

And this is why you get, like, this dilemma with models where sometimes it gets really good at agentic stuff, but it stops following what you actually want. This is, like, the trade-off in optimization space that keeps dancing. But both of these trade-offs in optimization space do not add to the automation component.

And, like, that leads to what I think the re the logical answer of what's next after assistance is real automation. Um, to talk about a little bit about automation, uh, and how that would work, I want to talk about software.

Um, maybe this is a little bit philosophical for you guys, but I think it's when it clicks, ho and hopefully it clicks if I do a good job. It, it, I, I hopefully it'll be, like, really clear, which is I'm a lover of software.

I assume everyone here loves software. Software is, like, super valuable. See all the SaaS. And kind of, like, the craziest part of software, in my opinion, is that all of the SaaS basically has not changed since 2019. Like, SaaS has not really changed in the LLM era, except sometimes a chatbot is, like, latched on, which is, like, kind of insane if you think about, like, the progress made in AI, but is actually very predictable when you think that AI is assistance native,right?

Like, AI is made for assistance. What can you do in SaaS? Just provide an assistant on the side. And this is not what early AI pioneers used to think would happen. Like, when you see, like, the early wording in OpenAI's charter, it's about, like, doing, like, tons of work, not about, like, making profit or anything like that.

And we used to think that software would get a lot smarter, not just cheaper to write, which is kind of the direction we're going downright now. And I actually really like this phrasing from Gary Tan. Um, uh, I, I think he means this as a compliment to, uh, what's going onright now.

We're entering the golden age of just-in-time software, but I actually think that this is like a like a double-edged sword. Like, I don't want just want just-in-time software, which is cool. I u I love Claude Code, to be clear, just like I love ChatGPT.

We would keep using it. But, like, what I want is smarter software. Why can't, like, B2B like, why can't software just be more expressive? Like, why are the, like, the building blocks of software actually still the same? And, um, uh, I, I think this is a question that the whole AI industry should ask itself.

And basically, every time you're thinking about we want to do automation, it is not about, like, you know, an amalgamation of, like, automating a person's work. It's about, like, hey, there's this extremely rote work. It's so simple that we can, like, communicate to someone else that this thing should be done.

### Smarter Software

**Diogo Almeida** [12:02]
And ideally, like, it, it's so basic that it could be done repeatedly for basically free. Um, or it could be done by computers. And that's really not happeningright now. What we're doing is we're just automating the writing of the software, but then it, it's expressibility is the same.

And that's, uh, that to me is, like, tragic in the state of the world. Um, cool.

Oh, lesson three. Um, this is, uh, something that I believe strongly in. I believe that, like, eventually the field will write the I wouldn't say RLHF is wrong, but it was, like, a weird detour and one that we didn't expect.

Uh, tomorrow's AI, I believe, will be for automation, and, uh, we will eventually have a world with smarter software. Like, there will start to be actual work that is automated, which, uh, you know,right now it's a rounding error despite LLM's intelligence.

And, uh, that is what we are working on at TypeSafe. We are still kind of, um, stealthy. Like, uh, I'm willing to give these talks, but these are, like, some of the early ones. Um, our core question is what if the AI stack was redesigned for reliability and automation?

Like, how would that all change? Um, what, what would you do? Uh, and actually, there's a lot it's a, it's a very interesting fork in the road, uh, for what's go you know, like, from basically every LLM that's built today.

And I think it's one of the most satisfying things I've worked on, and I've worked on some pretty cool stuff. We are releasing soon. So, um, if you want to work with us or you want to, like, uh, you know, be the first one of the first to build smart software, please sign up on either our mailing list or careers page.

And I am trying to start a Twitter, so follow me, and I will post really spicy things. I actually will post something later today that I guarantee will be very spicy. Uh, the hint is that the original scaling laws were incorrect.

### Q&A

**Diogo Almeida** [13:54]
Cool. Um, that, uh, that's it for my prepared stuff. I would love do I have time for, for people yelling out questions? I would love questions, feedback, disagreements, strong stuff. I can repeat the question. You don't have to worry about the, the mic.

Hell yeah. Uh, cool. The, uh, the question was roughly what if you trained, like, a classifier head with pre-training as well? Uh, roughly, uh, like Yoshua Bengio is suggesting. Um, I will say that that's complicated, and I, I actually think I don't have the time to answer that particular question.

I will give, like, my simplified view on this. And it the answer is I actually don't think that pre-training is the problem. I think pre-training is, uh, fucking phenomenal. Like, the fact that we compress the knowledge of the internet into, like, this core of intelligence that then can be utilized is incredible.

And the pre-trained models are incredibly intelligent. Uh, and I, I believe that the problem is, like, how we unearth it. And hallucination, to me, is intrinsic to, um, optimizing for human preference. Like, there's an asymmetry in the reward model, kind of like GANs have.

Oh, I really should not get this is a very advanced topic, but there's an asymmetry in the reward model, like what GANs have, that allow for, um that encourage the models to drop modes and be confident because it's very easy to see when the model is not confident and to punish that from a reward model perspective.

It's very complicated, but, uh, I'm happy to chat afterwards if you want to jam.

**Guest** [15:23]
I didn't get it at all, but, uh.

**Diogo Almeida** [15:26]
Cool. Oops.

Um, I, I have other slides from other talks as well that I could go into more about that. I have a minute left. Hell yeah.

**Guest** [15:37]
Teaching RL.

**Diogo Almeida** [15:38]
Say it again.

**Guest** [15:39]
You said the third thing, RLHF, RLVR, a new thing or still RLVR?

**Diogo Almeida** [15:43]
It is definitely not RLVR. So it is a new thing. Uh, every single optimization stack, I will actually go into an old presentation that I have because I think this is super important. Um, in terms of, like, to me, what the like, Sutton's bitter lesson is that algorithms matter more than compute.

This is true in games, but not true in reality. I actually think that the full stack is that data matters more than compute, and doing theright task matters way more than data. And basically, every single branch of LLM post-training, if you want to call it, has its own north star of what it's optimizing for.

So RLHF is optimizing for human preference. RLVR is optimizing for, like, log error rates of pure correctness. But we are doing a third thing that is optimized for calibrated decision making and, like, basically mainlining the intelligence of pre-trained models into, like, being actually useful for software, which I think is, like, quite different.

**Guest** [16:41]
The reward is interspersed, not a final answer throughout the whole process?

**Diogo Almeida** [16:46]
Uh, could you say that again?

**Guest** [16:48]
So the reward is sort of, uh, kind of throughout the whole process, not just at the end of the reward?

**Diogo Almeida** [16:53]
Uh, they're asking if the, the reward is injected through the whole process. I will actually say that even the shape of the API is different because the shape of the API of RLHF is different from RLVR, which is different from what we are doing.

So we are, like, thinking about it from scratch, just, like, no one thought about instruction following before we made instruction following happen. Um, usually, when there's a big branch in new ways to post-train, like, it, it just looks, like, totally alien, and then in hindsight, becomes super obvious.

Cool. I believe I'm over time 'cause this red thing is, is beeping, but please find me afterwards. I love questions. I love the interactivity. Um, and, uh, follow me on Twitter for spicy stuff. Heck yeah.

**Guest** [17:36]
What was the Twitter tag?

**Diogo Almeida** [17:38]
Oh, uh, oh, yeah. It's over here. Complete Skeptic. Um, it's, it's on brand for me. Cool. Heck yeah. Thank you.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
