AIAI EngineerJul 23, 2026· 23:55

Notion's Token Town — Sarah Sachs, Notion

Sarah Sachs, Notion's AI engineering lead and contract negotiator, argues that AI companies must stop competing on token economics and instead build model-agnostic products that win on data flywheels, orchestration, and security. She advises treating every model supplier as a competitor, because frontier labs charge a markup on a markup for tokens they sell for first-party use. Notion's auto model routes 75% of traffic through a Switzerland-like system that swaps providers underneath, avoiding vendor lock-in. Sachs advocates routing by cost per capability per second, using open weight models for the moderate middle, and reaching for CPUs over GPUs (e.g., no LLM needed to turn a CSV into a PDF). She highlights the 'lethal trifecta' of private data, untrusted content, and external communication as the next security challenge, and demos Notion agents scoping a task, tagging teammates, and opening a PR. Her core message: optionality is leverage, and the product must transcend tokens.

  1. 0:00Token Town
  2. 2:25System of Record
  3. 4:37Cost Barrier
  4. 7:23Product Strategy
  5. 12:43Optionality
  6. 15:04Open Weights
  7. 17:00Infra Choices
  8. 18:09Lethal Trifecta
  9. 19:01Orchestration
  10. 19:46Notion Demo
  11. 22:44Closing

Powered by PodHood

Transcript

Token Town0:00

Sarah Sachs0:22

Okay, hello—okay, before I get started, you guys, this is a huge keynote room. Can everyone, like, come forward? Because I'm talking to, like, 4 empty rows and dispersed people. Do me a favor, I'm spending 30 minutes telling you all of our secrets.

I can see you still. Thank you, thank you, thank you, thank you. We're just going to chat. It's a giant room, and there's 500 of us. This room is way larger than that. Thank you. Honestly, I knew you guys had it in you.

It's really not so hard. Thank you. I also sit in the back. I also work during talks. I get it. I totally get it. I did all day, but not for me. Okay. I'm going to start, but I'm going to still point at you if you're in the back.

Like, you. Okay. I'm Sarah. I lead our engineering teams for AI at Notion. Welcome to my talk. It's about Token Town. How do you go from—not go—from AI pilled to AI poor, okay? I know that today is all about software factories.

We're going to talk about that, but we're going to talk about how to do it sustainably. This is me. This is on my first day at Notion in a very sweaty subway. Like I said, I lead our AI teams at Notion, and I negotiate AI contracts for a living.

My team jokes I act like Anna Winter, so this is a nice—a nice image of me with AI Anna Winter hair after a press article referred to me that externally. And that's kind of the idea,right? How do you think about negotiating between different vendors, making sure that you maintain taste for your company?

I don't do it alone. This is launch day at one of our recent launches. This is just a subset. Any good engineering manager points out that we have a whole company of people building this. I'm just the one that gets to come talk to you about it.

So we've been building a lot. This is an example of our AI usage just in 2026. And we've been really proud of how we've been able to grow that usage. And I'm going to talk to you about how you can build an AI-native product and an AI-native company.

System of Record2:25

Sarah Sachs2:25

But this is just to give me some credit that we're doing it kind of well. Okay. So for those of you that don't know, Notion has always been that durable system of record. It's always been the place where you can collaborate with your peers.

But today, that point of collaboration is a little bit different. It's not just humans. Notion has always been the place for collaboration, and today that collaboration happens between humans and agents, humans and humans, agents and agents.

And we like to think about AI transformations going through this journey. And I'm sure some of you are looking at the slide and wondering where you are. AI as a thought partner is when we all started tinkering. We all started just going to the very first version of ChatGPT on Thanksgiving when it came out, 3 years ago, 4 years ago.

And we started saying, like, how can I send this email to my landlord to say that I shouldn't pay for a repainting,right? And then we'd copy-paste it, enter it into our email. Eventually, we started getting to a place where we could use AI like an assistant.

AI was able to maybe execute individual tasks. That's how Notion AI really took off in the beginning. And it was able to save employee time, but it functionally was limited in its capabilities based on what humans asked it to do.

AI as teammates is what we were really excited to launch almost a year ago now. But this is true in many products. We can do repetitive work and think about a process and have AI do that process. What I think is really interesting is when AI actually becomes that critical workflow where processes are interfacing with each other and you have entire systems running.

How many of you guys feel like you have AI as a system down? Or is you sad, you came up now? I'm kidding. Great. None of you. Exactly. We have found that no one has figured out how to do this well.

88% of people can't even get past AI as an assistant. And why is that? We have a thesis at Notion, it's because there's too much siloed data and not a durable system of record for that point of collaboration.

And we believe that for your software factory to work, for your company to work, and for your systems to work, you need that durable system of record, and that is Notion's mission.

So doing that is expensive. You see a lot of companies that try and commit themselves to this vision. And these are just a series of headlines all within a week of how that's painful. So you can put all of your money into a process to try and make a system, and you end up feeling like this.

Cost Barrier4:37

Sarah Sachs4:56

Right. You end up using a blowtorch to light what is actually a large cigar. But you kind of get the idea.

Cost is a structural barrier to entry. It makes it hard for you to serve products, it makes it hard for you to build factories, and it is ultimately, I would posit, one of the largest reasons why things do not happen at scale successfully today.

And I would argue, for anyone working at an applied AI company, it's something for them to be really familiar with to understand the trade-offs that they're making to build durable and exciting and enlightening product for their customers. But that's not really how the market is today,right?

I'm not going to name names here, but you guys have search engines. You can figure it out. Exhibit A. Every reasoning model gets upgraded. Amazing that per-token pricing is the same. What's not to love? You try it out.

It uses 3 times as many output tokens. Right? Exhibit B. A model gets upgraded, but it has an entire new digit. Right? Whatever markation system that model family likes, it's brand new. It's 40% more than its predecessor, which is being deprecated in the next 4 months.

These are real scenarios that we face at Notion. All of you are nodding because these are common. Pretty much monthly now. But here's the problem. Are you growing 40% in that time period? Are you making 30, 3X more revenue?

No. So how do you navigate this system? If you just auto-upgrade your model and everything that you're doing, you're giving someone a bad deal, either your customers or your investors, depending on how you charge and where you get your money.

Neither are good.

Fortune 5 million companies have the capability to navigate this. They can hire large consulting teams, have durable teams on their own, and build expertise on how to navigate these trade-offs. Most people don't. Everyone else has no ability to negotiate with leverage, and they're stuck in these scenarios.

Right? Part of my job as that Anna Winter joke is to think about advocating for the Fortune 5 million. The non-Fortune 500 companies that don't have the mass to have leverage and negotiate, but need to think about how.

And I'm going to share some of the lessons that I've learned when I have kind of large amounts of traffic behind me that I think scale to those who don't.

This is probably less of a secret now than it was when I started giving talks like this, maybe 4 months ago. Your supplier is your competitor. I know very few people who have convinced me that that's not true.

Product Strategy7:23

Sarah Sachs7:39

You will always be getting a bad deal on tokens with someone who builds them natively. Right? Sometimes the cost of goods served is extremely different. You're basically, they're serving a first-party product, and then you're buying those tokens at a huge surcharge and then selling them again at another surcharge.

That's not really value you can defend. You're getting a really bad deal. And if you tie yourself to one provider, you have no exit. If you build an AI product that you're selling with this structure, you are crossing your fingers and hoping that you are a viable business.

I do not encourage that.

This is really interesting. Dylan in SemiAnalysis posted this. I think it says 8 hours ago. It wasn't at this point. It was probably a month ago. They purchased a subscription plan, and they just highlighted,right, how different what frontier labs charge customers for first-party products are versus what they sell.

It's a bad deal. Don't play this game. Or try and let me know how you win. I don't recommend. Think about everyone else. Think about what that structure means and where you have expertise. I don't think that that's winning on the token economics.

I think it's about product. It's about building data flywheels and understanding your customers better than anyone else. Understanding when you need capability, when you need low price, when you need latency improvements. I promise you, you don't always need what is usually the slowest but the most capable model out there.

And then build compelling UI and orchestration, and I'll show you some examples of that, to justify the cost on the bad deal tokens that you do resell. The job is not to train. I mean, some of you might be training the best model, and I'd love to serve it and come talk to me afterwards, but most of you are not doing that.

Stop trying to win that game and think about the best product that uses many models. Help your customers. Help your team. Bet on the frontier, not on the lab. And we'll talk about what it looks like to do that.

This cost per capability per second trade-off is actually really intense. Citadel came out with this memo a while ago, maybe 2 weeks ago. I loved it. The idea is that for the economy at large, simpler models might be the most cost-effective productivity-augmenting pathway.

They talk about this bifurcation on frontier versus everyday usage. I really believe that. And for every product, the definition of frontier versus everyday, the definition of saturated capabilities or model capability overhangs, depends on your expertise on your product.

No one can replace that. And not all traffic is equal. It is a huge miss to send all of these to the latest Opus model. Some of these absolutely. Large-scale data analysis, when you do it on Notion, we'll recommend Opus.

Right? When you triage an email inbox, if we're charging you to do that on Opus, we're ripping you off and ourselves. Think about where your traffic patterns are. And then think about how frontier lab model providers are structured today.

I mean, it's functionally an oligopoly,right? And that's fine because they're racing to the top, and I think the top is really hard and really important. This is not to say that products don't have a place for frontier difficult tasks.

I want everyone to nod and understand that's not what this talk is about. Understand when you need those tasks, and it's not everything. The problem with those tasks is keep in mind how pricing is incentivized. You can figure out who these players are.

Either you are the best model, everything above what AI can't do today is your market, you can basically price it as high as you kind of want. If you're slightly behind that best model, all you need to be is like a dollar per million tokens cheaper, and you have the rest of the market.

You know that economic theory about gas stations where the best gas stations are the ones that areright next to each other because they cover east and west the most? Yeah? It's the same with model pricing, which means that price does not correlate with capability growth.

So for this complex task, understand what capabilities you need, but be the expert on what complexity is.

And keep in mind that who handles complexity changes. Oftentimes you'll see applied AI companies really be super outspoken on marketing with a specific lab. That's always kind of a red flag for me when they're not model agnostic. Because if you look at this graph, it basically shows that they're behind every month.

Right? The new model and the new model provider of the best frontier capabilities change. And if you hit your ride with one particular provider in exchange for, for instance, a larger discount, you're doing a disservice to your customers like half of the time.

Right? So really think about if that discount is worth not actually having a frontier product.

Optionality12:43

Sarah Sachs12:43

And remember that that optionality is your leverage. If you don't have the capability to walk at any point, you are stuck. And again, I think that's probably the most expensive decision you'll make, regardless of what discount you get or the engineering work to have model interoperability.

One option to navigate this is stay model agnostic. Have different models and capabilities in your system so that at any point, if pricing seems unfair or untenable, you are not out of business.

Notion's auto model does this really well. We have state-of-the-art models available always, but we also have an auto model there at the top that handles about 75% of our traffic. Right? We have the ability to switch between models in our product, and we also offer it to our customers that they have access to these models without vendor lock-in.

That's part of our AI Switzerland approach.

You guys love taking photos of slides. This is the slide. Okay. Model agnostic playbook. This is how you do it. Build for multimodal. It is hard to kill the cache and switch models mid-transcript. I understand that. We invest in that technology.

It doesn't even have to be per thread. Just think about your harness as model interoperability. Think about the cost per capability per second, not just the tokens. Here's a great example. We posted this review when we announced our partnership with Parallel as our web search provider.

If you were to look at just latency of a single call or just cost, Parallel might not be the cheapest. But if you have expertise in entire web search trajectories, you'll see how it differs. The granularity of this eval is what lets us make the best decisions for our customers because we understand all of the trade-offs on entire trajectories, not just single calls.

Switch fast and often. I think we talked about that. And give them something back. That expertise on use cases is also very valuable to frontier labs. We find that our evals and our early access program partnerships actually help us a lot with frontier labs and is something that we can exchange instead of extraordinarily large commits.

And I don't think the discount is ever worth the lost in optionality. That's a perspective you can choose to keep or not. The second option is moderate tasks. Understanding open weights place there. Open weight models are really strong enough to handle these tasks, and the possibility to RL on top of them has also kind of expanded the upmarket growth that they can cover.

Open Weights15:04

Sarah Sachs15:17

I view open weight models as basically lowering the barrier to entry on cost for our customers. And they also give you negotiation leverage. So it's kind of a credible alternative that's putting that downward pressure on pricing that if there's an oligopoly of 2 or 3 providers at the top is unavailableright now otherwise.

I think KIMMI 2.6 was probably the first time that we really saw a model that outperformed 5.2, GPT 5.2, GLM 5.2. Now is another 5.2, bombshell in the villa that also probably does best here. But it's no longer the case where open weight models are good for just SFT on small tasks.

Really think about without RL if they're capable enough for what you need.

And again, don't just think about external benchmarks. Be able to have expertise on your system. What are your tool errors? What's the actual latency that you need? Right? Here's an example of a benchmark that we posted. It's a little bit stale on purpose.

Right? But you get the idea.

Philip at Base 10 showed this slide once, and I've stolen it ever since.

Guest16:27

Where's Paul?

Sarah Sachs16:28

Thank you. Are you here? Buddy. Okay. We'll chat. Hi.

Well, he could come up and say it better. But the idea is that you don't have to be at the top. Right? I'm not trying to make a case that open weight is the best model out there. The case being made, however, is that the gap gets covered eventually.

So if the tasks that you're having today are good enough, then in 6 months they're probably covered by open weight. So be prepared now. And the last thing is CPUs over GPUs. We've recently launched something at Notion called Workers.

Infra Choices17:00

Sarah Sachs17:04

I don't think that the GPU is necessary for every job. A lot of the jobs that we have are actually serving discrete pieces of code. Like you don't need an LLM to turn a CSV into a PDF. You don't need an LLM to talk to Notion tool calls if we have a CLI.

You definitely don't need an LLM to do deterministic SQL queries. This is where people become token poor very quick.

And I think the last option here, besides open weight, CPUs, and optionality, is actually governance. There's a lot of AI governance. One is visibility. Understanding who's using the data. Understanding its maintainability and control. When you have model optionality, you can offer a lot more to your customers.

Here's an example of how that governance works in Notion.

So final tips again. Think about architecture. Think about open weight. And build value that transcends tokens. So we're going to depart Token Town. I know I said welcome to Token Town. We're going to spend the next 10 minutes really thinking about what to do next.

So I think the challenge of the next 6 months doesn't have to do with capabilities. I think it has to do with security. Let's start there. There's this concept called the lethal trifecta. Simon Wilson, I think, crafted this.

Lethal Trifecta18:09

Sarah Sachs18:21

If you have access to private data, exposure to untrusted content, whether it be through ingestion, MCP, email,right, and the ability to communicate externally, and that can include like payloads in a web search, the second you have that system, you're exposing risk.

And in fact, the more autonomous your system is, the more unsupervised this risk is. I think that this is what builds valuable product, not just capability. Same with sandboxes and computers. We talked about this, but it really is something that builds better determinism in your product and also better token economics for your customers.

And multi-agent orchestration. Understanding what agents see and do and what persists. I think persistence of enterprise knowledge is something that's actually really not discussed enough. It's starting to be with some recent launches.

Orchestration19:01

Sarah Sachs19:17

You know, oh, there is audio. So don't have your workflows look like this. And I think this is where most software factories are today. Right? It's like actually your entire engineering time just spends time babysitting the factory. Right?

I mean, I get it. Ours started off like this. Agent orchestration is one of the most difficult tasks of making factories work. So, okay. This is me telling T-Pain to tell people to buy Notion AI. And the reason I included this slide is I am going to sell Notion for a second.

It's my job. Always be closing. Always be selling. Always be hiring. Come find me. But I'm going to talk for a second about how Notion does this. Today we already have the ability to inspect tasks. And you can imagine any task that you look at in a Notion document, you can have Claude actually go ahead and scope out what you need.

Notion Demo19:46

Sarah Sachs20:07

We've launched this manage agent capability today. So if I go ahead to the top of this task, I can actually ask Claude agent to scope out the task. Right?

Ideally, it's working.

And you'll see it'll actually populate an entire spec of what needs to be done. In this example, it's not ready. It's going to ask me a question. Keep in mind this isn't a Markdown file. This is an active document.

Let's say I don't actually know the question, and I go ahead and I ask my team what to do. Imagine that you can kind of tag in your team into these systems. MJ's our PM.

So in this example, she doesn't know. Usually, she does. But multi-agent orchestration is important. Maybe Claude code isn't the best at customer voice, but Decagon is. Right? You can ask Decagon agents. We're proud partners of them as well to collect theright data that you need.

Okay. In this example, we think we know enough. We're going to go ahead and actually iterate through some of this flow. I'm going to skip ahead a little bit. We asked our TL what we needed. He replied. Again, it's a collaborative file, not just a Markdown.

And we can have Claude actually go ahead and spin up the PR. Hopefully, this is looking a little familiar now. This is kind of the vision of software factories. That's what we're trying to host. Okay. Claude put up a PR.

Maybe that's not enough. Maybe I want to go ahead and ask Codex what it thinks.

Great. Found two issues. You can think about this scaling in an actual factory. So today in Notion, you're actually able to orchestrate these agents together, and you're not committing to a lab. You're committing to the concept that AI is augmenting and automating what you do.

This is real. I asked Rajiv if I could post this. This is how it works today internally at Notion. Almost all of our polish and large feedback like this is actually coordinated through our software factories, both in terms of writing to theright teams and also having coding agents take the first step.

Vercel does this as well, from staging to shipping to closing. And we see massive ROI gains from our customers. That's over 3 minutes saved on a given task. Imagine that at scale. So I think we're trying our hardest to think about the factory lens.

Closing22:44

Sarah Sachs22:44

We cannot do this without optionality, and we cannot do this without conviction that we understand what models are required for which tasks. It's really wild out there, you guys. I get it. The market is really young. It's exceptionally opaque.

It's moving fast. I'm super grateful for communities like AI Engineer to bring us together and talk openly about these things and how we navigate it. I think we owe it to all of our customers to get itright and to be critical thinkers about how we navigate this together.

I'm chronically online, unfortunately. You can always DM me on Twitter. You can email me. You can find me after this. But thank you for yapping with me and thinking about this problem, and have a good day.