AIAI EngineerAug 19, 2026· 21:35

From Ambient Documentation to Clinical Intelligence — Chaitanya Asawa, Abridge

Chaitanya Asawa of Abridge argues that everything in healthcare sits downstream of the doctor-patient conversation, and that the administrative machinery around it can be automated. Abridge began with clinical documentation to end 'pajama time' and reached 300 of the largest US health systems in two to three years. For its contextual clinical decision support, quality is paramount because the generator-verifier gap is tiny; Abridge has physicians write independent rubrics, adjudicated into a final rubric, with expert-calibrated LLM judges scoring responses. At a run rate of 100 million medical conversations a year, Abridge decomposes each note into sections and post-trains smaller models per section, betting its unique dataset and narrow problems can outrun frontier models.

  1. 0:00Intro
  2. 2:09Healthcare stigma
  3. 2:58Career path
  4. 5:25Healthcare problems
  5. 6:37Documentation
  6. 8:35Conversation
  7. 11:14Engineering
  8. 12:54Evaluation
  9. 14:41Decision support
  10. 18:01Cost and latency
  11. 20:47Closing

Powered by PodHood

Transcript

Intro0:00

Chaitanya Asawa0:13

Thank you so much for everyone being here. We're going to get started in a second, um, but before we get started, I am curious: how many of you currently work in the healthcare industry in some shape or form?

Oh, that's amazing to hear. Uh, how many of you are clinicians by training? Okay, a couple. How many people in the room are engineers? Okay, awesome. Um, and then how many people have heard of Abridge before? Okay, awesome.

Uh, well, I'm going to let you hear, actually, from our users to start off on a little about Abridge.

Guest0:54

Full day of 22 patients, out by 4:30 PM, notes done, that's nice.

When I think about Abridge, I think the thing that comes to mind is it's really a cornerstone of how I practice medicine today. Um, there's just no way, um, I would do a clinic or see a patient without using it.

I can be present throughout my clinical encounters. I don't have to think about, um, "Oh wait, did I get that? Do I need to write that down?" because I know Abridge has my back and has everything ready for me.

Abridge makes me feel free, because I can really look at a patient, really listen, and not have to be thinking about, "What do I need to put in the computer?"

Full day of 22 patients, good healthcareright now. I know that I have to have those so much, way down in my patient burn.

Chaitanya Asawa2:09

Our marketing team produces really good videos, and so they always hype me up. Um, but the goal of this talk for me, and I know that we have a lot of engineers in the room, my goal is to talk about healthcare as a domain.

Healthcare stigma2:09

Chaitanya Asawa2:21

At least, I felt in the past that there was a lot of stigma around maybe the technical problems weren't as interesting in healthcare. And it is true, in some ways, there's some parts of healthcare that might not be as tech-forward.

A lot of things run on fax machines, for example. Um, but I want to give exposure throughout this talk of two things. One, Abridge's journey from clinical documentation to clinical intelligence and what that looks like. Um, and then two, I want to expose you to some of the technical problems we work on, um, and that have to be the that are truly frontier AI product, uh, problems that have the highest stakes.

A little about me. My name is Chaitanya, you can call me Chai. Um, my career has always been in AI companies and startups. I first started as in research engineering at a company called Vicarious, uh, which whose goal was actually to develop AGI, but they took very different methods.

Career path2:58

Chaitanya Asawa3:12

They wanted methods inspired by neuroscience and probabilistic graphical models, um, and they concretely worked on robotics. Uh, then I started working at this company called Glean, because I faced this problem in my workplace itself. How, uh, like, information scattered all over the place, context is everywhere, and it's so key to decision-making.

And Glean was building basically the ChatGPT for your workplace. As they're about six and a half years, as one of their earliest engineers, as we went from 10 people to over 1,100 people and work with some of the largest companies all over the world.

Um, but that journey was amazing. I love that product, I love that company, I love the people there. Um, but I've actually always really been interested in healthcare. I remember a decade ago, cover of Nature magazine was "AI to Detect Skin Cancer."

I was like, "Wow, is someone interested in AI?" That was amazing. At the same time, I went to the hospital for something, and I remember seeing, oh, coming back, it was a really minor thing, but I remember coming back and looking at the bill and I was like, "I'm not really sure what exactly I paid for."

Um, and so there's these known problems in healthcare, and I'll talk a little about some of them, access to care and cost, and then we had these amazing solutions, so I was like, "Why don't we bridge these things together?"

Um, and remember, this is about a decade ago. Um, and so I actually started this seminar where I invited speakers who were physicians, researchers, entrepreneurs, to, uh, to talk about the space. And what I learned was, while there was really, really cool technology, very little of it made its way into the clinic at that time.

And so, again, my, my journey went a different way. But as I, as I peeked my head out 10 years, uh, 10 years later then, actually our technology has gotten better than ever, as everyone knows, and the AI wave as it's taken over the whole world has also influenced healthcare.

And as you, as you saw towards the end of that video, Abridge, in the matter of two to three years, got its way into 300 of the largest health systems, uh, in the United States. Kaiser, Mayo, Johns Hopkins, Sutter, and so forth, and maybe, maybe you've visited some of these hospital systems.

And once you're inside the hospital systems, you realize there's so, so much more you can do. And I'll talk about that journey that we've had. I specifically work on, um, lead our engineering teams for clinical decision support and our agentic experiences that the technology has now enabled and how we can bring that to healthcare.

Healthcare problems5:25

Chaitanya Asawa5:26

But first, maybe, maybe some of the problems that inspire us as a company at Abridge. One, one of the things that we've noticed, uh, or many people, economists have noticed over the past few decades is actually, in many other industries, you actually see the cost of a good go down, and that's because the productivity has increased.

But in healthcare, we actually see administrative costs have only gone up over the past, uh, few decades, um, and productivity hasn't necessarily increased. Um, and a lot of our problems in healthcare, we solve with labor, but even that, we cannot actually keep up.

So there's this bit of this, like, productivity paradox you might have heard of, like, Baumann's cost disease. And it's part of partially because technology, I think, hasn't fully touched healthcare as much as it's touched other industries to increase that productivity.

A few other problems, you know, hospitals are shutting down, margins are actually razor thin for many health systems. Of course, some patients have massive, uh, medical, uh, debt. And then we actually and the problem that's another problem that's very near and dear to our heart is that we hear all the time that doctors are burnt out, and they actually often don't recommend it as a profession to, uh, to their children.

So what we started as, as a company was working on clinical documentation. So the idea here, if you're not familiar with it, is at, at the end of every patient visit, the, the clinician must create a note. Uh, our typical format is a SOAP note that has a couple different formats, like, what's the chief complaint of the patient, and a few other sections, and then what's the assessment plan, what do we do with this patient.

Documentation6:37

Chaitanya Asawa6:59

You have to do write this after every single visit, and there's some different variations on this de-depending on specialty. Um, typically, clinicians end up often doing it takes like two hours a day to write just write these notes, and you often do it what's known as pajama time after work itself.

Uh, and that's a common source of clinician burnout, spending all this time outside of work, and it's not the most fun part of the job. However, these documents are actually extremely high stakes, because these clinical notes are often used as a basis of billing, but also, uh, which is, of course, financial things are high stakes, but also have clinical impact.

And the reason for this is because these prior note these notes are used for the next clinician, or as you switch health systems. They use it they provide context to the clinician of the, uh, patient's longitudinal medical record.

So it's actually really high stakes to get thisright. Um, we started here because we it's a known pain point that's existed for many, many years, but finally the technology's caught up to do really, really high-quality medical notes that's actually personalized to the clinician.

This was an amazing wedge into healthcare industry, which has typically been technology reticent, because it led to they actually care a lot about getting these notes high-quality andright. It led to higher doctor satisfaction, uh, satisfaction, and such they could actually see more patients.

Um, and it can actually help create a higher record, um, that helps prevent, uh, as it relates to billing, auditing, and other reasons. So we started there, just this product alone scaled to 300 hospital systems. But I want to show you a little about where we're going next.

Conversation8:35

Chaitanya Asawa8:35

And, um, and the core thesis of the company is that everywhere in healthcare is everything is around the conversation. So we started over here with the conversation to clinical note. Everything else is downstream of that, whether you it relates to billing, whether it relates to things like clinical trial matching, or whether it relates to clinical decision support.

It's all about the conversation. That sacred doctor and patient conversation, and we've just built all this administrative machinery around that. But how can we bring it back to that conversation and actually automate some of that, uh, administrative machinery?

So to give you a tactical example of what this looks like, um, and I'll play this video.

Of where, where we're going from here.

Guest 29:28

We've been building a solution that allows the physician to interact with Abridge directly by using their voice. "Hi Abridge, is Nathan eligible for any clinical trials?"

Guest 39:41

He may be eligible for the Abridge HF study. He remains symptomatic despite maximal therapy, and most screening criteria are already met, but an updated echocardiogram is needed to confirm his ejection fraction and complete eligibility assessment.

Guest 29:53

"Allright, please order that echo for him."

Guest9:56

Done. Confirmatory echo ordered.

Guest 29:58

And then when I'm done for the day, I can just ask Abridge, "Hey Abridge, can you prepare my charts for tomorrow?" and Abridge is working for me.

Chaitanya Asawa10:10

Uh, pauseright there. One of the things that you'll, you'll notice is that we are thinking about how to, uh, revolutionize the entire visit, uh, for a clinician. From pre-visit, how, um, earlier in the we have suggested discussion topics.

Here's things that you can talk about, uh, with your patient, whether they're clinical or more, uh, billing-related. Um, we have after the visit, we actually create everything for you. The patient visit summary, the actual clinical note, and we actually pend orders, as you might have seen.

We're able to use we're able to do this all by reading all this context. We have access to all of the EHR context, so we know everything about the patient, the prior labs, the prior notes. We have access to the live conversation between the doctor and the patient.

That's where the quote-unquote debugging happens in healthcare, where you learn about what the patient is facing now. Um, and then we have access to world's medical literature that we can ground and clinical guidelines that we can ground all of our work in.

So.

Engineering11:14

Chaitanya Asawa11:14

I want to I want to switch now, given the context of where we're going as a as a product, I want to switch into some of the key technical and engineering problems we face and inspire you on some of the what I think are very much frontier AI challenges.

So if you've ever worked on an agentic product before, the regardless of vertical, um, the three KPIs that tend to matter are quality, and then latency and cost. In healthcare, I feel that we're actually playing on hard mode for all of these three KPIs.

This is a high-stakes scenario, especially when you're doing something like clinical decision support. You have to beright, because the downside is extremely high when you're wrong. When I used to work at Glean, you know, while I love that product, I could be wrong and it would have been fine.

Maybe we answered a question incorrectly. But in healthcare, if we answer something incorrectly, there's actually consequences and we entirely lose our trust. So quality needs to be absolutely high, and I'll talk a little about how we keep that bar high.

And then latency and cost also really matter for us when you're live in the conversation. You can't, uh, with latency, you can't act on information too late, and you have to act on the also at theright time for it to be useful.

And then finally, cost at the scale we're, we're doing this at. Um, and, and as an interesting aside, it actually relates to our we have a motto inside the company that our goal is to save lives, save time, save money for, uh, for the hospital system and for the healthcare industry as a whole.

And it actually I think it's funny that it really maps to the three KPIs that you care about in an agentic product. So talking a little about quality, how do we keep that bar high? For us, we really treat evals as the operating system, the lifeblood of the of the company.

Evaluation12:54

Chaitanya Asawa12:54

This starts from internal benchmarks and offline evaluation. Before we develop any product, you know, whether we're talking about clinical trial matching, clinical note, clinical decision support, coding, we start with a robust set of internal benchmarks. This is pre-deployment, and then we test that against, you know, things that we've actually seen in the wild.

Then we all have a staged rollout. We know that we need to make contact with reality, reality. Not everything offline will perfectly represent what happens in practice, and so we slowly roll it out. Uh, maybe it starts with the alpha group of clinicians that we trust and understand the stakes.

We rely to beta, maybe there's A/B testing at scale. And then even after it's fully rolled out, you always need continual monitoring. Again, the stakes are really high, and you cannot get away with just being like a prototype that you just ship out there and be like, "Yeah, I mean, I tested it on a few cases and it works."

How we do this is we always have expert-calibrated LLM judges. So we have clinicians embedded throughout the entire company. The clinicians are our domain experts, but not all of us are clinicians. I'm not a clinician. So how can we the rest of the company still move fast is by encoding that clinician judgment into LLM judges.

You know, I think a really great evaluation system has a property that it reflects the behaviors that you want in your product. At the end of the day, we are making a product for clinicians, and so who best other than our clinicians to actually create our judges that represent what they want.

And those judges, once you have that, create a feedback loop so that anyone, whether you're a clinician or not, can actually, uh, hill climb and learn from that. We also have a lot of online signals, whether how you're editing the clinical note and your typical thumbs-up, thumbs-down, uh, and star ratings and other free-form text.

So this is a general framework we use for our for all our products. I want to deep dive into the product that I work on, which is clinical decision support. So to give you an example of what clinical decision support looks like, and specifically we're building something novel, which is contextual clinical decision support.

Decision support14:41

Chaitanya Asawa14:59

Maybe a provider asks a question like, "Hey, does this patient meet the criteria for febrile neutrophilia?" Um, and so what we have to do here is actually a lot of context is under-specified in this question. We're so the first thing we're going to do is we're actually going to pull from the EHR data, previous, uh, previous lab values.

Then using that context, we're going to, uh, use the, uh, uh, we're also going to use the live conversation, and we're going to use clinical guidelines and medical journals to come use as the reasoning sources for combining all this context to actually answer the provider's question.

Now, but I want to focus on evaluation. Again, the stakes are really high here. We, we really can't get this wrong. So how do you tell whether or not an answer is correct? And sometimes I, I and this is a case where the generator and the verifier gap is really small.

What I mean by this is, in some problems in AI, such as like Sudoku, it's really, really hard to generate a solution to Sudoku, but it's extremely easy to verify, uh, once you do have the solution. And that makes, uh, that makes it much easier to hill climb against, uh, and build evaluation for.

But in a case like this, the generator and verifier gap is really small. If I had a really, really good generator, uh, verifier, then that would just be my generator itself. So how do I create a reference that isn't just a language model itself and ground itself to I have trust?

So what we do is we tackle this by having many, many different signals. Uh, we have a clinical quality judge, which I'm going to dive deep into, and then we tackle from, uh, we have many signals from a boundary and adversarial judge.

We have a clinical safety judge. And we also have judges that re-represent product ac-aspects, like tone and style, and that matters a lot as well for AI products. So all of these are different signals that try to get a piece of this like really, really hard-to-measure problem and guarantee it in the way that we want the product to be.

So diving into the clinical quality judge. So again, I said the generator-verifier gap is really small here. So what we need is we actually need human references to tell are we generating theright thing. But you can't just create a human golden response, because there is a lot of variability in the potential responses.

So what we did is we took a lot of real clinical cases. We had independent physicians create a rubric. So this rubric said elements of what we wanted in the response. So it's not, "Here's the exact response," because again, there's many infinite possible responses, but a good rubric elements that what a good res-response would look like.

And then we had a separate physician that actually adjudicated it, brought these two independent rubrics together, created a final rubric, and we actually had a fourth clinician do QA on these rubrics. Once you have these rubrics, and here, here's a sample rubric what it looks like.

Here's actually a question, and there, there's more context than the case itself. You have a rubric of what are the elements that a response should look like. Now we can actually have an LLM judge that compares our agent's responses to these rubric elements and does some basic semantic match to tell, "Hey, is our model performing well as we continue to hill climb, whether it's our agent architecture, our models, or search ranking algorithms?"

I want to talk a little now about cost and latency, two other really, really hard problems for us. So we do this, as we said in that intro video, we do this on the live in the conversation, and we do on the run rate of 100 million medical conversations a year.

Cost and latency18:01

Chaitanya Asawa18:13

How do we do this in a way that doesn't really break the bank for us? So one, one place that this problem comes up or is, is actually in generating the clinical note. So when you're generating the clinical note, there's many different sections to it.

There's a history of, uh, present illness, past medical history, and there's the assessment and plan. So one of the core insights for us is, rather than say using a foundation model to generate all of this, is we can actually break decompose this problem into simpler, smaller workflows.

Healthcare is actually many specific workflows. You don't need, you know, Fable 5 to actually solve all of your clinical notes. We, we don't need frontier-level intelligence for every problem. So we actually post-train a lot of smaller models for different problems, such as different actually even to the granularity of different sections in the clinical note.

And that lets us use much smaller, uh, models, because it's a mo-more specific problem and at, uh, at much cheaper cost and latency. And we have this data flywheel that we have this unique dataset of 100 million medical conversations a year.

And as far as we know, no one else has such a ho-uh, large dataset. So our key insight is having aright to tr win in training models. There are problems where the quality's already maxed out, and so you should train models then to reduce quality and latency.

But there are other problems where the quality isn't maxed out, and people say, "Oh, the frontier model will just steamroll you." Our key insight is we can actually potentially beat the rate of change on the frontier model if we have theright, uh, to win by having theright data that they may not have and the focus on a problem that they may not be focusing on.

And that lets us still maximize quality. Another, uh, problem that I'll quickly touch on is in-visit orders. So, uh, doctors really aren't a big fans of pending orders, but often they'll mention orders during the visit itself, medication or non-medication orders.

So what we have is while we're listening in the visit, as the, uh, clinician says, uh, order, we actually queue it up in the background, uh, and let them, uh, actually sign it off in the EHR. But you can imagine if we did this in a very naive way, like every few seconds are just listening for orders, that would really break the bank.

Um, and so a lot of our tricks are like, how do we find theright events in the conversation to actually trigger heavier models that will actually do the order matching, because you need to match the order, not just is the order said, but does it match and reference lists of orders that are approved by the system and are relevant to the conversation.

So we have a number of different gates that are cheaper and faster that let us trigger actually larger models and hand off to them for actually doing the end-to-end work.

Um, but the last message, this was, of course, a very quick talk, but the message I want to leave you with is healthcare is a domain that needs frontier AI and actually puts it to the test at, uh, higher stakes.

Closing20:47

Chaitanya Asawa20:56

In the past, I, as an engineer myself, was wary about working in healthcare. Does, does like, does healthcare technology actually work? Well, Abridge has proven this at scale, for sure, and I as hopefully I gave you a taste of some of the frontier problems that we work on.

So thank you so much. My name is Chaitanya again. You can, uh, and feel free to connect with me at Twitter or LinkedIn. Thank you.