CPO to IC0:00
Joining us on stage is the co-founder of Instagram and a member of technical staff at Anthropic, Mike Krieger.
How's everyone doing? I mean, good morning?
Good morning.
Nice.
Uh, Mike, thank you for releasing "Fable" just in time for us.
Exactly for the conference. We timed it.
Um, we're— we're so glad to have you. You're one of the preeminent builders, and you're leading labs at Anthropic. How has your model usage changed as you've, you know, seen models internally grow?
Yeah. I mean, for me it's been, like, both the model shift and then my role shift. So I— for, like, the first 2 years I was at Anthropic, I was chief product officer, and then I kept seeing people build with the models, and the FOMO just kept increasing because I was, you know, using the models as much as possible.
But for example, on product strategy, I would write a strategy doc and then have Claude critique it, and maybe you can use a workflow, but it's not quite the same as, like, building in that pure way. And I was, like, spending all my weekends trying to build with it, and then I realized, okay, I actually just need to shift.
It's, like, way too interesting a time. And it's actually an interesting trend I've seen now, like, several people that were CTOs at other places are, like, now joining as ICs at Anthropic in other places. But I made a role shift, and it was actuallyright around the time where we started getting sort of internal snapshots of what became Mythos and Fable.
And what was really interesting watching that sort of shift was, um, that kind of change between I have an idea, I'm going to, like, sort of break it down in my head, much more how I would do engineering normally, and then kind of iterate through these different steps, to moving to much more of the paradigm of, I'm going to describe the goal, like, go off and work on it, and then, like, we can talk about what tradeoffs.
You, you know, surface some questions along the way, but then figure out where you landed and where we can go from there. I find it's hard. I don't know if people have this experience where— and I know Fable's only been re-enabled for a couple of days.
Fable's definitely way, way smarter than me, so sometimes it'll finish work and be like, "Here's the tradeoffs I made." I'm like, "Can you explain it to me like I'm a little dumber than you are because I need you to, like, sort of break this down for me?"
But that's been one sort of big change, is sort of moving from that task delegation to, like, express the end state and then have it go and cook on it.
Be unreasonable2:35
Yeah. We're all learning how to delegate better. Tariq did us a huge favor yesterday. We— do you want to read it in the newspaper? It's, you know, we have write-ups of talks now in, like, the next day's newspaper.
He said, "Be unreasonable." In what ways, you know, have you been more ambitious with your prompting?
Yeah. I love that. I mean, I love that framing. We actually just hit this today. One of the labs initiatives I have is internal products, and somebody was like, "Hey, it doesn't work the way I want it to, and can you make some changes?"
And I realized I'm just going to go ask Claude to do this. Like, why don't you ask Claude? And this was a non-technical person. And so I actually think as an industry, or even as a product team, we have to teach people to be more unreasonable in their usage, and it's sort of hard to imagine.
And I think that— if I can digress for a second on product design, I thinkright now the, like, kind of first generation of AI products, we put them too much in a box and constrain their, their sort of access to tools or kind of degrees of freedom, which means it was much harder to be unreasonable,right?
When you say, "Do this thing for me," and then it would be like, "Oh, I can't."
You can't.
I can barely, like, I can write code, but I can't really run it, or I can kind of introspect my environment, but not really. And I think as you see our own, like, product progression, even with things like coworker, like, you know, does every single, like, knowledge worker need a virtual machine that can write Bash?
Like, on the face of it, no, but then when you realize, "Oh, actually, that way it can remediate an issue where, oh, I tried to parse a PDF using our built-in PDF parser. I hit this yesterday," and it was like, "Oh, I can't parse it this way."
Well, okay, well, I can probably write a script that can do this as well. So I think that's it. My most unreasonable thing, though, was one of our labs projects I wrote in Python, like, near and dear to my heart.
All of Instagram is in Python. I think they're finally converting it to PHP now that they have, like, models that can do it.
Oh God.
I know. It's going to be some tokens. And for deployment, I realized that Claude coded, like, figured out a better deployment story with Bun, and I was like, "Okay, I need to port this whole thing from Python to TypeScript."
Like, as a, you know, if I put on my, like, 2010s engineering hat or even my early 2020s, like, that's a dumb idea. Like, who would ever port, like, at that point, you know, a couple hundred thousands of lines of code?
Weekend port4:51
But I was like, I think this is doable now. And I basically created this dynamic workflow setup, and over the weekend had it port the whole thing, like, verify it, double-check it, then read both code. Like, just basically churn and churn and churn, and then came back Monday to complete a workflow that was a ported version of that thing.
So that probably ranks on, like, the more unreasonable things. Like, yeah, just port this entire Python code base to TypeScript, get it working, get it deployable in, you know, a weekend.
Yeah. I mean, a lot of people are talking about the Bun Zig to Rust version.
Yeah.
I think a lot of people are also like, "Well, it's a compiler, it's a runtime, it's got lots of tests, easy to do. Can you port Instagram, which you know very well, to PHP like that? Like a product?"
Yeah. I mean, I think the product side, it's even— I don't know if easier or harder. One of the things we did at Instagram, this is when Python 3 came out and we were able to add type hints for the first time, and it was— people had a lot of internal conversations, like, "Are we going to run out of steam on Python?"
And my perspective was always like, "I think we can take this way further than we think we can, but I think types are going to help us not sort of be in our own way." And we built this thing called MonkeyType, where we basically, like, captured runtime types, like, basically the types that were actually getting used in production, and then mapped those back to the types in the code base.
And I think because of that sort of pattern, I think there's really interesting ways in which if you're doing sort of conversion or sort of cross-compiling using LLMs, you can also lean on production data a lot more or run sort of, like, segmented tests.
I think that, like, there's a lot of things you can do there. But yeah, I think it's— I mean, the sky's the limit there as well. I think the hardest part is always finding the boundary around where you can start doing it incrementally without trying to boil the whole ocean and, like, swap it overnight.
Yeah. I mean, your users are your tests ultimately. And, you know, I also read another article in the newspaper about how you can just use rollouts, and sometimes you don't really know what you're going to need it for, but when that infrastructure exists for your experiment and to roll things out, it's enabled so much.
Yeah. I mean, I always found— this was advice we got. It was, like, we launched Instagram, and the happened to be the first week everything melted because we didn't really know what we were doing on the backend side of things.
Scaling lessons6:48
And coincidentally, that week there was, like, a lunch that one of our investors just scheduled, like, not even for us, it was just an infrastructure lunch, and we ended up spending— we totally, like, monopolized that conversation because everybody had their own opinion about how we could fix our scaling.
And, like, the two pieces of advice I got there is, like, 2010 that I've, like, will forever retain is, like, basically, like, pre-measure everything that you think you might even remotely need because the worst thing is an outage.
We were like, "Well, is this, like, number normal or is it high?" And, like, "Oh, I don't know because I don't have data until I just added this metric." And the other one is being, like, really thoughtful about knobs and feature flags.
So even, you know, early Instagram, we had, like, a very simple but really effective, like, way in which you could do, like, ramp-outs and rollouts. And dynamic config too, where, you know, a lot of our runtime configurations had to be changed, you know, in a matter of seconds so that we could handle load.
And being able to, like, do that in a first-class way was really important. I'm seeing that definitely in AI as well where, you know, we're making all sorts of different trade-offs, and having that kind of runtime configuration is super key.
Yeah. My favorite scaling story for Instagram, by the way, I think it's, like, your launch day when you DDoS yourself with email.
Yes.
Which people should look up that story if you haven't seen it. I wanted to go into tags. Very, very major ship. It's how 60-something percent of your code is written today.
Tagging Claude8:06
Yeah.
How did you square that with everything you just said where it's, like, very dynamic? Like, you don't actually ship one app. You ship one app with 3,000 flags.
Yeah.
And, like, well, what are you working on today? I don't know. Like, it's for this segment of the population.
Yeah. Yeah. I mean, I think there's a bunch of things. So, like, with— I was really excited. I was talking to Swix earlier. Like, I'm really excited that we have tag out there because it is how we've been working for a while.
And I would get up on stages and people would be like, "How do you work at Anthropic?" And I'd be like, "Oh, yeah, we use these things, like, that are not quite Claude code, but, you know." But it's hard to describe it.
But, I mean, if you, like, got to poke into Anthropic, like, you would see, of course, Claude code usage for things that are, like, more interactive, or if you're kind of iterating on a particular sort of specific thing where you want a lot of, like, high sort of bandwidth back and forth.
But most usage is actually much more delegating via tagging and via tag. And you can say, like, "Here's the—" and the reason it's really interesting is how multiplayer it is. And it reminds me sort of of, like, actually, like, Midjourney, like, the fact that everybody was on Discord seeing how other people were using it.
I think it actually, to your earlier question, really helps with that unreasonableness or ambition where the first time you see somebody tag Claude and be like, "Hey, you know, don't just fix this bug, but, like, now you are responsible for this part of the code base, and I want you to monitor this feedback channel and proactively take on tasks and then fix them, and then also take, like, you know, if this API changes, do that."
Like, I saw somebody do that. I was like, "Oh, wait, I've been totally underutilizing this thing. I've just been using it as, like, a glorified Claude code in Slack." Like, that's definitely a totally, like, sort of noob version of it,right?
And then the more advanced version is really trying to start thinking of it as a teammate that is actually sort of holds context, has memory, and can be proactive. And that's just really changed how we operate internally. It's much more like this multiplayer, async, proactive way than it is a, you know, most people off in their own CLIs.
Are you bottlenecked by code review and Git? Obviously, there is Claude review, but someone usually still looks at it. Is there a world in which you just merge it in?
Yeah. We're— it's a really good question. We are definitely still bottlenecked on reviews, especially for things that are, like, touching some architecture pieces. And it's actually more subtle than just being bottlenecked on review because that's, you know, okay, we can carve out time differently.
It's, like, bottlenecked on human ability to even, like, fully conceptualize what we're doing. So one of the reasons we built Claude code artifacts that we shipped a couple weeks ago was partially for that, which is you would send somebody a PR, and then they'd be like, "I don't know, man.
Review bottleneck10:37
This is, like, 2,000 lines of code. Like, it looks like code to me." And what we started doing instead is sharing much more, like, "Here's a Claude code artifact. Like, here's the explanation. Here's the intention of the change.
Here's the trade-offs that were made." And, like, I think that's going to much more be the trend by which we communicate, which is the code is ultimately, you know, verifiable using some things, but actually, like, discussing intent and trade-offs and then measuring in production is, I think, at least the direction of travel we've gone in.
I don't review. When I get a pull request, I wish I could say I've reviewed every line of code. I definitely do not. I, like, actually talk to Claude about the code and say, "Allright, like, these are the questions that I would have.
Can you go investigate it?" So it is kind of Claude-powered code review, but still human-driven, and for the really important ones. And for the ones that are, like, cosmetic visual changes, it's much more like, "Look, like, we'll fix forward if we need to fix forward," you know?
Yeah, totally. I think a lot of people here are trying to figure that out too. I wanted to talk also a little bit about Anthropic Labs in general. Nilay Patel, who you've probably met before, loves to ask the question, like, "Draw the org chart."
Yeah.
Like, how, like, people, you know, you ship your org chart. Like, I think it's important. Like, everyone knows Claude code. Now you've got tags. How are you structuring the labs?
Persevere or pivot11:53
Yeah. It's a good question because what we were trying to wrestle with was you want sort of people to be supported. Like, you know, I think the death of the engineering manager discipline has been greatly exaggerated. Like, I think there's still a lot of coaching and interpersonal pieces and personal development that I think is still really, really important.
But especially in a labs-type group where, like, our whole cadence is two-week reviews where every project goes up for, we call it persevere or pivot. So basically, every project goes up for review, and either it's time to, you know, keep going, persevering, or, you know, it's time to pivot it or even shut down.
And, you know, we've shut down projects basically every single one of those cycles. And it's like, the more you do it, the less it's just like, "Oh, no, my project has shut down. I failed." It's like, no, that is definitely the intention of the labs team is to prototype quickly, try to ship internally, maybe get it to early access, and if it doesn't work, wind it down.
But because of that kind of, like, rapid iteration, it means that if you align the org chart too much to the individual projects, you're going to end up, like, reorging every two weeks, which would be a total nightmare.
And so we've actually ended up with this interesting setup where, like, the pod or the team that is working on a given, we call them bets, within labs definitely just draws upon, like, allright, somebody from product, somebody from the eng team, you know, I'll jump in when it's a product I'm particularly interested in, and I'll come in and work together with the team on it.
And that's the unit for that time. And there is the concept of a bet lead or a directly responsible individual. But the interesting thing is that they don't manage usually any of the other people, which kind of breaks the kind of previous way in which a lot of these things were done.
But I think it leads us to be really flexible when you say, "Okay, actually, this project is not going to work out. Let's disband and keep going, and it's not a big deal." And the eng manager is much more playing the, like, make sure every individual is assigned to the thing that they're most excited about and that they're working in the best way possible.
Now, what we do sort of solidify is when there's a product that has, like, legs. Like, Claude Design, for example, started in this sort of ad hoc sort of group way. And then now that, like, we've shipped it, it's gotten traction, we've done, like, a big second release in June, like, it's becoming, like, we've hired people for that specific team, and it has more of a structure.
Claude Design13:43
So it's, like, loose until it gets solidified down the line.
What's the future of Claude Design? I think a lot of people are very interested in it's one of your biggest launches this year. Where does this go?
I think for me, I mean, the things that are holding back Claude Design from being even better is better interaction with our other surfaces. So, you know, I was designing something or I was talking to Claude code the other day.
I'm like, I want a really much more seamless, like, what I'm talking about, the design for it, you know, interactive design back to that. And I think in general, it's, I mean, this goes back again to kind of unconstraining Claude.
Like, the fact that our surfaces don't talk to each other as well as they could, I think, really holds back a lot of interesting ideas around what we could do. So I think that's one, like, kind of major area that we're looking at.
And then the other one is people, like, the lines between a Claude design and an app get blurrier and blurrier over time. Like, I've seen people, of course, there's no, like, persistence, but build, like, fully functional, like, even games, which is definitely not what we design Claude Design for, but you can do it.
It's just HTML and JavaScript. So blurring those lines even further and thinking through, like, what is the path from a, like, fully featured design that looks really well to really good to something that is maybe more like an artifact where you're actually able to go and, you know, persist data and share it with others and build from there.
So I think that those lines get really interesting over time too.
Yeah. A big part of design is having taste. I actually asked Fable what Fable wants to ask you. And this is what Fable came up with. You deleted almost all of bourbon to get to Instagram, which is, like, you had a whole, you know, solo, mo, whatever thing, and you went to Instagram.
What would you delete in AI? Or more spicily, what would you delete in Claude?
Project Unship15:37
Oh, I like the spice. I think, I mean, we have it's interesting. We have one of our Slack channels, like, Project Unship, which is, like, what is in the productright now? And, I mean, this is hard at Instagram.
The Instagram, we some things that had, like, 4 to 5 percent usage. You're like, "Oh, that's really not very many." But then you have, like, 20 features that each have 4 to 5 percent usage. It's like the classic Microsoft Word problem of, like, everybody uses some disjoint subset of the functionality.
So that's always the challenge. Now, I think we're a younger product, so hopefully we have less of those things. Like, we unshipped styles, I think, recently where it was, like, used by a small percentage of people and was not really AGI-built in a lot of ways.
It was, like, very sort of prescriptive in the way that it worked and skills were a much better application, something like that. So I think you have to be willing to take the primitives of, like, one generation of AI and, like, unship them or at least, like, supplement them or supplant them with the next one as well.
I think the biggest thing as I look at it, and I've been spending some time, like, outside of labs on some of this is, like, man, like, we're asking people to make, like, code versus cowork versus, like, chat distinctions.
And, like, one, they don't interoperate well and they can't delegate to each other. And two, I think the average person off the street could not explain to you why those surfaces are all different. So I think deleting some of the product complexity within our code or our product, I think, is a thing that would serve well.
Also because then Claude can do what it needs to do and do well. Like, there's nothing more frustrating than having a cowork session where you're like, "Great, I've mapped out exactly what I want you to build," and then be like, "Can you please, like, create a paragraph that I can paste into Claude code?"
Like, that is some 2020, you know, kind of workflow there that really shouldn't exist anymore.
Startup advice17:17
Yeah. I think drawing lines on what you don't want to do and also sort of leaving room for others is interesting. A lot of people, today is, like, the startups day for AIE. You are obviously very sympathetically aligned to startups, but there's some anxiety in the room because tomorrow Anthropic could wake up and publish some Markdown files that destroy my industry.
So why should we not all just give up and join Anthropic? Like, why bother starting any other company?
I mean, I actually joined. One of the main reasons I joined Anthropic was because I saw how much this was, like, you know, the models weren't that good at code yet, but they were getting there. Like, how much it would unlock, like, hold, like, next generation of startups.
Not because it was going to solve their ideation or their taste, but because, like, it would make experimentation way simpler and would get you to move faster. And I still, like, really believe that. And, I mean, it's the reality of, you know,
and we saw this with, like, Instagram. Like, we would get questions in investors like, "Well, what happens when Google launches a Photos product?" It's like, Google is going to launch a very Googly Photos product, and it's going to have to be bound by the integrations that they already have.
And it's going to be, like, it's going to play to their strengths. And I think that is going to be true. I'm not, like, giving advice on how to compete with Anthropic, I guess, in a way, but, like, it's actually not because we're also a platform, which is, like, there is so much, I think, room to be, like, laser obsessed with your particular vertical or your industry or a group of people that you know really well in a way that, like, none of the labs are ever going to get to that level of understanding and, like, therefore get that kind of adoption and user love and build that up.
Now, it's definitely harder in the age where, like, the models can just do a lot. And so there's, you know, some of these things can be, like, skill-ified and, like, maybe don't need their own dedicated product. But I think it's, like, the hard stuff is still hard.
It's, like, understanding the needs of people, like, figuring out how you're going to reach them, listening to them and iterating on them really quickly. Like, it is still the case that, like, a group of four or five people obsessed with a problem is going to move faster than those same people at any other kind of organizations that are, like, you know, subject just to the complexity.
I just mentioned the, like, the fact that we have, you know, a lot of different products that kind of interoperate. Like, that's an interesting constraint that we have to work through. It's an advantage in other ways,right? So, yeah, I'm still, like, very long and bullish on startups.
And it's just it papers over the fact that, like, writing code was never the, like, the limiting part. You know, maybe it was on the timeline perspective, but it was never, like, the thing that was going to, like, make or break your startup.
It's really that space and user understanding.
Yeah. Domain knowledge.
Yeah.
Vertical AI19:44
Today is also our day for Vertical AI. One of our returning speakers and top speakers, Chris Lovejoy, was always talking about Vertical AI. He was from Interior in the healthcare space. And then recently, I invited him back and turned out you guys just hired him for your healthcare efforts.
We also, our next big one is also finance. You know, we have an AI and finance track. You guys just had a huge finance event in New York City. And where our next AIE is sort of finance-focused. What are you seeing there?
Any, you know, any potential for Claude? Obviously, a lot of Excel spreadsheets.
Yeah. No, I think that there's a lot in there too. And that's, like, an area where you could see the model get clearly better at it, like, sort of generation to generation. And there's, you know, there's some good sort of vertical-specific finance startups that have, like, done their own evals, which has also been interesting to track.
And it's not like we're, like, sort of playing to the eval, but it is a useful sort of barometer on, like, is this actually getting better at these finance use cases? I think the interesting blend that's going to happen is this mix of, again, the model having the flexibility to, like, dive in and create just-in-time analyses or dashboards or workflows with, like, some sense of, like, what is the not immutable, but at least, like, verified sort of set of data.
And so, like, having all of that be totally free-form, I think, is a recipe for confusion and is, like, not what most companies in the financial services space want. So finding thatright sort of cut line where you have verifiability and audit logging and sort of data provenance here, but not in a way that constrains the kinds of applications that you can build on top, I think, is a lot of the art that we're seeing in that space as well.
And I think, you know, if you solve it well, you can get the best of both worlds. The hard part is a lot of the systems that were built to do the verifiability/auditability are kind of almost by design not super flexible in terms of agentic workloads on top.
So I think there's opportunity at both sides of the stack there.
Burnout and feelings21:41
Yeah. I think I also agree. We'll be exploring that in New York. The last thing I want to end on is on mental health, which we don't talk about enough in technical conferences. You've seen a lot of hypergrowth.
People are just always refreshing their timelines, and it's exhausting. How do you advise people who are working 996 to avoid burnout?
Yeah. I mean, I think this is a hard one. I mean, and it is, I'm sure you all are experiencing this because you're all working in this industry. Like, it is, you know, multiples more intense and things move much more quickly.
Like, on Instagram, like, the two things that we were thinking about was, like, what is Apple going to announce at WWDC, and is it going to, like, totally mess us up or boost us,right? So that's, like, once a year.
Or, you know, a competitor launches every three or four months,right? And it is definitely not that. It is Anthropic. We do our weekly all-hands. It's usually on Wednesdays, and we have a slide that's, like, the week in AI, parentheses, and it's only Wednesday.
And, like, inevitably, like, some competitor has shipped a new model and, like, there's been, like, a new product, and maybe there's some interesting thing happening on the regulation side. Like, things are moving really, really quickly. I think the way I try to stay at least relatively sane, one is, like, actually carving time off.
And I think the Anthropic co-founders do a good job of, like, saying, like, look, like, burnout. If you burn out, like, you're kind of done. And I've seen it happen, unfortunately, to people I'm really close to, and then it takes a long time to recover from that.
So actually encouraging people, like, there's no job that is so important that you can't be offline for a couple of days. So I think that's, like, a big key, like, piece in there. So, like, let's strongly believe. And if it is, you're probably doing something wrong, and you talk to somebody who could be a mentor to figure out how you can unblock that.
And then I think the other one as well is, like, I love sports and, like, this is the notion of, like, you're never as good as, like, your best game, and you're never as bad as your worst game.
I think that's also really true. Like, I know, like, in AI, there's, like, the, you know, it's so over, we're so back thing. Like, that, like, if you internalize that that cycle is always going to be at play in some way, you realize, like, it's never that bad.
Like, Ben Horowitz's book is The Hard Thing About Hard Things has this chapter on, like, we're effed, it's over. And, like, that feeling as a startup that probably many of you have had at startups where you're like, oh, I can't believe this thing happened.
Like, we're never going to, like, recover from this. We definitely had it at Instagram a couple of times. And then you get through it, and like, it defines, like, defines the company when you can actually go through that.
And I try to remind myself and the team here, even within Anthropic, which is like, look, this is a fast-moving, but it's also a long game. And it's like, we're never, it's never just about today's model launch and reaction or this product launch or something else.
Like, you're playing and you're building, and you just have to trust that you're building, like, the team and culture that is going to get through those things and have that sense of perspective. Even if perspective is saying, like, look, three months ago we were in a similar position.
Maybe it's not a year, it's just a matter of months, but it's still, like, zooming out and not thinking things, not letting your internal sort of, like, sense of self and success be so driven by the day-to-day.
Yeah. Has anyone, any coach or mentor said something to you that you repeat to yourself that gets you through the tough times?
I think the biggest one was this, like, sense of, like, if you're feeling something, it's really often the case that other people on the team are feeling it too. So this is, like, advice I got from my coach around just being, like, just verbalizing emotions.
Like, even saying, like, hey, I'm feeling really stressed out about this, or yeah, I'm really sad that we are shutting down this labs initiative. I literally had this meeting a couple months ago where I was working really hard on something, and I kicked off the meeting like, I'll kick it off like, I'm really sad, like, and frustrated.
Like, I wish this thing had worked out. And I think that holds the space for other people to be like, yeah, I'm pissed off too, or like, I'm sad too. And, like, I think giving that advice around, like, not, yeah, I think if you can get yourself to be open and vulnerable, it often, like, lets other people verbalize that.
And then you can, from there, you can be like, great, what are we going to do about it? Like, you know, it's much easier to start from that place.
Yeah. We actually kicked off AIE with a session from Carol Robbins, who runs Touchy Feely at Stanford. And I can't think of a better way to end than encouraging people to talk about their feelings, manage their mental health, and keep shipping.
Yeah.
Thanks so much, Mike.
Thanks for having me.





