Intro & live demo0:00
Thank you for coming here to my talk to watch me talk about agents on the canvas. Um, the first thing I'm going to do, though, is—before I have to record my screen—the first thing I'm going to do is I'm going to ask my agent to do something on the canvas.
And what I'm going to do is say, "Hey, my colleague Spencer just emailed me a link to a Notion document for a really cool demo we could build with the tldraw desktop app. Can you, like, find that document and then can you build it on the desktop app?"
Thank you.
Okay, so that's going to build, and then we're going to come back to that later and hopefully it'll work. Um, hi everyone. My name is Max Drake. Thanks so much for coming. I work on agents on the canvas at tldraw.
tldraw app & SDK1:03
I'm a product engineer there. So first things first, am I qualified to be giving this talk? I like to think so. I've been doing, like, agents on the canvas stuff since before ChatGPT came out. I think it's really cool.
I think there's, like, so much UX stuff you can do with when you get LLMs, you have them working in space. And I think it's really interesting. I've been doing it for about as long as you can have been doing it.
More recently, I've been talking about this a lot. Here's some proof. And yeah, so I work at this company called tldraw. Can I get a quick show of hands? Has anybody ever heard of or used tldraw before? Yeah?
Okay. Awesome. So yeah, the thing that you've probably used, if you've used tldraw, is this appright here. So this is all tldraw. This is a free infinite canvas whiteboarding app. You know, we have selections and arrows and resizing and, you know, all the things that you need in a whiteboard.
tldraw is also the company that makes this app. It's based in London. It's where I work. But the last thing that tldraw is, which is, I think, in my opinion, the most important, is it's the infinite canvas SDKs that powers this app.
And so what that means is that, you know, this is kind of the tldraw, the SDKs, the engine that powers a lot of infinite canvas experiences. Because it turns out it's really hard to get that kind of stuffright.
And so if you ever want to build a Miro competitor or a slide designer, or if you're, like, Replit, Replit has their whole new agent canvas stuff built on top of tldraw. And so the reason we built tldraw in the first place was that we were running into this issue, or people were running into this issue, where they had this idea for this, like, really great killer canvas app.
And they went to go build it, and everybody would run into the same problem where they would run into they would have trouble making the actual canvas part of the app. And they would, you know, try to deal with resizing and selection and, you know, all the matrix math.
And the issue is that they wouldn't be able to build their actual app itself. They would get stuck on the canvas. And so we built tldraw to kind of be the engine that could power that could be the canvas so that they could focus on the actual app.
When LLMs came out, we, like a lot of other people, saw that this was going to be this weird new type of software. I don't know if anyone, you know, I'm sure a lot of you were building in 2022, and it was really exciting.
And a lot of people, it was the exact same thing. People had the idea, had an idea for this cool app that would, you know, involve LLMs on the canvas, having them manipulating things in space. But then they would try to build it and they'd get stuck.
There were no best practices. People didn't really know how to do it. And so at tldraw, we realized that we need to make it easy for people to build with agents, with LLMs on the canvas. And also, so tldraw, the SDK, as well as the app, has multiplayer built in with, like, live sync.
It's really nice. There's cursors. There's you can see your collaborator's cursors and selections and viewports. And I think all of the things that make just the canvas in general a really great place for interacting with and collaborating with your colleagues also make it a really great place for interacting and collaborating with agents.
Why canvas agents struggle4:07
And I hope I'm going to be able to show you guys some of that in the demos that come up. So before we talk about agents on the canvas, really quickly, I want to talk about agents not on the canvas.
I'm sure you guys have all used an app that looks like this, you know, Claude Code. And I'm going to really oversimplify here, but basically, part of the reason why these apps are so good and why they work is because they're, you know, the medium in which they're working, writing code, is essentially the medium in which they were trained.
You know, it's text in, text out. That's how they were trained. And when we work with them, we give them a prompt and they write code. It's text in, text out. You know, again, oversimplifying, but that's essentially how they work.
I don't know if you guys have ever, you know, tried to get your agents to do, like, UI stuff and try to get them to align something. Found that they could not do that whatsoever. Because it turns out agents are really, really bad at working in 2D space and understanding 2D space.
And it actually requires, like, a lot of engineering work to get them to do it. And that's kind of the project that we've been embarking on at tldraw recently. And so the first thing we had to do, this is an older project, but the first thing we had to do is get them to teach them, teach the agents, or at this point, not agents, LLMs, to understand the canvas and understand kind of what they're even looking at.
Teaching models space5:08
And so we had this project called Teach, where we taught. So I'm going to, I'm going to prompt this really quick. I'm going to say, "Hey, make the mouse blow out the candle."
Yeah. So that's going to take a second. This is, this is an older, older project. But basically, what we had to do is we had to kind of, like, teach the LLMs how to take the, like, the screenshot that we give it and the JSON and all of the other information about the canvas and have, oh.
Yeah, there. Okay, so yeah, that's some, that's some wind. It's, is it, sometimes it gives us smoke as well. Yeah, and we got a little smoke as well. So, so we basically had to take it, how to, like, and I want to be very clear, this is not, this is not like a special mouse shape.
These are just, like, you know, these are just shapes on the canvas. This is, and so the work behind this, it's a single shot prompt, but we basically, we tell the agent how to interpret both via screenshots and via the data, what is actually on the canvas, like what it's looking at, which is actually, you know, it's not a trivial problem.
And then also how to, we teach it how to actually act on the canvas and to understand how the actions that it produces will affect the canvas. So, you know, it got it, you know, it made the, it made the smoke, it made the, it made the wind, it got the positionsright, and it understood what it was doing.
So we, we got this. We kind of figured out how the, like, kind of we got, we taught it what the canvas is. But this was like a single shot, single prompt kind of thing. And so the next thing we built is the tldraw agent starter kit, which basically turns that and wraps in a harness that lets an agent work agentically on the canvas.
The code is also MIT licensed. You can find it on, you can find it on the website. So here's a little, here's a little cat. I'm going to make this a little bigger. But what I'm going to say is, "Hey, so somewhere else on the canvas, there are some friends for the cat.
Agent starter kit7:01
Can you please bring one of them over to the cat? Her favorite color is red."
And so I'm going to zoom out. I'm going to show you guys what's actually going on. So you can see the view of the agent. There's some, there's some potential friends over here. And, you know, if you read the, allright, so, and basically what's going on is that the agent has kind of, like, we've given it a prompt, and using the information it has about the canvas, it's going to kind of, like, make some goals for itself.
You can see there's some to-dos in the corner here. It's, it changed its view in order to see what was, you know, the other stuff that was on the canvas. The same way that if you ask a coding agent, you know, you ask it, you know, where, where do we define this thing in the codebase?
It can go, it can search, it can find it. So this is kind of, like, turning that single shot prompting experience into this kind of, like, agentic thing that you can have, you know, it can autonomously set goals and, and work towards them.
Fairies & multiplayer8:12
The next thing we did, we did this, this project called Fairies. And so we basically, we had this agentic experience, but we realized that, you know, tldraw and, you know, the canvas in general is so collaborative. It's so multiplayer.
And we wanted to basically, we wanted people to be able to work together with their agents. And we also wanted the agents to be able to work together. So this is, this is a fairy. There's also, if you guys want to scan this QR code, you can actually, this is multiplayer.
You can join if you want. It requires a Gmail signup, but you don't need to pay for tokens. This is, this is what the link is. So basically this is a, this is a fairy. This fairy's name is Joan.
They don't like being, they don't like being grabbed. You can, you can throw them around. You know, we added a lot of really important stuff. You can, you know, you can, you can change its hat. You can change the color.
And this seems silly, but it's actually really important. And I'll talk about this a little bit more later, but actually understanding when you get a high-level view of when you see your agents working on the canvas, it's important to know which one is which.
And so differentiating them is actually important, which is why, of course, we added the leg slider.
But so, you know, I can say, like, you know, I can, I can, I can say hey to it. And I can say, you know, something like draw a cat. And I can have it work. But the most important thing here is that fairies have friends,right?
And they can, here we go.
And they can, fairies can work together. And so we kind of designed this, like, multi-agent collaboration system that works on the canvas.
And I'm going to actually, I'm going to go to, I think one of my, yeah, I think so. My colleagues' agents are here working, making this, this really great scene. I'm going to bring mine over. Summon. And I'm going to give them a slightly different prompt.
So I'm going to select them all. And now I have a group chat of the agents,right? And I'm going to say, "Hey, I have a board meeting coming up in, like, 10 minutes, and I don't have any of my figures.
Can you draw up, like, a little memo for all of my financial data for fiscal year 2025? Thank you." Okay, so what's going to happen there basically is this kind of, like, creates this multi-agent, you know, coordination thing.
We have one of the fairies is writing, writing out a plan. You can see it. And again, the animations are kind of cute and funny, but it's actually really important. I don't have to read a chat or go through, you know, imagine if I have 10 agents working.
I don't have to read a chat in order to know what's actually going on. I can look at the state of the agents, and I can actually, you know, I can see what's happening. So we, we have a task here that's been defined.
It seems like, you know, the fairy is waiting for that to finish.
One is bored.
Yeah. So that one's, one's bored. That one's waiting. So this is the orchestrator fairy. What it's done is it's assigned the, it's assigned the task, and now it's waiting for the other ones to start and finish it. And it's going to get notified.
It's going to get prompted in order to review. It seems like for whatever reason my internet's not working, but thank, oh, never mind. So yeah, we have one, we have this one. So yeah, we have one fairy who made the, made the task, one fairy who's working on it.
And so this is this kind of, you know, multi-agent coordination system on the canvas. I have, you know, you can see my colleague has his agents over here. They're working as well. And so you can kind of collaborate with people and with agents in this environment.
Agent dependency graph11:41
And I don't know, I think that's really cool. The, so the next thing, so the problem with fairies is that they're kind of trapped in the canvas. And all the stuff you've seen before, this requires, if you want to build something like this, this requires, like, the, you do, like, opt in and have your entire harness be a, like, canvas harness.
And the downside of that is that it makes it really hard to have any of this work with stuff, like, outside in the real world. The fairies are, the fairies are trapped in the canvas. And so I built this experiment.
We had a little hackathon internally. But first, as a quick motivation for that, at tldraw, whenever we have an, whenever we're getting closer to a launch, we, like, abandon all of our task tracking software and we make just one massive dependency graph of how, like, so this is what an actual, this is a real thing from when we launched fairies, actually.
And so this is what it looks like when we're, like, really, like, when shit is hitting the fan at tldraw when we're launching something. And I really like this interface because it kind of lets you, this is not like a special app.
This is still just tldraw.com. You can, you know, move your shapes around and things like that. But I really like this because it both, it lets you see, like, what depends on what. It lets you know what's coming next.
It lets you get a high-level overview. You know, these are all green because we finished them, but, you know, you can imagine during the project, some of them are in process. And I really want, I really wanted something like this, but something that I could actually, that could actually do the work itself.
And so I prototyped this thing. It's called the Tech Tree app. And basically it's similar to this. It's a dependency graph. But each of these tasks is a coding agent that you can kick off and you can have your agent kind of, like, be running and doing them autonomously.
The project itself that's working on, it's this little, this is just kind of like a demo app. But this is, can I, yeah, so this is, this is a little fun, you know, multimodal input thing. I haven't written any of the code for this.
This is, this is all written by agents, but I can manage all of the work is being done in this desktop app or in this app here. And so I can do something like, I can see this one has finished building some gesture controls for the canvas.
So I can open the PR. And unfortunately, sorry, Jeffrey, I am just going to merge this. I am not going to have it be explained to me. But so this is, and so yeah, great. Awesome. It looks good.
And then, you know, eventually this is going to get marked as complete. And this is also multiplayer, which is really cool. And you can have people working together. Yeah, so that's finished. And you can also prompt from, like, inside the app.
You can draw and have a prompt. So I can basically, I can just take all of this and I can draw a little, like, so this is my prompt and I can wrap it in a task. And I can, you know, call it facial animation canvas control.
And then I can assign that to Claude and I can just hit run. And so now that's working as well. And so this is kind of like, you know, this is kind of something similar to Conductor or OpenAI's Symphony, where you're using a kind of, like, one abstract interface above what the actual, in order to, like, manage your multi-agent coordination and things like that.
And the thing I like about this also is that because this is multiplayer, one of my colleagues can come and join and add tasks and edit things and see the work that's been going on. So it's much more collaborative than, like, your own instance of something.
Here's the moment of truth. Let's see if that demo that I had it build in the beginning worked.
Desktop scripting15:12
Allright. It's, it hasn't built the fluid simulation yet. It's been working for 13 minutes.
That's actually fine. So basically this is the tldraw desktop app. Something that's really cool here is that we have, so this is running locally. It's working on files. We'll see if it finishes. We'll let this run. But basically what this does is this, this basically exposes the editor instance of the tldraw app that's running here.
And it has a server that lets any agent, for example, my Claude code, write just plain JavaScript against the, against the editor. And, and it basically, it, you know, it's code mode if you've ever used code mode. But you can basically turn your tldraw desktop app into a, like, scripting environment.
And the, one of my colleagues actually is, I'm going to, this is, this is the kind of off the rails bit of the canvas here or of the, of the talk. So here's something my colleague made using the same thing.
So this is, he has a tldraw desktop app in the corner here, and he's using it as his window manager. And what he did, the way he did this was he just told Claude code to, because Claude code has access to your actual computer, it's not locked into the canvas.
It basically, you know, it made some rectangles and it probably wrote some Apple script or something to actually re, you know, move the things around. And so you can kind of make all of these, like, ephemeral UIs and have them actually be doing things in the real world.
Another really cool one that he did was, if this loads, it's Pong on the desktop. Let's hit it with a little refresh there and see if it works. Yeah. So this is, he's got in the corner here, you know, you have, you have tldraw running.
This is the desktop app, and it's using the windows in order to play Pong. And so again, like, kind of crazy, but there's, you know, maybe it seems a little silly, but
let's see if this, this worked. Oh, it's still working, man. It was usually much faster. But I think this stuff is so cool because this lets you kind of, you know, do all of the weird kind of, like, spatial interfaces that you can do on the canvas.
You get all of, like, the primitives of the canvas. But you can, like, you can have your agents working kind of, like, in the real world. It has access to real data. If I scroll up, I'll show that
if, you know, this is my Claude code and it found, it got the Gmail, it got the Notion doc, it found the spec, and it's going to implement it.
But yeah, so to sum up, I think that agents working on the canvas is so cool. And I think that there's, like, so much we can do if we use, like, the agent, the canvas as a place to work with agents.
Wrap-up18:01
And I think the place part of it is really important because, you know, when we do, you know, with remote work collaboration, we do a lot of stuff online with each other and we collaborate with people on the canvas.
And I think that the, yeah, the canvas can be a place where we collaborate with agents. And I'm, I'm, I'm vamping because I'm trying to see if this is finished, but I don't think it's going to finish. But thank you so much.





