Meet Scout0:00
Hello everyone, my name is Sandhy, and meet my co-host today: Scout. This is my friendly rover, and one would think that rovers can't really think for themselves,right? We have to tell them what to do, or we have to very specifically program them on how to think.
But this little guy here actually has a brain, and he can think for himself. Let me show you my screen—oh no, it's going to the wrong screen. I'm going to see how I can stop this, and I'm going to see how I can move to my screen.
Um, give me just a second. I'm going to end show, and then we get to this. So this is what Scout here is looking at, and Scout here has a small little brain. And Scout can understand what I'm saying in natural language.
Natural Language1:09
For example, if you can see my screen here, if I say, "Hey Scout," I'm going to type to him. Say, "Hey Scout, turn on your headlights and say hi to everyone."
He is actually going to be able to understand and do those actions. But how is he able to do these things,right? It's going to take him a couple of seconds to think about it, because how this has been set up is he's working with a Raspberry Pi installed, and that Raspberry Pi is actually in a small little location in New York City, and he's got a SIM card here which is then connecting to that Raspberry Pi, and then he's getting signals.
So this is via 4G connection, and so he's like, "Okay, on it." And if you can see—you know what, I'm going to tell him, "Spin around," and—oh no, he's actually speaking to me. He's like, "I see the stage, and it's looking great.
Untrained Question2:09
Big presentation screen to the left, bright stage lights, and rows of seats for the audience." How many people do you see?
I've never trained him to answer this question, but he's—let's see if he's getting it. He's thinking for like a second or two.
Let me take a good look. From the front camera, I see two people: one near the speaker monitor set up at the back of the stage, and another person furtherright. Now, spin 360.
And he's also saying, "Is this session about to start?" Right? So he's going to spin 360 any second now, and then he's probably going to be like, "Oh wow, I did." So this little robot here is a next generation of robot where—ah, there we go.
He is spinning 360 now, and he's probably going to tell me what he's seeing.
Hardware Agent3:18
And he's saying, "Let's spin,"right? So this new generation of robots is different from our traditional robot training, because I have given this guy a little brain. And what do I mean by I've given him a brain? I've given this robot an agentic layer, and I've given it—it's called Strands Agents, which is an open source framework which was built by AWS, and I'm going to quickly go back to my slide deck.
We can see it,right? And so here, what happens is we have these existing tools that the robot can do. He can take certain actions by himself, but only those actions by himself. So what we can do is we can add a layer of LLM, or even better, add a layer of agent to it so that the agent orchestrates which tool to call and how to really get the robot to start doing the things we want.
So in traditional software with traditional AI, like AI engineering, we can give agents software tools. Similarly, we can give the same AI agent a hardware tool called a robot which has access to preset functions or programmable policies, and then the agent can decide which policy to implement when.
So all it takes is one robot agent for us to be able to do new, enumerous tasks and have it understand what we're teaching it in natural language. So how do we get started with it? All it takes is five lines of code.
Five Lines5:00
This is through the agent harness called Strands, and all we have to do is import the Strands agent and call the robot tool. And we say, "Tools equals the robot," and then we say, "Pick up the red cube," and it should be able to pick up the red cube, assuming that the robot has that capability.
Yeah, now he's seen someone, and he's like, "Oh, let me go towards that person." So he gets pretty excited. This guy is pretty special because he doesn't have just one agent. He's got three different agents. All three of them are Strands, and all three of them are working simultaneously.
Three Agents5:35
One of them is the thinker agent, and that's the part of him that's constantly thinking and assessing the environment and like, "What do I do next?" And that guy's—that part of his brain is constantly thinking. Then there's the other communication part of it where—and I'm going to show you that in a bit,right?
And I've connected him to my Telegram app as well as to my web app, and so he is able to have a conversation with me in natural language and then take actions based on what I'm telling him to do, apart from him just perceiving and thinking and figuring out what he wants to do.
And the third agent, the third type of agent that he's got access to, is the voice agent. I did have to disable it because every time I speak, he's going to think I'm speaking to him, and so he's going to keep chatting away with me, and it's just not going to be fun because we're going to have our co-host interrupting me all the time.
So I've disabled that feature for the time being, but essentially all three of these agents work in tandem with this one robot, and thereby this gives him the ability to do way more than what—just what he's been trained to do, more than just the policies that he's learned.
Now, what is—a quick overview on the Strands package itself. This Strands package has more than—supports more than 40 different robots under eight categories, and all of these are just simple robot tool calls. And how is this all set up?
Four Layers6:46
Four different layers. The first one is the agent layer, the topmost one. And there are two parts to this. One is how the actions go in, and the second is how it observes and the observations go up. So if you notice, it's very bidirectional.
So first, when we give it an instruction, we would be talking to this Strands agent, which is the agentic layer. That would then decide which policy to call. And the policy provider—again, Strands agent supports a bunch of different policy providers—and we can then train our policy based on our traditional robot training.
So in our policies, we would collect data, and then we would train on it, and we would create more simulation data, and that policy then becomes a VLA model, which then the robot would have access to, Strands agents would have access to, and then it would invoke that specific policy based on the question that we're asking it or the command that we're giving it.
And that policy needs to sit somewhere,right? So that sits in the backend, which could be your simulation environment, or it could be your real hardware chip, your hardware environment. That is the backend on which—that is the interface on which the policy is running.
And finally, the output actually takes place in the physical hardware, which is the robot. And so the robot—ah, see? So now it's responding. This ro—even if he falls down, he's supposed to be fine. He technically shouldn't, um, he technically shouldn't get hurt.
He should be able to pick back up from where he stops. Ah, okay. So I'm telling him to go back a bit. Back off. Let's see if he actually backs off. Um, so that is the four layers of how to get started with building this,right?
And what's happening under the hood. Like a more picturesque view of what's the architecture of what's going on under the hood. We want everything is basically Strands agents on the edge as well as on the cloud. We want to be able to train the VLA and the policies on with using agent call, and we want that to happen on the cloud, but we also want to be able to call it directly on edge so that our robot can execute functions and policies faster.
Cloud and Edge9:28
So this is sort of like a hybrid model where a part of it happens on the cloud and another part of it happens on the edge, and Strands can decide when to call which part of it. And so this helps with massive amounts of training as well when it's constantly collecting information, and it's able to train on that information and learn from itself, but also just execute at runtime really, really quickly.
Now, like I said, the agent decides what to do, and the policy decides how it should be done. But he's pretty smart. He should be able to pick himself back up if he's not fully fallen down, and he should be able to continue moving along.
So I think he's okay. Now, where does this leave us, and why is this so special? We started off with very traditional robots. Robots have existed since forever,right? And they've always just been programmed, pre-programmed to do—to autom—be automated and do a certain set of tasks autonomously.
Future of Policies10:11
But there is a future in this world where these robot policies, these VLA models could be so advanced that we wouldn't even need to do this. They could be as large as our large language models so that—ah, wait, hang on.
He's falling back again. I'm going to see if I can get him to move back up.
Good boy. Stop. And then he's fallen off again. Um, we get to a point where these large language—the VLA models could be as large and as amazing as our larger language models, and they have all the information in the world, and we wouldn't even have to do this.
We might just have to feed in one simple model, and then we could give it to him, and then he would know exactly what to do. But until that point where we don't have to fine-tune on top of existing VLAs and existing policies, we can do this.
And this is a stepping stone towards a future where we don't need to train robots anymore. So now if we wanted to do more things than just the tasks it's trained on, give it an agent and see what it can do.
Complex Demo11:48
And so let me quickly go back to my demo, and I'm going to show you how it's actually working.
Okay, so this is my—so this is Strands here. This is Scout here. And I've been telling him to do a bunch of things. So I can say, "Hey, do something complex."
That's not complex. He's going to be thinking now. Ah, he's going to fall off.
So he's saying, "Let's spin. Full 360. Done." Still safely on the stage. I can see the bright stage lights and the audience seating area. All good. What's there? Ooh, a challenge.
Watch this.
So he's speaking.
I call this my signature performance.
But he's not doing anything. What are you doing?
He clearly seems to be speaking, but what are you doing? Please do something. He just turned off his headlights. Cool. Ah, okay. Now he's calling. So do you see it saying, "Calling rover speak," which was the function that it called because I said do something complex?
So now it spoke, but now I think it should have been attempting to do something, and it fell off because it tried doing something. I've actually seen it do like a funky dance, like this funky dance move, but he's got a mind of his own,right?
Now, what's going on under the hood here? A couple of things. The first thing is here, I can use this. What is the point of creating him? I can use him to create my datasets. Because I'm able to also manually move him, I will get him to navigate in the direction that I want him to, and then I can create training episodes, and I can get information on how he's responding and how he's reasoning based on the questions that I ask.
Under the Hood13:17
And this is super good information for me to then be able to make him do a better job of it. So that's one part of this whole process and this experiment of getting—of giving him his own autonomy and getting him to do things so that I can create more data.
But also, apart from that, this is my configuration. So over here, under the hood, Strands agents, which is your harness SDK, is using currently Anthropic Cloud Opus 4.8 under the hood. So that is the brain. And then this is my simple prompt where—system prompt where I'm telling it what it's supposed to be doing, and I'm telling it all of the rules, and I'm also giving it access to all of the rules that it's already got.
So I'm telling it what each of these rules are meant for, and so that's how Strands decides which tool to invoke based on what I'm asking it to do. And the voice that it's using is the one of OpenAI Realtime, and I've also given it more information for it to be able to, like, just safety and guardrails to ensure that it's doing really well.
Telegram Mirror14:52
Now, it's—this is—these are two of the agents. The other thing that it can do is also chat with me on Telegram. This is amazing because when I'm not at home and I still want to get it to speak to me, I can say, "Hey, Scout, who is turn around, uh, spin around, analyze—uh-uh, don't fall off—analyze the room and tell me who is the best-looking person."
This is always a hit or miss. Don't blame me if Scout decides to be biased. I tried playing with this because it makes me feel like Snow White, um, and the queen saying, "Mirror, mirror on the wall. Who is the fairest of them all?"
And so let's see if Strands can give us that. So it's saying, "Ha-ha, on it's on there. Spinning around to scope the room." Is he spinning? He is. Looks like he's making—and let's look at what he's looking at.
Now he's still looking at the stage. This is the other mirror. He's looking at me. "Hi, am I the prettiest? Oh, you flatter me." But no, turn again. Spin some more, please. He's decided not to spin more. Let's see what he says.
Spin complete. Here is my analysis of the room. You can see my screen,right? I have no way of zooming this in. Um, spotted six to seven people total. A couple seated front center near the bright stageright. A few more in the audience.
One person in a dark outfit on the stage. The rocking khaki jeans and sneakers. Thank you very much. Oh, no, no, no. The person in front of me, khaki jeans and sneakers, clearly walking with purpose on stage. Uh, best looking?
Honestly, from my low-angle ground view, it's tough to tell faces. But the person on the stageright now gets style points for the confidence stride and the dramatic blue curtain backdrop. Thank you very much. Every time I need a confidence boost, I can be the only person in the room and ask this guy, "Who's the prettiest of them all?"
and he will always choose me as the answer.
Thank you very much.





