Intro0:00
Good morning. I know a lot of you in this room; it's great to see you. Welcome to the voice track at AI Engineer World's Fair. For those of you who don't know me, my name is Kwindla Hultman Kramer.
I work at a company called Daily. We make developer infrastructure for real-time audio, video, and AI, and we're the team behind Pipecat, which is the most widely used framework for building voice agents today. Pipecat is open source and vendor-neutral.
It's used by companies like AWS, and NVIDIA, and Anthropic, and thousands of startups, and scale-ups, and enterprises. And today I'm going to talk about what kind of agents we're building today, including voice agents, but not just voice agents, and what I'm interested in building next.
Bush's Vision0:58
And I'm going to try to put all this in the context of the roughly 80-year history of digital computing so far. So we've got a lot to cover. We're going to go fast. But we're going to start in 1945 with an essay called As We May Think, written by an engineer, an academic, a civil servant named Vannevar Bush.
Bush deeply understood technologies ranging from analog computers to photography to radio to radar, and As We May Think is an extraordinary piece of writing. The essay predicts the development of, among other things, document display on a screen, and document scanning in OCR, and speech to text, and text to speech, and programming languages, and hypertext, and search engines, and data networks.
Something like the GoPro camera, something weirdly like the Amazon Kindle Store, and voice interfaces, and brain-computer interfaces. And I've been thinking a lot about As We May Think lately because Bush wrote this essayright at the very beginning of the computing age.
And I think it feels to most of us like we're workingright at the beginning of a new age: the intelligence age. So what will we build? Well, at the moment we're building agents, and we're having a lot of fun doing it.
Agents & Web2:19
And a lot of the AI engineering work we're all talking about this week is focused on building a full, coherent software stack for AI agents. Here is Satya Nadella talking a couple weeks ago on a crossover episode of the No Priors in Latent Space pod about the challenges of building agents in 2026.
That'sright. So in some sense, you kind of want the harness to define the models, the data, and the tools, and so that you have a loop across those three. And so what we're trying to, first of all, make sure is each of our products that we build,right, whether it's GitHub Copilot or the security copilot, the stuff we showed with MDash, or even the Discovery for Science, it doesn't matter.
All of them are multimodal harnesses with tools access so that you can do this progressive disclosure of tools even, so that they're token efficient. And then you're feeding it with very rich context.
So if you were here last year at AI Engineer World's Fair, you could draw a through line from the things we were talking about last year to loops, and tool calls, and context engineering, and the stuff we're focused on this year, to some emerging ideas.
You can hear that in Nadella's clip just there. I think of this as kind of agents plus plus, like multimodal harnesses and software copilots embedded in every single piece of software, organization-level harnesses. So how do we go from agents to agents plus plus to the next thing beyond agents?
Well, the last time we had this kind of massive change in how we write software and what we write software for and to do was the early days of the World Wide Web. And I was around for the early days of the World Wide Web.
I was a baby programmer in 1995. And the thing we talked about all the time in 1995, the way we talk about agents today, is web pages. I spent a lot of time writing HTML by hand and building web server software in C, and indexing and search software in C, and authoring tooling and management infrastructure for web pages in Perl.
I was as excited about HTML in 1995 as I am about agents today. And the web page is still with us, and it's still important and useful. But today we talk a lot more about web applications and native mobile applications than we talk about web pages.
So just like we went from web pages to full-blown web and native mobile, clearly we're going to chart a path to a new, fully AI-native software that comes after agents and agents plus plus. So let's keep going back in time in order to think about this future.
Here's a timeline Vannevar Bush lays out in As We May Think. He talks about the abacus, which was both an immensely useful device for doing practical, everyday mathematical calculations, and also an incredibly important theoretical tool that led to ideas like numeric place value and the concept of zero.
Computing History5:03
And Bush talks about the massive jump from the abacus to the state-of-the-art electromechanical keyboard calculating machines that he had in 1945. And then he posits that we're about, or he is about, to
witness and help create an equally large leap to what he calls the arithmetical machine. And then he goes a step even further than that, and he invents or designs in that essay a device he calls the MIMX. And we have a little bit of an advantage over Bush in 1945.
We've seen 80 years of computing play out. So we can modify his timeline a little bit. We can go from the abacus to the stored program computer to 40 years later the personal computer, and 40 years after that this AI agents era that we're all collectively helping to invent and create and bring into being.
So the question for me is, what did we build to go from those very first digital computers in the 1940s to the personal computer in the 1980s? Well, in the 1950s, the big job was to figure out more effective ways of transmitting human intent to these new computing machines.
We built the first programming languages. We wrote the first compilers. And the theoretical underpinnings here were figuring out how to combine the elegance of mathematical formalisms with something a little bit more like natural language. And then building on that, in the 1960s, the challenge was to make these machines interactive, make these machines capable of a two-way dialogue with humans.
The '60s also saw the birth of graphical programming with systems like Ivan Sutherland's Sketchpad. And the '60s were an amazing era for science fiction. Even though almost nobody had access to a computer, the computer became a big part of the popular imagination.
The idea of a computer really resonated with people. And ideas matter. For example, here is the idea of the computer in Star Trek.
On record.
Recording.
Come.
Captain's log supplemental. Engineering officer Scott informs warp engines damaged but can be made operational and re-energized.
Computed and recorded, dear.
Computer, you will not address me in that manner. Compute.
Computed, dear.
I love the background sound of punch cards going through a punch card reader, so like you know the computer is working even though it's talking to you about what it's actually computing. There were, of course, a bunch of other talking computers in science fiction of the '60s.
And next year after this, the Kubrick movie that was an interpretation of Arthur C. Clarke's 2001: A Space Odyssey had the HAL 9000 computer. This is a much, much more dystopian view of a talking computer than the Star Trek computers.
And by the 1970s, computers had become powerful enough that the next big job was designing abstractions that could scale to much larger amounts of data and much more powerful computing substrates. We got relational databases, which introduced new theoretical underpinnings for data manipulation.
And we got declarative languages, which leveraged those new theoretical insights. And programming languages in general continued to evolve in what, to me at least, are really amazing ways. We got small talk and object-oriented programming in the '70s. And all of this set the stage for the personal computer in the 1980s.
PC Revolution9:03
The Macintosh shipped in 1984. Windows 1.0 shipped in 1985. And Microsoft's mission statement was a computer on every desk and in every home. And incredibly, Microsoft delivered on that mission statement. And we got a computer on every desk and in every home because these new personal computers delivered real, amazing, tangible benefits.
Take VisiCalc, for example, which was the first spreadsheet program, a truly new abstraction for doing computation, numerical computing, two-dimensional, interactive, so durable and so useful that probably most of us in this room use a direct descendant of VisiCalc regularly, Google Sheets, or Microsoft Excel, or whatever.
Or put another way, this was a spreadsheet in 1957, and this was a spreadsheet in 1985. And I think a lot about VisiCalc these days too because I think VisiCalc is an example of how transformative new technologies can be in the way of delivering a capability that used to require a lot of specialized people and specialized knowledge and making it generally accessible.
And I think VisiCalc is potentially a counterargument to the argument, or the fear, or the concern that AI is going to lead to mass unemployment because VisiCalc didn't put accountants out of business. Instead, it made much, much, much, much more accounting-like work possible.
And it made new categories of work possible that we couldn't even really conceive of when a spreadsheet or doing a screen's worth of calculations, as we think about it today, took a room full of people. So if we were here in the Moscone Center in 1985, and these two interfaces, the Macintosh system 2 and Windows 1.0, were state-of-the-art, what would we have said the world would look like in 10, or 20, or 30, or 40 years?
Well, we actually have a really great example of a prediction from that time. Like As We May Think, another famous document in the history of human-computer interaction, concept video from Apple made in 1987 called Knowledge Navigator. This is very much worth tracking down online and watching all of if you haven't seen it.
I'm just going to play about 20 seconds from the middle.
You have three messages. Your graduate research team in Guatemala just checking in. Robert Jordan, a second semester junior, requesting a second extension on his term paper. And your mother reminding you about your father's surprise birthday party next Sunday.
So the video shows a foldable tablet, a touchscreen interface, a conversational voice assistant with a really strong personality, access to both global and personal information, real-time video generation, real-time computer vision, seamless video call integration, delegation of complex tasks for autonomous execution, and what we might call today continual learning.
And it's really, really clearly influenced by As We May Think, but it's also quite different. It really is updated for 40 years of progress, and it really does sort of pre-sage this AI agents era we're in now in a way that Vannevar Bush's MIMX didn't and maybe couldn't.
Networked Future12:17
The Knowledge Navigator video divides our timeline, I think, quite neatly in half. And hold that thought because we're going to come back to it. The 1990s were about the network, first local area networks, and dial-up, and then the internet and the web.
And with the benefit of hindsight, I now think that the single most important thing about the web was that it was multimodal from the very beginning, more even than the GUIs of the 1980s. The web anticipated that text, and audio, and video, and data were not different things to be used in different programs.
They belonged together. And in a real sense, the web was an attempt, and a conscious attempt on the part of a lot of people building the web to make that Knowledge Navigator video real. Then in the first decade of the new millennium, the big job was to make all of this computing stuff mobile and continually connected, to put this new multimodal networked computer in your pocket, literally to give a supercomputer to everybody in the world that they could carry around in their hand.
And as with the 1960s, there was an efflorescence of futurism on screen in the first few years of the new millennium. And I think it was because computers you could carry around with you, and cameras everywhere, and a kind of Moore's law for pixels making screens super cheap really gave us a chance to think through what we thought the future would look like in a new way.
A lot of stuff we could almost but not quite build was cohering in the minds of people working on these machines. And the best and most famous Hollywood computers from that era were created by John Undercoughler for the films Minority Report and Iron Man.
Here's Minority Report from 2002.
It's no longer there.
Time frame?
13 minutes.
Hey, Chief. Investigator from the Feds here.
Yeah, I don't need some tweak from the Fed poking aroundright now.
John, the word's out on your calendar. I left you a message at your house.
Check in with the papers they had it forwarded. See if the neighbors knew where they went. Check all relations.
Checking neighbors and relations.
But John.
What?
Just get him some coffee. Tell him some stories how I save your ass every day and you can't live without me.
I got coffee. Thank you.
Danny Whitworth.
Tweak from the Fed.
Oops.
Come. So the gestural interface in Minority Report was implemented on screen as special effects, but it was actually based on John's PhD work at the MIT Media Lab. And in a real sense, this was real technology. John had brought the UI out of the small screen and into the world with us in a bunch of really interesting and lovely ways.
John also consulted on Iron Man, which is a very different view of the future than Minority Report, which was Spielberg working in like the American Kubrick dystopian tradition. Iron Man is really squarely in that Star Trek goofy futurist tradition.
But I think you could see the common elements in the UI depicted on screen. It's still from the same era.
Wake up, daddy's home.
Welcome home, sir. Congratulations on the opening ceremonies. They were such a success. As was your Senate hearing. And may I say how refreshing it is to finally see you in a video with your clothing on, sir.
Sorry, got too late.
You.
Sure.
Swear to God, I'll dismantle you, I'll soak your motherboard, I'll turn you into a wine rack.
If you say that you.
I co-founded a startup with John in 2006 to make the Minority Report interface into a commercial product. This is our demo reel from 2012, six years into that work.
Agents & Beyond17:11
This long project to build the multimodal, multi-device, multi-screen, multi-player, ubiquitously connected computer is still what I'm working on 15 years later. In 2010s, we built out the Cloud, which laid the groundwork for the infrastructure and data centers and data capacity we would need to scale up AI training and inference, which brings us to now.
We're building agents. And we're starting to think about Agents++. But I think we can also start to think about the next thing, the AI native software that is to agents what today's internet is to the web pages of 1995.
And one way to think about the story is this. We went from the calculator to the computer to the personal computer to the global cloud computer. And now we actually have the ability to build the MIMX, and Jarvis from Iron Man, and Knowledge Navigator from that 1987 video for real, completely working.
And by building those things, we'll figure out what we want to build next. A couple of weeks ago, the team at Tavas released a reimagined Knowledge Navigator video, this time entirely built on real and available technology. I'll just play another 20 seconds of this.
But like the original Knowledge Navigator video, this is worth tracking down and watching in full.
Good evening, Hasan. I adore that houndstooth jacket you're wearing today. Anything I can help with, or would you like to review tomorrow's schedule?
Thanks so much, Tom. Yeah, let's review tomorrow's schedule and see how busy it is.
Opening your calendar now. Here is the quick version since it is late. Tomorrow morning is slammed. Investor meeting at 9:00, then back-to-back one-on-ones and internal meetings until 1:00.
So the full four-minute video is one take, completely real. And when you watch it, it really does feel both like the Knowledge Navigator video from 1987, familiar but built on real technology, and like something brand new. And I'll just close with a massively multiplayer game project I've been working on with some friends as a canvas to really think about what AI native software can be.
Gradient Bang18:56
This game is built from the ground up with LLMs as the core of every interaction. At every moment in the game, there are hundreds of inference calls happening, and we couldn't have built anything like this even a year ago.
Oh, sorry.
Welcome to Gradient Bang, a multiplayer game that showcases real-time agent orchestration. Gradient Bang demonstrates several patterns for AI sub-agents, such as asynchronous non-blocking context compression.
OK, make a note for later. We are going to eliminate Haley Tightly from existence.
Noted.
Long-running sub-agents that share context.
Eagle is on five trade loops. Hawk and Raptor are on five exploration loops each. Your fleet is busy.
Progressive skills loading.
How much does your average ship cost?
Ships range quite a bit, Captain.
Dynamic user interface generation.
Show my task history.
Certainly.
Hide the map.
OK.
And conversational voice.
No, I don't want to exchange it. I just want to sell it for cold, hard cash, please.
I'm afraid the galaxy doesn't allow you to be shipless and hitchhike.
So I went long. I have to wrap up, but I will say that the first version of this new draft talk was an hour, so I have a lot more things I'm super excited to talk about with all of you.
Closing20:32
So if you are interested in this stuff, come find me. We have a booth on the show floor. I'm online everywhere. And I'm excited to build agents because agents are awesome, but also to build the next next thing too.
Thank you.





