AIAI EngineerAug 7, 2026· 43:21

Local Models: Trust, Control, Optimization — Carter Abdallah, NVIDIA

NVIDIA's Carter Abdallah, Prime Intellect's Vincent Weisser, Arcee's Lucas Atkins, and NVIDIA's Chris Alexiuk argue open-weight models are the trustworthy foundation for enterprise and local AI. Atkins separates trust from safety: when Anthropic pulled Fable, enterprises chose Chinese open models for guaranteed availability, and open models are inspectable unlike closed APIs. Arcee pretrained a 400B model in six months; Weisser cites a customer that specialized an open model for finance in a week or two, beating Opus at a fraction of Haiku's cost. Alexiuk calls open weights the fix for 'mismanaged genius' and expects capable local models on MacBooks within a year; the panel predicts Fable-level open models within a year and hopes local-model use rises from a rounding error to 10–15%.

  1. 0:00Intros
  2. 6:11Trust
  3. 11:41Control
  4. 22:36Optimization
  5. 32:16Predictions
  6. 41:25Outro

Powered by PodHood

Transcript

Intros0:00

Carter Abdallah0:13

Thanks all. I hope everybody had a great lunch, and you got to check out some of the amazing demos that we have. We're going to begin the panel, the first panel of the afternoon here, where we're going to be talking about, of course, the engines that are actually powering the stuff that, you know, could remotely be used for things like local, sovereign, any kind of ownership over your own artificial intelligence, and of course the engine powering those, in addition to the hardware is the models themself.

And so, for this panel, we have excellent guests. We have Vincent, who's the CEO and founder of Prime Intellect. We've got Lucas, the CTO of Arcee AI. And we've got Chris, who is the Senior Product Research Engineer on the Nemotron family of models at NVIDIA.

Now, what's really cool about working in this industry is, uh, really cool companies like this, we all get to work together. And so this is one panel where we all directly get to work together on both models, infrastructure, some of the ways that we think that the direction of the industry should go, and each of us kind of play a different role in that stack.

But I want to leave it to you guys to introduce yourself and be able to talk about, sort of, the charter that you see, the problem of the stack that you guys are working on.

Vincent Weisser1:21

Awesome. Should I kick it off?

Lucas1:22

Kick it off.

Vincent Weisser1:23

Yeah. So I'm Vincent, as you mentioned. And then, really the goal with Prime Intellect from the beginning was, like, to ensure that basically frontier intelligence will be open and accessible, not just the models, but also the full stack to train the models.

So, kind of like this was, like, our motivation from the beginning. And we've, yeah, like, worked also together with a lot of gentlemen here on stage, like, on the one side, it's like, we work with folks like Lucas and Arcee to help them train frontier open models.

We help also, like, NVIDIA on the Nemotron coalition help out their frontier open models. And I think, like, I'm actually think both, like, Nemotron and Trinity might be the best, like, two open modelsright now outside of China. So I think it's actually, like, we need to fact-check that, you know?

But this is actually from my, uh, I think they might be.

Lucas2:13

Our marketing says that, yeah.

Vincent Weisser2:16

But, yeah, so that's the high level.

Lucas2:19

My name's Lucas Atkins. I'm happy to be here, and thank you for joining. Very similar to Vincent, Arcee was, you know, founded with the idea of

domain-specific, owned models are going to be needed. You know, we were founded early 2023, jumping on the custom model train quite early. You know, you have all these people who are excited about AI and all the things these new generation of LLMs can do, but they're using these monolithic, very expensive closed APIs for, at the time, and still, like, very narrow tasks that don't require, you know, at the time, it was $100 per million tokens out.

And through doing that, we were building on top of open models, and we were releasing a lot of our tooling in the open. And we noticed that in the United States, and in the West, you know, in general, we were starting to lose leadership in the open model space.

A lot of it was coming out of China, and that's amazing. I love those models. We learn a lot from them. We're close with a lot of the people building those. But when you're working with large enterprises and companies and geopolitics gets involved, whether you like it or not, you have people that become concerned about where those models are coming from.

And we decided that, you know, we had a good group of people, and we had a good group of partners like NVIDIA and like Prime Intellect, where we could probably try to pretrain ourselves. So last year, we did that.

We kind of reoriented the entire company towards, let's figure out how to pretrain a, you know, a 400 billion parameter model in six months. A lot of people said it was impossible, and in many ways it was, but we figured it out.

And now we are an open model lab, working with our wonderful partners and our customers to build Western open models that are permissive, and you can own those and customize them or run them wherever you want. And that's kind of where we're atright now.

So thanks for having me.

Chris Alexiuk4:19

Yeah. And so I'm Chris Alexiuk. I work at NVIDIA as a product research engineer, and I support the Nemotron family of models. I think it's, you know, we've talked a lot about why we do Nemotron, but just to say it a few more times, you know, AI should be open.

Open as in weights, data, training methodology, training frameworks. Really respect a lot of the work that the two other peops up here do, because they believe that very strongly as well. But the Nemotron family of models is focused on being as open as humanly possible.

So we have this understanding or belief that in order for AI to continue to grow and be useful to everybody, it has to be done in the open so that we can build off of each other, we can compound on each other.

And part of what we do, because Team Green, this is always true, is we think that the rate that you can squeeze tokens out of models is very important. So we kind of have this mantra that, like, faster models are smarter models.

And so a lot of the decisions we make when designing a model like Nemotron is built around how fast can we make it go, as especially you are going to see in the next however many months, local AI takeoff.

We need to make sure that models are well supported on hardware that doesn't just exist in massive buildings, you know, thousands of kilometers away from you. And so that's, you know, for AI to be very useful, it should be quick and open.

So that's kind of the vibe of Nemotron.

Lucas6:00

Who makes those buildings with the massive question?

Chris Alexiuk6:03

Oh, that's a lot of excellent people in the world that use a lot of excellent hardware from a pretty cool company. Yeah, I heard. I heard anyway.

Trust6:11

Carter Abdallah6:11

Yeah. And I'm Carter Abdallah. I'll be your moderator for today. Something that, you know, we all kind of talked about is this, you know, building on top of each other, learning from others, whether it is people, you know, across the big pond of the Pacific Ocean from us.

But really it is kind of like a collaborative sort of research effort. And I imagine that a lot of the people here in this room share that sentiment. But as it was brought up during the, you know, inaugural panel this morning in the State of the Union, there is a growing sentiment, potentially on the other side, that paints open source to be something that is actually more chaotic, that there's less trust involved.

And I think trust ultimately, as Lucas, you and I were talking about before, depending on who the party is and depending on what lens you're looking at it from, I think it kind of means different things. But ultimately from the end consumer, the somebody who's using this intelligence, or somebody who's, you know, more of a business and is actually customizing something to maybe monetize tokens in their business.

Can you comment, we'll start with you, Lucas, a bit on how open source and open source models are actually key to building that trust, so that when these people walk out of this room and somebody does come at them with that other angle, they can sort of steel man this side.

Lucas7:23

Certainly. You can weaponize any term, and certainly trust has been weaponized, that word. And the reason I say that is because it means something based on the context and with your speaking about it. Often, in AI, people like to conflate trust with safety, and those are not the same thing.

And I'm happy to speak on safety, you know, later on. But when it comes to trust,

I think that, you know, you hear a lot from closed model providers or politicians or people out in the space who are advocates for or against open source, that you can't trust these open models because you don't know what went into them.

Well, the same is true for these closed models, even more so. The benefit of open models is that we can very easily validate what is inside of them. There is a whole bunch of files with a whole bunch of matrices in there, and you can view them, and you can see the code that is running these models.

You have implementations from Prime RL, VLLM, SGLang, the provider themselves. These models are inherently trustworthy. You know much more about what's going on when you hit and talk to these models than you ever will what's going on when you hit an arbitrary API.

Now, that being said, certainly there is fear that people can reduce, you know, you can't trust that these models are writing safe code. Well, again, that is the same thing with any model. You need to use your judgment, and you need to make sure that you have the proper, you know, safeguards in place, and you're viewing the outputs of these models as the outputs of an inherently random system that we are working very, very hard to make less and less random.

I think that a telling thing is a lot of people said, well, you can't trust Chinese models, you can't trust Chinese models, you can't trust Chinese models. That was often, for the last few years, meant you can't trust open models.

Well, as soon as Anthropic had to put Fable away, and people realized that, oh, our access to these frontier systems might not be universal anymore. There's probably going to be a lot of checks and balances. You had a tremendous number of enterprises and developers and companies start going to these new Chinese models because they could trust that they would always have access to them.

And so when it comes down to trust, and the way I view that word as it relates to open models is, do I know that what I am running, and can I be as sure as possible that when I send something to this model, that I am going to get the output that I expect?

And the only way, currently, to be 100% sure that what you are getting is what you are expecting is by hitting an open model, either that you are running yourselves or you're working with a partner like Prime Intellect or Arcee or NVIDIA to validate.

So that's my take on the word trust.

Chris Alexiuk10:17

I think, too, something you mentioned is, like, we don't get to know a lot about the data that goes into these models. And that's something that I'm really happy, you know, that we're trying to do, which is not something, it's not, you know, the incentives don't exist for everyone to do this,right?

So it's not something that I think is mandatory or should be mandatory, thanks to the things that Lucas mentioned, which is that it's rather straightforward to validate what data did go into a model without seeing the data sources originally.

But I'm happy that NVIDIA continues to release data sets along with our models, release environments along with our models, to make sure that even if you can't go through the work of determining what went into the model, which you can do with the weights alone for the most part, you have, like, a spreadsheet you can look at that says, here's, you know, a couple trillion tokens of this data set, a couple trillion tokens here.

And I think that helps to educate people on why it's much easier to trust open models than models that we don't get access to any of that.

Lucas11:20

It also helps people see what that data looks like.

Chris Alexiuk11:22

Yeah.

Lucas11:23

You know, if you don't have someone releasing it openly, when someone says data is going in, I mean, data can take many different shapes. You can, but you can go to Hugging Face, you can go to NVIDIA, or you can go to Prime Intellect or Arcee's Hugging Face, you can look under our data sets, and you can see exactly what that looks like, and that can help you inform your priors on it.

Carter Abdallah11:41

Yeah, I think that trust also, you know, there's some angle of a reputation. Do I believe that your intentions are pure? And I think that a lot of people, again, in this room, believe that intelligence is kind of this next layer of almost, you know, infrastructure for us to progress as a species.

Control11:41

Carter Abdallah11:56

And I believe that everybody should have intelligence. So on that front, I want to hand it over to Vincent, because you kind of have this almost like founding thesis that this stack should be the open,right? The open superintelligence stack.

You want everybody to have a lot more intelligence. Can you talk about how this is kind of moving into the era of control, but beyond just data sets, how important it is to have the knobs and dials of this industry also be available in an open way for people who are building this?

Vincent Weisser12:29

Yeah. Like, I think it's like a really important point is to, to some extent, it's like be able to take those open models, like customize them, be able to, like, build on top of them. And I think, like, all the different components that go into it, like, especially from, like, the pretraining to mid to post training, I think, like, need to be more accessible,right?

Like, so more people can also, like, take those amazing models and, like, make them work for their specific use cases. So I think when we started, like, we also took a look at the whole stack that was out there and tried to figure out, like, what is missing for ourselves to train open models and for, like, helping our partners to do so.

And a lot of this was around the RL and post-training stack. So we basically went deep into building out, like, a lot of infra around that, like, around our environments, evals, around, like, making it much more accessible to do post-training also, because it's, like, the most economically viable way to maybe, like, customize those models, to take an open model and to have, like, a specific eval and environment and specific domain and dimension that you want to improve it on.

And this is kind of like what we are really doing with Prime Intellect now is, like, enabling people to post-train specialized agentic models. So being able to take models like Trinity, for example, from Arcee or Nemotron or others and specialize them, post-train them for the use cases that ultimately enterprises care about.

So a good example of this was, like, a company like, for example, RAMP was able to, like, take an open model and, like, specialize it to automate finance within, like, a week or two to get, like, better performance than, like, Opus at a fraction of the cost of Haiku.

And I think really this parade of frontier of, like, being able to create these specialized models that are much better than the frontier, but also faster, cheaper, I think it's like a key thing enterprises care about increasingly is really, like, just making it work for their use cases, basically.

Lucas14:09

If you go back to trust, it's how you can make your CFO trust you by knowing exactly how much something's going to cost all the time. That is increasingly becoming very important is you hear a lot, you know, all these companies have unbelievably large token spend and they're having to cut back on their Opus usage because they burned through it all in a couple months.

And that is going to continue to be a problem because, yes, the cost of an individual token has come down drastically. You can look at it, you know, the difference between GPT-4 when it first launched and GPT-5.5 is much, much cheaper per token, but at the same time, the amount of tokens in an individual session has gone up exponentially as well.

So we're kind of, we're spending more

on a total session. And so the ability to bring in-house or at least work with partners to ensure that you are controlling your cost and you're not at the whims of when a company releases a newer model that might be better, but also more expensive, they might deprecate a model.

Owning that and being sure that, you know, same way is what input goes in, you know what output's going to come out. In the same way, when an input goes in, how much it's going to cost, having assurance on that's really important too.

Vincent Weisser15:24

Yeah, and maybe like one thing to add to this is like almost like, I like this new term of, like, instead of speaking about token maxing, you know, speaking more about, like, the outcome maxing of, like, ultimately it's like you want to have, like, more than a dollar worth of value come out of, like, a dollar of input.

And I think this is sort of like Jevons paradox of, like, if you can create more value for, like, your GPU, basically, I think this is sort of like how you'll get, like, the most adoption also of, like, agentic models.

Like, if they can, like, be able to create as much value as possible. And I think the cheaper those models get, the more usage they'll get, like, for those specific use cases.

Lucas15:58

It's funny you say that. I have a, and I think a lot of us in this room, but especially on this panel, believe this to be so that the most meaningful AI applications in the next couple years, even this year, are the ones where the harness and the model and the product, they all kind of blend together.

If you think back to, at least for me, the first, like, truly game-changing agentic experience I had was when Deep Research from OpenAI. And that was because they spent a tremendous amount of time doing reinforcement learning on O3 with test time compute to do these longer running research tasks that people had tried previously, but they were kind of just doing a for loop over search.

Whereas I kept coming back to Deep Research. And, you know, you saw for a very long time that OpenAI and Anthropic and Google, when they'd release a new product, they'd release a custom version of their model for that product.

And if they're doing that, if their off-the-shelf GPT-5 isn't good enough for, you know, their Atlas web browser, why should it be good enough for our apps? And that's why I appreciate the work that, you know, Vincent and NVIDIA are doing for giving people the tools to customize their own model and allowing us to focus on how we get a good model to start from.

So it really is, you know, it's extremely important as you look at developing applications and experiences over these next few years that you're taking into account that you can make the model do something that maybe your harness isn't fully able to do alone.

Chris Alexiuk17:25

Well, that's something I think that's really important to just, like, reiterate,right? I mean, like, Nemotron is great, I love it. Trinity is great, I love it. Like, we design a model that's supposed to be as good as it can be across a number of harnesses,right?

You can see this in the technical report. The idea is, like, we want the model to work as well as it can in Py compared to, you know, Hermes compared to whatever you're using,right? But, like, you're not using all of these tools at once.

You're using one of these tools. And so when you have open models, you can do things like Noos Research can create a post-train of whatever model for their harness,right, that you know will be extra good. And, you know, this thing from, you know, I can't remember who originally wrote it, but this idea of, like, the mismanaged genius,right, where we're leaving a lot of, like, a lot of important capability on the table because we're just not, we're not fitting the models into the harness,right?

You can do a bunch of stuff with closed models, like you can change your prompts and your skills and all kinds of other neato things,right, but nothing will let you get the level of customization or custom ability that you can achieve with open models.

And I think that is something that is going to become increasingly and increasingly more important, especially thanks to folks like the others on the panel where, you know, I can just straight drop, like, my favorite coding and, you know, agent environment, spin up the CLI and suddenly my model feels way better with very little effort,right?

Like, that is something that is already at our fingertips and it is only going to get easier and easier as time goes on.

Vincent Weisser19:11

Yeah. And it's maybe also the most concrete, like, info to builders and audience, like, call to action of, like, if you kind of want to build the next, like, Claude Code, the next, like, cursor of Perplexity, I think the easiest way to get started is, like, take the best open model, like, and then post-train it on your harness, like, that you care about,right?

Like, basically, like, create a product that is, like, truly AI native,right? And I think this is sort of, like, I think one of the most exciting, like, unlocks for builders, like, out here.

Lucas19:37

That's a big thing about the framing of control is similar to, like, you know, when the cloud explosion started to happen in the, you know, the late 2000s, early 2010s, and the social media world kind of took off and apps became extremely popular and more, you know, cloud-based and, you know, managed by these bigger companies, you know, the data that they were collecting from you, whether anonymized or not, was how they were monetizing their platform through ads or in other ways.

And a very similar thing has always been happening, but I think it's becoming clear to people in the space is that the data that people get from you using these models is how these companies largely make their models better, whether it's through actually training on that data or by using it as a signal for what data they go out and find or generate to train.

And with closed models, there are terms of service that keep you from being able to, well, I could get into a discussion about what terms of service is an agreement between you and the provider. It's not illegal. Anyway, but you shouldn't be training on a, you know, a Claude Opus output or a Fable output or a GPT-5 output.

And they do a lot to try to obfuscate to make that not great for you. If you're using an open model, you can save all of those traces. All of those traces of you using it inside of your harness that will allow you over time to, if you say, hey, I want to go train a custom model, you can take all of that and again, either use it to directly do, like, fine-tuning on a smaller model so you're not spending as much, or to have a model help you find signals so that you can go out and use verifiers or Nemo RL or Nemo Gem to create these environments so that you can hill climb and make your models better.

So as much as using open models is like owning your stack, owning your intelligence, it's also owning your outputs,right? Owning your data. That's going to be extremely important too.

Chris Alexiuk21:31

I do want to, I do want to plug the license for a second.

So AI is very different than traditional software, which is why recently Nemotron, as well as Trinity, I know, has adopted the open MDW model data weights license. The idea is, like, we need a way to really make it very clear in the license that you can use the outputs to produce a model.

You can use the outputs to train, write all of these TOSs and stuff like that that have language that is meant to dissuade you to do that. We wanted to make sure there's a license that exists that not encourages you, but makes it crystal clear that it is permitted, it is permissible.

And I think, you know, the licenses maturing,right, to fit the use case better should be extremely positive signal for the way that the ecosystem is thinking about open models to the fact where even the lawyers are on board.

Lucas22:33

Do you know how much lawyers cost?

Chris Alexiuk22:36

Yeah. Yeah, it's as we move on to the, you know, kind of the third topic, which here is, of course, optimization. Something that, you know, we've implicitly said, but haven't said it quite explicitly yet is effectively that I think that for a lot of people there's this preconceived notion that when you're deciding to use an open model for whatever the use case, there are the trade-offs that come in the form of performance at the benefit of getting things like, you know, maybe data sovereignty and so forth.

Optimization22:36

Chris Alexiuk23:04

But what we are now talking about is that with theright customization and optimization, depending on the use case that you're and the harness that you're applying it to, you can actually exceed and build the model against the tool to get better performance than even frontier models.

I'd love to hear a little bit more about the, because I think that another thing that we would probably agree on is that the current level of intelligence already has so much left to diffuse into society. And so where are those areas where that diffusion is happening in the specific industries?

I know, for example, things around, again, kind of fundamental pieces of infrastructure, whether it's like browser use and how you can start to train models to be able to use, you know, the internet better when looking at a computer and so on and so forth.

So what are some of those examples to where you think that the post-training of open models will see new use cases basically unlock compared to just paying full price for the frontier models?

Vincent Weisser24:08

Yeah, I think, like, I can start on this. Like, I think the power that we've seen with a lot of different customers is it's really kind of this idea that, like, if you want to make a specific use case work, like, we can take the example of, like, if you want to figure out, like, a way that agents can actually automate your text, like, the most, like, concrete way you can do it today really is, like, build an RL environment for that use case, like, train on that and then deploy it into production with those users,right?

Like, let's say with, like, a million accountants that then now use this agent to ultimately get it towards full autonomy. It's a bit like, almost like Tesla's levels towards full autonomy where, like, you kind of need to deploy it into, like, do the last mile of actually, like, training for that specific use case, but then also deploying it to those specific users,right?

So, like, there's a reason why, like, a chatbot isn't good at self-driving because, like, it's not trained on that, it's not deployed into that context,right? And, like, I think it's the same even for these specific, like, knowledge work use cases where it's like if you want to have the perfect, like, financial agent, it's much more likely that you'll be able to get there if you have, like, RL environments for that use case.

If you deploy it into production, for example, as a bank,right, like, to millions of customers, then if you're, there's, like, one god model chatbot, like, and I think this is sort of like what we've seen now with a lot of verticals and customers that, like, there's, like, a huge unlock there to really go into these, like, specialized domains, post-train on them, deploy into them, and then continuously learn from production traces.

So we work with, like, some also big AI natives on things like computer use. We're ultimately having, like, millions of traces from production data really can help you to continuously improve those agents. And I think this kind of applies to almost every single domain.

And I think it's sort of the white pill for, like, the AI application builders and the AI startups to actually have a huge opportunity to build kind of their modes and to get to this data flywheel of, like, specialized models, even in a broader sense, like, just going after, like, let's say, computer use agents,right?

And I think, yeah, this is something where I think we're just seeing a lot of, like, movement, especially now with, like, open models catching up to the frontier. And I think the other piece is, like, optimization where, like, I think, like, GM is a great example, like, also, like, Trinity or Nemotron is, like, you have the whole ecosystem sort of, like, driving down the cost and optimizing it further,right?

Like, we are very able to work, like, very closely with all the teams here, but then also, like, deeply also with NVIDIA and with teams like VLM to really drive down the cost and make the, for example, inference and training for models like GM or, like, models like Trinity or Nemotron extremely efficient so you can basically drive down the cost, like, further and further.

And I think this is something you don't obviously get with the closed APIs where, like, they have, like, a huge margin on top. Like, they might drive down the optimization, but then might not pass through those savings. So I think in general, like, the open models are only getting through the open ecosystem, like, more and more efficient, like, and cheaper and cheaper to run and train on.

So I think there's, like, this element as well.

Chris Alexiuk27:07

I think too, like, a couple of things that I want to make sure we're very clear about is, like, most people probably do not need frontier level intelligence for, like, 90% of their tasks,right? Like, not to say that you're not doing cool smart stuff, not to say that I'm sitting here trying to do not cool smart stuff, but, like, a lot of the time these models are just overkill or they have, like, this really smooth, you know, capability horizon that means they're also quite good at chemistry.

But, like, most people are using models to do one or two things very well. And open models let you choose those one or two things and then make the model just very good at those things at the expense of, at the expense, sorry, of almost everything else.

And that is great. I mean, that's exactly what we should be doing,right? To use this model that is hyper-generalized and able to, you know, perform well across, like, 90 different axes is dope and cool, but it is not really, you know, using the model effectively.

It makes sense for someone who is trying to ensure that everyone can use this one endpoint to do their task, but it makes much less sense when you're a person who's trying to do that task yourself.

What was just said about efficiency is also deeply true,right? I mean, the idea that you are all here at a local AI summit, presumably you are running AI locally, presumably you would like it to be faster and better, and presumably many of you are quite cracked engineers,right?

This is a whole room of people who is going to contribute in some small part to making the ecosystem just a little bit faster, just a little bit more efficient. And while it's true that closed companies can afford to hire great amazing teams of people, as we saw with Linux over the whole time that it's existed,right?

Linux is the thing that runs the internet, it runs networks, it runs all of these services that require it to be hyper-optimized in a way that I think you can only get when you have people who are trying to run as resource-constrained as possible.

And all that to wax poetic and say, this idea that, like, local AI and open models and the most efficient version of the model ecosystem is necessary to do it in the open. I think it's, in fact, not possible to do it behind closed doors because you're shutting too many people that could make that one small contribution out of the room.

Lucas29:48

I think that that it's important to state too that I don't think any of us agree that or are of the mind that closed models or frontier, you know, what OpenAI and just to name names, you know, Anthropic and others are doing it is not extremely beneficial or that don't use them.

I certainly, I use those models near every day. It's just that it's where does it fit in the future of this ecosystem? And just like, you know, Chris is alluding to, you can think of open models and self-hosted or kind of owned intelligence or LLMs as, like, the Linux layer, which you're beginning to see kind of take place.

Linux runs enterprises, you know, it runs the cloud. We're seeing a very similar thing take place with hyperscalers and Neo Clouds and providers like Fireworks, Togethers, Basetends, Modals. But just like Macs are one of the best ways to get work done individually in the same way that maybe using OpenAI and ChatGPT is the best way for you to do the vast majority of simple check my email, help me rewrite, you know, check for grammar, those kind of things.

It's accessible, it's easy, and for, you know, your average consumer and individual, it's pretty cheap if you're using, like, the $20 a month plan. In the same way that, you know, Microsoft helps, you know, the world of medium-sized to large businesses run on Microsoft and Windows because, you know, they're not as expensive as getting everybody a Mac.

And you're going to see a similar world play out there for some closed and open providers themselves. So it's all an ecosystem. You know, I don't want to give the impression that, you know, I think anytime you log into ChatGPT or Claude that you're committing a sin, only that as you are, you know, this is a conference for AI builders, AI engineers.

As you're looking at the best way to engineer your product or your service, that there is another layer you can go down into and it's becoming way more accessible than it used to be.

Chris Alexiuk31:54

I think that's a great point. And I think that the relationship between closed frontier models and open models will be one that is, it's constantly there,right? I think that we have, it's never, you'll never get the headlines to apply nuance and say that both will coexist and gain more usage and are going to be useful to everybody.

But that is kind of the de facto state that will not only currently exist, but will continue to exist. I want to spend the last few minutes here to really give the audience something that only you guys potentially can answer.

Predictions32:16

Chris Alexiuk32:33

Oftentimes I reflect about my time at NVIDIA and I think I feel as though I have a clear vision outside into there's definitely still a fog of war out there, but I have a vantage point that many people don't have.

And you guys, because the positions you are in as well. And so what is the thing if we are looking forward towards AI engineer world fair 2027 that you think if you were to make a bold prediction, let's say, let's not be conservative around the intelligence in the open source and ground it with some frame of reference, what do you think we can look forward to by this time next year?

Vincent Weisser33:16

I think one key aspect obviously that people are closely tracking is, like, sort of like just, like, capabilities of open frontier models. And I think they'll keep being very close to the general frontier. Potentially, like, with, like, now the speed or, like, the of those closed frontier models, like, slowing down, I think, like, they'll catch up even more.

And I think the most concrete thing that I think will be very exciting is, like, seeing the world move from sort of chatbots and now coding agents to, like, just general knowledge worker agents, I think, over the next 12 months,right?

Like, to see more and more of, like, kind of, like, everyone across every knowledge worker domain, like, adopt agents in their workflows, which I think, like, developers have with coding agents have probably done better than any other domain in the world.

But I think we'll see over the next 12 months, like, a lot of, like, domain-specific, like, knowledge worker agents. But then also I think domains like computer use agents and I think others will take off. Like, I think in a similar way that, like, coding agents have taken off, I think we'll just see almost, like, in some ways you could say, like, almost like the general intelligence for, like, the knowledge and digital domain before then hopefully moving on to physical.

And I think very concretely, like, I think we'll, like, in 12 months, I think it's pretty likely that we'll have, like, better than Fable, Meteor's level capabilities in open models. And I think there's, like, a huge opportunity to ultimately enable a huge crop of new, like, AI startups and companies.

So it's in some ways, like, you want to almost, like, write the levels of capabilities, like, to some extent, like, Cursor really only took off when, like, Opus was good enough to do coding,right? So it's like this is when, like, Cursor inflected.

And I think we'll see, like, hundreds of these inflections for, like, startups getting started, like, now, like, over the next year or two once, like, open model. And I think we've seen this literally a month ago with, like, GM 5.2.

I think, like, it was, I think, legitimately one of those moments when people were like, okay, this is now, like, similar to the Opus inflection point, feels like an inflection point for open models to be, like, extremely strong and ultimately enable a ton of new businesses.

And I think this will only continue, like.

Lucas35:24

Not so bold prediction is that Prime Intellect and Arcee are going to have a combined valuation of a trillion dollars. That's obvious. But I think that this is going to be a huge year for

this is probably going to be the most consequential year for, like, the future of how AI gets distributed. The Fable and GPT-5.6, you know, embargo, if you will, has left a lot of open questions, you know, no pun intended, about open models and where how this intelligence gets distributed and at what capability level it starts to be politicized and kept back.

And so sovereign intelligence is going to be very important. I think that if you were to take, you know, the general population of AI users, people that are using it every day, so, you know, upwards of a billion to two billion people if you, you know, include ChatGPT and Google and whatnot, maybe 0.00001% have ever used an open model, you know, I think, or, you know, run it on themselves with a multitude of different tools.

And I would hope that the work that we're doing and the community's doing and the way that we're advocating for open science and open models and open discussion,right, that's probably been the most frustrating thing about the last couple of weeks is that all of these conversations around capabilities and who gets to use them and who doesn't have been happening behind closed doors.

My hope and I hope that I can predict that we will be able to have 10% to 15% of people that have ever used AI have used a model locally on their system and that that becomes a very important part of ensuring that you have access to what you need.

And so it's a prediction. It's also something that I know all of us up here and you out there are going to try to fulfill. And I hope that we can continue to advocate for that because if we're quiet, if we just let these things play out the way they are, open models will, you know, will be put under the microscope in the context of untrustworthy, unsafe.

And as much as there's work and vitriol and weaponized terms being out there advocating for that, we need to be combating as much of that, if not more, with the reasons that it deserves to exist.

Chris Alexiuk37:54

Yeah, I just, I could not plus in affinity what the last part, what Lucas said more. I think this is going to be the most consequential year for open intelligence that will, at least from where I sit, determine the future of a summit like this,right?

I think it has a potential to look very different in two radically opposed ways. As for bold predictions, I think that we will not be needing to go to an API for most of the tasks that we all do each day with AI.

I think it's likely to assume that you'll be running a model that is sufficiently capable in, let's call it day-to-day work on your MacBook within the year. It's already extraordinarily close, so not maybe not that bold of a prediction, to be honest with you.

I also think that we're going to continue to see models become the future of AI, so not model,right? Swarms of or specialized systems of models, I think, are going to be increasingly important. And lastly, on the open model front, I think we're going to see some very large architecture shifts, especially as we start to crack things like diffusion models for text a little bit more to get us models that are better suited for the hardware that we have in our houses.

And then last, mimi one, I think you're going to buy computers with agent operating systems on them instead of traditional operating systems, similar to, like, buying a Spark preloaded with Hermes or whatever. I think that's likely to occur.

Lucas39:44

I also predict that come September, when the next iPhone comes out, you're going to get a lot of texts from family members asking about this magical new Siri. So a lot of people who have not engaged with AI are about to in a very real way.

And the response to that's going to be very, very cool. And so just like when DeepSeek came out, I'm sure a lot of y'all got questions about what's this DeepSeek thing? There'll be another one in September. Get ready for it.

Vincent Weisser40:07

Yeah, I think that this might be actually one of the consequential, almost, like, unlocks,right? It's like, I think combination of, like, basically open models getting good enough as well as, like, the on-device compute getting strong enough to serve the equivalent of, like, today's frontier models,right?

Like, in a year or two. Like, basically, if you can, like, run Opus, like, at decent speeds on your, like, phone or laptop, I think the majority of humanity will probably, like, run local models. Like, and I think this probably applies more to the consumer than to the heaviest, like, enterprise agents.

But I think it seems pretty likely to me that, like, there will be this inflection point, even then, like, almost, like, similar to a new platform shift where, like, you can almost, like, tap into the local compute of a phone or laptop and then, like, start a next generation of almost, like, AI-enabled applications without, like, that can ultimately really, like, leverage the local compute of, like, device or on-device compute.

Lucas40:58

You can run a 4 billion parameter model on your phoneright now that is way more useful than GPT-4 was when it came out. And I think it's important for us to continue to focus on how do we best utilize that in the most meaningful way possible, as well as chase the newer capabilities that will come from things like, you know, drug discovery and scientific exploration.

It's going to be a fun couple of years.

Outro41:25

Chris Alexiuk41:25

Absolutely. If I were to try and summarize, I think that, you know, we're going to learn a lot more about how these local models are built incredibly in the systems and their relationships as they, you know, interact with frontier models over the next panels.

But this panel really shows that I think that we are at an inflection point to where if you think about how, you know, not even a short six years ago, it was AI was really for the research crowd and not really many people cared about it.

And then, of course, it came into the public consciousness with ChatGPT. But now there's this next thing, which is that open source is now really starting to enter the public consciousness, but very few people have touched and played with it and have had that aha moment.

And it sounds like we have the potential to do that this year and sort of guide the future wisely. But ultimately, it's up to a lot of the builders in this room as well to leverage that and represent, you know, this important inflection point that we're in on the side that hopefully brings, you know, intelligence, more intelligence to all of us, which is ultimately, I think, what everyone in this room would agree is sort of the direction of progress.

Vincent Weisser42:28

And, you know, everyone has said it on a panel previously, so I'll just also say it, which is that, and both of you have already said it, in fact. Like, you guys are extraordinarily important to this goal. Every one of you who is in this room and your friends and whoever, whatever communities you're part of, without you guys, we lose the fight,right?

So thank you for showing up, and I can't wait to see what we all build together.

Carter Abdallah42:53

And with that, thanks, Vincent and Lucas, Arcee and Prime Intellect, and of course, Chris from NVIDIA.

Lucas43:01

Always, always, of course.