← All AI Engineer talks

AI Engineer World's Fair 2026

Local Models: Trust, Control, Optimization

About this talk

A Local AI panel featuring publicly scheduled representatives of NVIDIA, Prime Intellect, and Arcee AI examines why builders want local and open models that provide ownership, durable access, privacy, and greater control than centralized APIs. Panelists describe collaboration around NVIDIA Nemotron and Arcee Trinity, discuss trust in open ecosystems and model provenance, and argue that efficient inference, accessible hardware, and developer participation are essential to advancing open frontier intelligence.

Chapters

  1. 0:00Panel introduction: local AI and participating organizations
  2. 1:21Prime Intellect and Arcee AI on open models and ownership
  3. 4:19NVIDIA Nemotron: openness, speed, and local hardware
  4. 6:06Trust, open-source concerns, and durable access to models
  5. 15:57Integrated AI products, cloud tradeoffs, and developer optimization
  6. 32:29Predictions, community participation, and panel closing

Talk transcript

  1. 0:00

    [upbeat jingle] So I hope everybody had a great lunch, and you got to check out some of the amazing demos that we have.

  2. 0:18

    Uh, we're gonna begin th-the panel, the first panel of the afternoon here, where we're gonna be talking about, uh, of course, the engines that are actually powering the stuff that, uh, you know, could remotely, uh, be used for things like local, sovereign, any kind of ownership over your own artificial intelligence, and of course, the engine powering those

  3. 0:36

    in addition to the hardware is the models themself. And so for this panel, we have excellent guests. We have Vincent, who's the CEO and founder of Prime Intellect. We've got Lucas, the CTO of Arcee AI.

  4. 0:47

    And we've got Chris, who is the senior product research engineer on the Nemotron family of models at NVIDIA. Now, what's really cool about working in this industry is, uh, really cool companies like this, we all get to work together.

  5. 0:59

    And so this is one panel where we all directly get to work together on both models, infrastructure, some of the ways that we think that the direction of the industry should go.

  6. 1:07

    Um, and each of us kind of play a different role in that stack. Uh, but I wanna leave it to you guys to, to introduce yourself and be able to talk about sort of the, the charter that you see, the problem of the w- stack that you guys are working on.

  7. 1:21

    Awesome. Should I kick it off? Um-

  8. 1:22

    Kick it off.

  9. 1:22

    Yeah. So I'm, I'm Vincent, as you mentioned, and really the goal with Prime Intellect from the beginning was, like, to ensure that basically frontier intelligence will be open and accessible, um, not just the models, but also the full stack to, to train the models.

  10. 1:36

    So kind of like this was like our, uh, motivation from the beginning. And we, we've, um, yeah, like, worked also together with a lot of, uh, gentlemen on the, on stage.

  11. 1:44

    Like on the one side, it's like, uh, we, we, we work with folks like, uh, Lucas and Arcee to help them train, uh, frontier open models. We, we, we help also, um, like NVIDIA on the Nemotron coalition help out their, uh, frontier open models.

  12. 1:57

    And I think, like, I'm actually, um, think both like Nemotron and, and Trinity might be the best, like, two open models right now outside of China. So I think it's actually, uh, like we, we need to fact-check that, you know.

  13. 2:08

    But this is actually from my, uh... I think they, they might be.

  14. 2:11

    It's [laughs] our, our marketing says that, yeah. [laughing]

  15. 2:15

    But, um, yeah, so, so that's the high level.

  16. 2:18

    Um, my name's Lucas Atkins. I'm happy to be here and, and thank you for joining. Um, very similar to Vincent, Arcee was, uh, you know, founded with the idea of, um, domain-specific owned models are, are, are gonna be needed.

  17. 2:33

    Um, you know, we were founded early twenty twenty-three, uh, jumping on the custom model, uh, train quite early. Um, you know, you have all these people who are excited about AI and all the things these new generation of LLMs can do, um, but they're using these monolithic, very expensive closed APIs, uh, for at the time and still,

  18. 2:54

    like very narrow tasks that don't require, uh, you know, at the time it was a hundred dollars per million tokens out. Um, and through doing that, we were building on top of open models, and we were releasing a lot of our tooling in the open.

  19. 3:07

    Um, and, uh, we noticed that in the United States, um, and in the West, you know, in general, we were starting to lose, uh, leadership in the open model space.

  20. 3:17

    A lot of it was coming out of China, and that's amazing. I love those models. We learn a lot from them. Uh, we're close with a lot of the people building those.

  21. 3:24

    But when you're working with large enterprises and companies and, um, geopolitics gets involved, whether you like it or not, you have people that become concerned about where those models are coming from.

  22. 3:36

    And we, uh, decided that, you know, we had a, a good group of people, and we had a good group of partners like NVIDIA and like Prime Intellect, where we could probably try to pre-train ourselves.

  23. 3:46

    Uh, so last year we did that. We kind of, uh, reoriented the entire company towards let's figure out how to pre-train a, you know, a four hundred billion parameter model in six months.

  24. 3:55

    Uh, and a lot of people said it was impossible, and by m- in many ways it was. Uh, but we figured it out, and now we are an open model lab, uh, working with our wonderful partners, um, and our customers to build, uh, Western open models that are permissive.

  25. 4:10

    Uh, a-a-and, uh, you can own those and customize them or run them wherever you want. Uh, and, and that's kind of where we're at right now. So thanks for having me.

  26. 4:19

    Yeah. And, uh, so I'm Chris Alexiuk. I work, uh, at NVIDIA as a product research engineer, and I support the Nemotron family of models. I think it's-- You know, we, we've talked a lot about why we do Nemotron, but just to say it a few more times, uh, you know, AI should be open, open as in weights,

  27. 4:37

    data, training methodology, training frameworks. Uh, really respect a lot of the work that, uh, the, the, the, the two other peeps up here do because they, they believe that very strongly as well.

  28. 4:48

    Uh, but the Nemotron family of models is focused on being as open as humanly possible. So, uh, we, we have this understanding or belief that, uh, in order for AI to continue to grow and be useful to everybody, uh, it has to be done in the open so that we can build off of each other, we can

  29. 5:06

    compound on each other. Uh, and, uh, part of what we do because, uh, team green, this is always true, uh, is we, we think that, uh, the, the rate that you can squeeze tokens out of models is very important.

  30. 5:20

    So, uh, we kind of have this mantra that like faster models are smarter models. And so a lot of the decisions we make when designing a model like Nemotron is, uh, built around how fast can we make it go.

  31. 5:33

    Uh, as especially you are gonna see i-in the next however many months local AI take off, uh, we, we need to make sure that models are well supported on, uh, hardware that doesn't just exist in massive buildings, you know, thousands of kilometers away from you.

  32. 5:51

    Uh, and so that's, uh, you know, for AI to be very useful, it should be quick, uh, and, and open. So that's kind of the, the vibe of Nemotron.

  33. 5:59

    Who makes those buildings with the massive processors?

  34. 6:02

    Oh, that's, uh-

  35. 6:04

    A lot of excellent people in the world that-

  36. 6:06

    Okay

  37. 6:06

    ... uh, they use a lot of excellent hardware from a, a pretty cool company. [laughs] Yeah, I heard. I heard anyway.

  38. 6:11

    Yeah. A- and I'm Carter Abdallah, I'll be your moderator for today. Uh, something that, you know, we all kind of talked about is, is this l- you know, building on top of each other, learning from others.

  39. 6:20

    Whether it is people, you know, across the, the big pond of the Pacific Ocean from us. Um, but really it is kind of like a collaborative sort of research effort, and I imagine that a lot of the people here in this room share that sentiment.

  40. 6:30

    But as it was brought up during the, you know, inaugural panel this morning in the State of the Union, there, there is a growing sentiment, um, potentially on the other side that, that paints, uh, open source to be something that is actually, uh, more chaotic.

  41. 6:43

    That there, there's less trust involved. And I think trust ultimately, as Lucas, you and I were talking about before, um, depending on who the party is, and depending on what lens you're looking at, at it from, I think it kind of means different things.

  42. 6:55

    Yeah.

  43. 6:55

    But ultimately from, from the, the end consumer, the somebody who's using this intelligence, um, or somebody who's, you know, more of a business and is actually customizing something to maybe monetize tokens in, in their business.

  44. 7:07

    Um, can you comment, we'll start with you Lucas, a bit on how, uh, open source and open source models are actually key to building that trust-

  45. 7:16

    Mm-hmm

  46. 7:16

    ... so that when these people walk out of this room and somebody does come at them with that other angle, they can, they can sort of steel man this side.

  47. 7:22

    Certainly. Uh, you can weaponize any term. Uh, and, and certainly trust is, i- is, has been weaponized, that word. Um, and the reason I say that is because, uh, it means something, uh, in, based on the context in with your speech, you're speaking about it.

  48. 7:36

    Uh, often, uh, in AI people like to conflate trust with safety. Um, and those are not the same thing. Uh, and I'm happy to speak on safety, uh, you know, later on.

  49. 7:47

    But when it comes to trust, um, I, I think that, you know, you hear a lot from closed model providers or, uh, politicians or people out in the space who are advocates for or a- against open source that you can't trust these open models 'cause you don't know what went into them.

  50. 8:05

    Well, the same is true for these closed models, uh, even more so. Uh, the benefit of, uh, uh, of, of open models is that, uh, we can very easily validate what is inside of them. [laughs]

  51. 8:17

    Uh, they are... You can... There is a whole bunch of files with a whole bunch of matrices in there, and you can view them, and you can see the code that is running these models.

  52. 8:25

    You have implementations from Prime RL, VLLM, SGLang, the provider themselves. These models are inherently trustworthy. You know much more about what's going on when you hit and talk to these models than you ever will what's going on when you hit an arbitrary API.

  53. 8:41

    Now, that being said, um, certainly there is fear that people can, um, uh, reduce, uh, you know, the... You can't trust that these models are writing safe code. Well, again, that is the same thing with any model.

  54. 8:55

    You need to, you need to use your, uh, your judgment, and you need to make sure that you have the proper, um, you know, safeguards in place, and you're viewing the outputs of these models as the outputs of, uh, uh, an inherently random system that we are working very, very hard to make less and less random.

  55. 9:13

    I think that a, a telling thing is a lot of people said, "Well, you can't trust Chinese models, you can't trust Chinese models, you can't trust Chinese models." That was often, uh, for the last few years meant you can't trust open models.

  56. 9:24

    Well, as soon as, uh, Anthropic had to put Fable away and people realized that, oh, our access to these frontier systems might not be universal anymore, uh, there's probably gonna be a lot of checks and balances.

  57. 9:37

    You had a tremendous number of enterprises and developers and companies start going to these new Chinese models because they could trust that they would always have access to them.

  58. 9:47

    Um, and so when, when it comes down to trust and the way I view that word as it relates to open models is, do I know that what I am running, and can I be as sure as possible that when I send something to this model, that I am gonna get the output that I expect?

  59. 9:59

    Um, and the only way, uh, currently to be 100% sure that what you are getting i- is what you are expecting is by, uh, hitting an open model, uh, either that you are running yourselves or you're working with a partner like Prime Intellect or Arcee or NVIDIA, uh, to validate.

  60. 10:14

    So that's my take on the word trust.

  61. 10:16

    I think too something you mentioned is like we don't get to know a lot about the data that goes into these models, and that's something that I'm really happy, you know, that, that we're trying to do.

  62. 10:27

    Yeah.

  63. 10:27

    Which is not some... It's not... You know, the incentives don't exist for everyone to do this, right? So it's not something that I think is mandatory or should be mandatory, thanks to the things that Lucas mentioned, which is that it's rather straightforward to validate what data did go into a model without seeing the data sources originally.

  64. 10:43

    Uh, but I'm happy that NVIDIA continues to release data sets along with our models.

  65. 10:48

    Yeah.

  66. 10:48

    Uh, release environments along with our models to make sure that even, even if you can't go through the work of determining what went into the model, uh, which you can do with the, the weights alone for the most part.

  67. 10:59

    Uh, you, you have like a spreadsheet you can look at that says, "Here's, you know, uh, a, a couple trillion tokens of this data set, a couple trillion tokens here."

  68. 11:07

    And I think that, that helps to educate people on why it, it, it's, uh, much easier to trust open models than, uh, than models that we, we don't get access to a- any of that.

  69. 11:19

    It also helps people see what that data looks like.

  70. 11:22

    Yeah.

  71. 11:22

    You know, if you don't have someone releasing it openly, when someone says data is going in, I, I mean, data can take many different shapes. You can... But you can go to Hugging Face, you can go to NVIDIA, or you can go to Prime Intellect or Arcee's, uh, Hugging Face.

  72. 11:34

    You can look under our data sets and you can see exactly what that looks like. Um, and that can help you, uh, inform your priors on it.

  73. 11:40

    Yeah, I think that trust also, um, you know, there's some angle of, of a reputation. Do I believe that your, your intentions are pure? And I think that a lot of people, again, in this room believe that intelligence is, is kind of this next layer of, uh, almost, you know, infrastructure for, for us to progress-

  74. 11:55

    Mm-hmm

  75. 11:55

    ... as a species, and I believe that everybody should have intelligence. So, uh, on, on that front, I wanna, uh, hand it over to Vincent because, uh, you kind of have this, um, almost like founding thesis that you...

  76. 12:05

    This, this stack should be the open, right? The open super intelligence stack. You want everybody to have a lot more intelligence. Um, can you talk about how, uh, the-

  77. 12:15

    This is kind of moving into the era of control, but, uh, beyond just data sets, how important it is to have the knobs and dials of th- of this industry also be, uh, available in an open way for people who are building this.

  78. 12:28

    Yeah, like I, I think it's, like, a really important point is to sa- um, so make sense, like, be able to take those open models, like, customize them, be able to, like, build on top of them.

  79. 12:36

    And I think, like, all the different components that go into it, like, especially from, like, the pre-training to mid to post-training, I think, like, need to be more accessible, right?

  80. 12:44

    Like, so more people can also, like, take those, uh, amazing models and, like, make them work for their specific use cases. So I think when we started, like, we, we also took a look at the whole stack that was out there and, and tried to figure out, like, what is missing for ourselves to train open models and

  81. 12:58

    for, like, helping our partners to do so. And a lot of this was around the RL and post-training stack. So we basically went deep into building out, like, a lot of infra around that, like, around our environment evolves, around, like, making it much more accessible to do post-training also because it's, like, the most economically viable way to

  82. 13:15

    maybe, like, customize those models, to take an open model and, um, to have, like, a specific eval environment and d- the, um, specific domain and dimension that you want to improve it on.

  83. 13:24

    And this is kind of, like, what we're really doing with Prime Intellect now is, like, enabling people to post-train specialized agentic models. Um, so being able to take models like, um, Trinity, for example, from Arcee or Nemotron or others and, and specialize them, post-train them for the use cases that ultimately enterprises care about.

  84. 13:40

    So good example of this was, like, a c- a company like, for example, Ramp was able to, like, take an old model and, like, specialize it to automate finance within, like, a week or two to get, like, better performance than, like, Opus at a fraction of the cost of HighQ.

  85. 13:53

    And I think really this Pareto frontier of, like, being able to d- create these specialized models that are much better than the frontier, but also faster, cheaper-

  86. 14:02

    Yeah

  87. 14:02

    ... um, I think is, like, a key thing enterprises care about increasingly, is really, like, just making it work for their use cases, basically.

  88. 14:08

    If you go back to trust, it's how you can make your CFO trust you- [laughs] ... by knowing exactly how much something's gonna cost all the time. Uh, that's, that is increasingly becoming very important is, uh, you hear a lot, you know, all these companies have unbelievably large token spend and, um, they're having to cut back on their

  89. 14:24

    Opus usage because they burned through it all in a couple of months. Um, and that is going to continue to be a problem because, yes, the cost of an individual token has come down drastically.

  90. 14:35

    You can look at it, you know, the difference between GPT-4 when it first launched and GPT 4.5 is, is, is much, much cheaper per token. But at the same time, the amount of tokens in an individual session has gone up e- e- exponentially as well.

  91. 14:47

    So we're kind of, um... We're spending more, uh, on a, uh, on a total session, and so the ability to, uh, bring in-house or, or, or at least work with partners to ensure that you are, uh, controlling your cost and you're not at the whims of, uh, when a company releases a newer model, uh, that might be

  92. 15:06

    better but also more expensive. They might deprecate a model. Um, o- o- owning that and being sure that, you know, same way is what you s- what out-- input goes in, you know what output's gonna come out.

  93. 15:16

    In the same way, um, when it, you know, an input goes in, how much it's gonna cost. Uh, uh, having assurance on that's really important too.

  94. 15:23

    Yeah, and maybe, like, one thing to add to this is, like, almost like I like this new, uh, term of like instead of, uh, speaking about token maxing and all, speaking more about, like, the outcome maxing of like-

  95. 15:31

    Yeah

  96. 15:31

    ... ultimately, it's like you, you want to have, like, more than a dollar worth of value come out of, like, a dollar of, of input. And I think that's the sort of, like, Jevons paradox of, like, if you can create more value, uh, for, like, your flop, for your GPU, basically-

  97. 15:44

    Mm-hmm

  98. 15:45

    ... I think this is sort of, like, how you'll get, like, the most adoption also of, like, agentic models. Like, if they can, like, be able to create as much value as possible.

  99. 15:52

    And I think the cheaper those models get, the more usage they'll get, like, for those specific use cases.

  100. 15:57

    It's funny you say that. I have a, um... And I think a, a lot of us in this room, but especially on this panel, bel- believe this to be so that, uh, the most meaningful AI applications in the next couple of years, even this year, are the ones where the harness and the model and the product, they

  101. 16:10

    all kind of blend together. Um, i- if you think back to, at least for me, the first, like, truly game-changing agentic experience I, it had was when, uh, Deep Research from OpenAI, and that was because they spent a tremendous amount of time doing reinforcement learning, uh, on o3 with test time compute to do these longer running research

  102. 16:31

    tasks that people had tried previously, but they were kind of just doing a for loop over search. Um, whereas, uh, I kept coming back to, to Deep Research. And, um, you know, you saw for a very long time that OpenAI and Anthropic and Google, when they'd release a new product, they'd release a custom version of their model

  103. 16:47

    for that product. And if they're doing that, if their off-the-shelf GPT 5 isn't good enough for, you know, their Atlas web browser, why should it be good enough for our apps?

  104. 16:57

    Um, and that's why I appreciate the work that, you know, Vincent, uh, and NVIDIA are doing for, for giving people the tools to customize their own model, um, and allowing us to focus on how we get a good model to start from.

  105. 17:09

    Uh, so it really is, um, it, it, you know, it's, it's extremely important as you look at developing applications and experiences over these next few years that you're taking into account that you can, uh, make the model do something that, um, maybe your harness isn't fully able to do alone.

  106. 17:24

    Well, that's something I think that's really important to just, like, reiterate, right? I mean, like, N- Nemotron's great. I love it. Trinity is great. I love it. Like, n- we design a model that's supposed to be as good as it can be across a number of harnesses, right?

  107. 17:39

    You can see this in the technical report. The idea is, like, we want the har- the model to work as well as it can in Py compared to, you know, Hermes compared what- whatever you're using, right?

  108. 17:48

    Uh, but, like, you're, you're not using all of these tools at once. You're using one of these tools. And so when you have open models, you can do things like, uh, Noose Research can create a post-train of whatever model for their harness, right?

  109. 18:03

    That you know will be- ... extra good. And, you know, this, this, uh, this thing from, from, from, uh, you know, I can't remember who, who originally wrote it, but this idea of, like, the mismanaged genius, right?

  110. 18:14

    We're, we're leaving a lot of, like, uh, a lo- lot of important capability on the table, uh, because we're just not s- we're not fitting the models into the harness right.

  111. 18:25

    You can do a bunch of stuff with closed models. Like, you can change your prompts and your skills and all kinds of other neato things, right? But nothing will let you get the, the, the level of customization or customability, uh, that you can achieve with open models.

  112. 18:40

    And I think that is something that is going to become increasingly and increasingly more important, especially thanks to folks like the others on the panel, where, you know, I can just straight drop, like, my favorite coding and, you know, a- an agent environment, uh, spin up the CLI, and suddenly my, my model feels way better, uh, with

  113. 19:00

    very little effort, right? Like, that is, that is something that is, uh, already at our fingertips, and it is only going to get easier and easier as, as time goes on.

  114. 19:10

    Yeah. And it's maybe also the most concrete, like, even for the builders and audience, like, call to action of, like, if you kind of want to build the next, like, Claude code, the next, like, Cursor or Perplexity, I think the easiest way to get started is, like, take the best open model, like, and, and then post-train it

  115. 19:25

    on your harness, like, that you care about, right? Like, basically, like, create a product that is, like, truly AI native, right?

  116. 19:30

    Yeah.

  117. 19:30

    And I think this is sort of like, I think, one of the most exciting, like, unlocks for builders, um, like out here.

  118. 19:36

    That's a big thing about the, the framing c- of control is, is, um, similar to, like, you know, when, when the cloud, um, explosion started to happen in the, uh, in the, you know, the late, uh, 2000s, early 2010s, and the, um, social media world kind of took off and apps became extremely popular and more, you know,

  119. 19:56

    cloud-based and, and, you know, managed by these, these bigger companies. Uh, you know, the data that they were collecting from you, whether anonymized or not, was how they were monetizing their platform through ads or, or, or, or in other ways.

  120. 20:10

    Uh, and a very similar thing, uh, has always been happening, but I think is becoming clear to people in, in the space is the, the, the data that people get from, from you using these models is how these companies largely make their models better, whether it's through actually training on that data or, um, by using it as

  121. 20:25

    a signal for what data they, they go out and find or generate to train. And, uh, with closed models, there are terms of service that keep you from being able to, um...

  122. 20:35

    Well, don't-- I, I could get into a, a discussion about what-

  123. 20:38

    Yeah

  124. 20:39

    ... terms of service is an agreement between you and the provider. It's not a legal... Uh, anyway, but you, you, you shouldn't be training on a, you know, a Claude Opus output or a Fable output or a GPT-5 output, and they do a lot to try to obfuscate to, to make that not great for you.

  125. 20:53

    If you're using an open model, you can save all of those traces, all of those traces of you using it inside of your harness, uh, that will allow you over time to if you say, "Hey, I want to go train a custom model," you can take all of that and again, either use it to directly do, like,

  126. 21:08

    fine-tuning on a smaller model so you're not spending as much or to have a model help you find signals so that you can go out and use verifiers or NeMo RL or NeMo Gym to create these environments so that you can hill climb and make your models better.

  127. 21:21

    So, um, as much as using open models is like owning your stack, owning your intelligence, it's also owning your outputs, right? Owning your data. That's gonna be extremely important too.

  128. 21:31

    I do wanna, I do wanna plug the license for a second. So [laughing]... [laughs] Uh, so AI is very different than traditional software, uh, which is why recently, uh, Nemotron as well as, uh, Trinity, I know-

  129. 21:45

    Yeah

  130. 21:46

    ... uh, has adopted the OpenMDW model, uh, data weights, uh, license. Uh, the idea is, like, we, we need a way to really make it very clear in the license that you can use the outputs to produce a model.

  131. 21:59

    You can use the outputs to train, uh, right, all of these TOSs and stuff like that, that, that, that, that have language that is meant to dissuade you to do that.

  132. 22:08

    Uh, w- we wanted to make sure there was a license that exists that im- not encourages you, but makes it crystal clear that it is, it is permitted, it is permissible.

  133. 22:17

    Uh, and, and I think, you know, the licenses maturing, right, to fit the use case better should be extremely positive signal, uh, for, for the way that the ecosystem is thinking about open models, uh, to the, to the fact where even, even the lawyers are on board.

  134. 22:32

    Do you know how much lawyers cost? [laughing] [laughing] Yeah.

  135. 22:35

    Yeah. It's, uh... As we move on to the, you know, kind of the third topic, which here is o- of course optimization, um, something that, you know, we've implicitly said but haven't said it, it quite explicitly yet is, uh, effectively that I, I think that for a lot of people, there's this preconceived notion that when you're deciding

  136. 22:52

    to use an open model for whatever the use case, um, there, there are the trade-offs that come in the form of performance at the benefit of getting things like, you know, maybe data sovereignty and so forth.

  137. 23:03

    Um, but what we are now talking about is that with the, with the right customization and optimization, depending on the, the use case that you're... and the harness that you're applying it to, you can actually exceed and build the model against the tool-

  138. 23:17

    Mm-hmm

  139. 23:17

    ... to, to get better performance than even frontier models. Um, I'd love to hear a, a little bit more about the... 'Cause I think that, uh, another thing that we would probably agree on is that the, the current level of intelligence, um, already has so much, uh, left to diffuse into society.

  140. 23:35

    Yeah.

  141. 23:35

    And so where are those areas where that diffusion is, is, is happening in the specific industries? I know, for example, things around, uh, again, kind of fundamental pieces of infrastructure, whether it's like browser use, um, and how you can start to train models to, to be able to use, uh, you know, the, the internet better when looking

  142. 23:53

    at a computer and so on and so forth. So what are the, so what are the some of those examples to where you think that the, uh, post-training of open models will, uh, see new use cases basically unlock compared to just paying full price for the-

  143. 24:06

    Yeah

  144. 24:07

    ... the frontier models?

  145. 24:08

    Yeah, I think, like, like, I can start on this. Like- I think the, the power what, that we've seen with a lot of different customers is, is really kind of this idea that, like, if you want to make a specific use case work, like, we can take the example of, like, if you want to figure out, like,

  146. 24:22

    a way that agents can actually automate your text. Like, the, the, the most, like, concrete way you can do it today really is, like, build an RL environment for that use case, like train on that, and then deploy it into production with those users, right?

  147. 24:33

    Like, let's say with, like, a million accounts that then now use this agent to ultimately get it towards full autonomy. It's a bit like, almost like Tesla's levels towards full autonomy, where, like, you kind of need to deploy it into, like, do the last mile of actually, like, training for that specific use case, but then also, um,

  148. 24:48

    deploying it to those specific users, right? So, like, there's a reason why, like, a chatbot isn't good at self-driving because, like, it's not trained on that. [laughs] It's not deployed into that context, right?

  149. 24:58

    And like-

  150. 24:58

    Yeah

  151. 24:58

    ... I think it's the same even for these specific, like, knowledge work use cases where it's like, if you want to have the perfect, like, financial agent, it's much more likely that you'll be able to get there if you have, like, RL environments for that use case, if you deploy it into production, for example, as a bank,

  152. 25:12

    right? Like, to millions of customers than if you're, uh, there's, like, one got model chatbot. Like, and I think this is sort of like what we've seen now with a lot of verticals and customers that, like, um, there's, like, a huge unlock there to, um, really go into these, like, specialized, um, domains, post train on them, deploy

  153. 25:29

    into them, and then continuously learn from production traces. So we work with, like, some, um, also big AI natives on things like computer use, where ultimately having, like, millions of, of traces from production data really can help you to, to, um, continuously improve, uh, tho- those agents.

  154. 25:44

    And I think this kind of applies to almost every single domain, and I think it's sort of the, the white pearl for, like, the AI application builders and, and the AI startups to actually have a, a huge opportunity to build kind of their modes and, and to get to this data flywheel, um, of, of, like, specialized models,

  155. 26:01

    even in a broader sense. Like, just going after, like, let's say computer usage agents, right? And I think, um, yeah, this is something where I think, um, we're, we're just seeing a lot of, like, movement, especially now with, like, open models catching up to the frontier.

  156. 26:14

    Um, and I think the other piece is, like, optimization, where, like, I think, um, like GM is a great example. Like, also, like Trinity and Nemotron is like you have the whole ecosystem sort of like driving down the cost and optimizing it further, right?

  157. 26:25

    Like, like we are very happy to work like very closely with, with all the teams here, but then also, um, like deeply also with, uh, NVIDIA and with teams like VLLM to really drive down, uh, the cost and, and make the, for example, inference and training for models like GM or like models like Trinity and Nemotron extremely

  158. 26:42

    efficient. So you can basically drive down the cost like further and further, and I think this is something you don't obviously get with the closed APIs where, like, they have, like, a huge margin on top.

  159. 26:51

    Like, they might drive down the optimization, but then might not pass through those savings. So I think in general, like the open models are only getting through the open ecosystem, like more and more efficient, like, uh, and, and cheaper and cheaper to run and train on.

  160. 27:04

    So I think there's, like, this element as well.

  161. 27:06

    I think too, like a, a, a couple things that I, I, I wanna make sure we're very clear about is you, you-- like most people probably do not need frontier-level intelligence for like ninety percent of their tasks, right?

  162. 27:18

    Like, uh, like n-n-not to say that you're not doing cool, smart stuff, uh, not to say that I'm, I'm sitting here trying to do, uh, not cool, smart stuff, but like a lot of the time, uh, these models are just overkill or they have like this really smooth, uh, you know, uh, capability horizon that, uh, means they're,

  163. 27:35

    they're also quite good at chemistry. But like most people are using models to do one or two things very well, uh, and open models let you choose those one or two things and then make the model just very good at those things at the expense of, at the expense, sorry, of almost everything else.

  164. 27:53

    And that, that is great. I mean, that's exactly what we should be doing, right? Uh, to, to, to use this model that is hypergeneralized and able to, uh, you know, perform well across like ninety different axes is, is dope and cool, but, uh, it, it is not really, you know, using the model effectively.

  165. 28:13

    It, it makes sense for someone who is trying to ensure that everyone can use this one endpoint to do their task, but it makes much less sense when you're a person who's trying to do that task y-yourself.

  166. 28:25

    It, uh, what, uh, what, what was just said about efficiency is also deeply true, right? Uh, I, I mean, the idea that you are all here at a local AI summit, presumably you are running AI locally.

  167. 28:40

    Presumably, you would like it to be faster and better, and presumably many of you are quite, uh, quite cracked engineers, right? Uh, th-this is a whole room of people who is going to contribute in some small part to making the ecosystem just a little bit faster, just a little bit more efficient.

  168. 28:56

    And while it's true that closed companies can, uh, afford to hire great, amazing teams of people, as we saw with Linux over the, uh, whole time that it's existed, right?

  169. 29:07

    Uh, Linux is the thing that runs the internet. It runs networks. It runs all of these services that, uh, that require it to be hyperoptimized in a way that I think you can only get when you have people who are trying to run as resource-constrained as possible.

  170. 29:22

    And, uh, uh, all that to wax poetic and say this idea that, like, local AI and open models and the most efficient version of the, of the, uh, model ecosystem, uh, is, is necessary to do it in the open.

  171. 29:37

    I think it's, in fact, not possible to do it behind closed doors 'cause you're shutting too many people, uh, that could make that one small contribution, uh, uh, uh, out of the room.

  172. 29:48

    I, I think that that-- it, it's important to state too that w- I don't think any of us agree that or, or, or are of the mind that closed models or frontier, you know, what OpenAI and Anthropic, to name names, you know, and Anthropic and others are doing it is, uh, not extremely beneficial [laughs] or that not-- don't

  173. 30:07

    use them. Um, I, I, I certainly, uh, I, I use those models near every day. It's, it's just that it's where does it fit in the, uh, in, in the future of this ecosystem?

  174. 30:18

    Um, and just like y- you know, Chris is alluding to, you can think of open models and self-hosted or kinda owned intelligence or LLMs as like the Linux layer, which you're beginning to see kinda take place.

  175. 30:30

    Linux runs enterprises. You know, it runs the cloud. We're seeing a very similar thing take place with hyperscalers and neo clouds and providers like Fireworks together as base tens models.

  176. 30:40

    Um, but just like Macs are one of the best ways to get work done individually, uh, in the same way that maybe using OpenAI and ChatGPT is the best way for you to do the vast majority of simple, check my email, uh, help me rewrite, you know, check for grammar, those kind of things.

  177. 30:59

    It's accessible, it's easy, and for, um, you know, your average consumer and individual, it's pretty cheap, um, if you're using like the $20 a month plan. Uh, in the same way that, uh, you know, Microsoft, um, helps tr- you know, tr- uh, the, the world of medium-sized to large businesses run on Microsoft and Windows because, uh, um,

  178. 31:19

    you know, they're not as, uh, e-expensive as getting everybody a Mac. Um, and you're gonna see a similar, uh, world play out there for, for some closed and open, open projects themselves.

  179. 31:28

    So, um, it's all an ecosystem. You know, I don't wanna give the impression that, uh, you know, I think any time you log into ChatGPT or Claude that, that you're committing a sin, only that as you are...

  180. 31:39

    You know, this is a, a conference for AI builders, AI engineers. Uh, as you're looking at h- the best way to engineer your product or your service, that, uh, there is another layer you can go down into, and it's becoming way more accessible than it used to be.

  181. 31:54

    I, I think that's a great point, and I think that the relationship between closed frontier models and open models will be one that is, uh... It's constantly there, right?

  182. 32:04

    I think that we have, uh, it's never-- You'll, you'll never get the headlines to apply nuance and say that both will coexist and gain more usage-

  183. 32:13

    Yeah

  184. 32:13

    ... and, uh, are gonna be useful to everybody. Um, but that is kind of the de facto state that, that will n-not only currently exist but will continue to exist.

  185. 32:21

    Um, I wanna spend the last few minutes here to, uh, e- really

  186. 32:29

    give the audience something that, uh, only you guys potentially can answer. Um, oftentimes I reflect about my time at NVIDIA, and I think, I, I feel as though I have a, a, a clear vision outside into, um, uh, there's definitely still a fog of war out there, but I have a vantage point that many people don't have,

  187. 32:46

    and you guys because the positions you are in as well. And so w- what is the, the thing, if we are looking forward towards AI Engineer World Fair 2027, um, that you think s- if you were to make a, a bold prediction, let's say, let's not be conservative, around, uh, the intelligence, um, in, in the open source,

  188. 33:07

    um, and ground it with some frame of reference, what do, what do you think, uh, we can look forward to by, by this time next year?

  189. 33:15

    Like, uh, I think one key, um, aspect obviously that people are closely tracking is like, um, sort of like the, just like capabilities of, um, open frontier models, and I think they'll, they'll keep being very close, uh, to the general frontier.

  190. 33:29

    Potentially even like w- like with like now, um, the, the speed or like the, uh, of those closed frontier models like slowing down, I think like they'll, they'll catch up even more.

  191. 33:38

    And I think the, the most concrete thing that I think will be very exciting is like seeing the world move from sort of chatbots and co- now coding agents to like just general knowledge work agents, I think over the 12, the next 12 months, right?

  192. 33:49

    Like, to see more and more of like kinda like everyone across every, um, knowledge worker domain like adopt agents in their s- um, workflows, which I think like developers have with coding agents have probably done better than any other domain in the world.

  193. 34:02

    But I think we'll, we'll see over the next 12 months like, um, a lot of like domain-specific, like knowledge worker agents. But then also I think domains like computer use agents and I think others will take off.

  194. 34:11

    Like, I think in a similar way that like coding agents have taken off, I think we'll just see almost like i- in some ways you could say like almost like the, uh, general intelligence for like the knowledge w- and digital domain, um, before then hopefully maybe even moving on to physical.

  195. 34:24

    And I think very concretely, like I think we'll, like in, in 12 months, I think it's pretty likely that we'll have like better than, uh, Fable, uh, Meta's le- level capabilities in open, uh, models.

  196. 34:35

    And I think there's like a huge opportunity, uh, to ultimately enable a huge crop of new like AI startups and companies. So it's in some ways like you want to almost like write the, the levels of capabilities.

  197. 34:47

    Like to some extent, like Cursor really only took off when like Opus was good enough to do coding, right? So it's like this is when like Cursor inflected, and I think we'll see like hundreds of these inflections for like startups getting started like now, like in, over the next year or two, um, o- o- once like open

  198. 35:01

    model... And, and I think we've seen this literally a month ago with like GM 5.2. I think like, uh, it was I think legitimately one of those moments when people were like, "Okay, this is now like similar to the Opus inflection point.

  199. 35:13

    Feels like an inflection point for models to be like extremely strong and ultimately enable a ton of new businesses." And I think this will only continue, like...

  200. 35:23

    Uh, not so bold prediction is that, uh, Prime Intellect and Arcee are gonna have a combined valuation of a trillion dollars. [laughs] Um, that's obvious. Uh, but, but, but I think that this is gonna be a huge year for, um, uh, this is probably gonna be the most consequential year for like the future of how AI gets distributed.

  201. 35:42

    Um, the, the, the Fable, uh, and GPT 5.6, um, you know, uh, uh, uh, embargo, if you will, um, has left a lot of open questions, um, you know, no pun intended, about open models and, and where, um, how this intelligence gets distributed and at what capability level it starts to be, uh, politicized and, and, and kept

  202. 36:05

    back. And so sovereign intelligence is gonna be very important. I, I think that, um, if you were to take- You know, the general population of AI, um, users, people that are using it every day, so you know, upwards of, uh, a billion to two billion people if you, you know, include ChatGPT and Google and whatnot, uh, maybe

  203. 36:24

    .0000001% have ever used an open model, you know? I think it, it... Or r- or, you know, run it on themselves w- w- with the multitude of different tools.

  204. 36:33

    Um, and I would hope that the work that we're doing and the community's doing in the way that we're advocating, um, for open science and open models and open discussion, right?

  205. 36:47

    That's probably been the most frustrating thing about the last couple weeks is that all of these conversations around capabilities and who gets to use them and who doesn't have been happening behind closed doors.

  206. 36:55

    Uh, my hope a- a- and, and, and, and I hope that I can predict that we will be able to have, uh, a, a 10% to 15% of people that have ever used AI have used a model locally on their system, and that that becomes a very important part, uh, of ensuring that you have access to what

  207. 37:13

    you need. Um, and, and so it's a prediction. Uh, it's also something that I know all of us, uh, up here and, and you out there are gonna try to fulfill, and I hope that, uh, we can continue to advocate for that.

  208. 37:24

    Because if we're ver- if we're quiet, if we just let this, uh, things play out the way they are, uh, open models will, you know, w- will, will be put under the microscope, um, in the context of untrustworthy, unsafe.

  209. 37:38

    Uh, and as much as there's work and, and, um, vitriol and weaponized terms being out there, uh, uh, advocating for that, we need to be, uh, combating, uh, as much of that if not more, uh, with the reasons that it deserves to exist.

  210. 37:54

    Yeah. I just, uh... I could not, uh, plus infinity what, what the last part, what Lucas said more. I think this is going to be the most consequential year for, uh, open intelligence that, uh, will,

  211. 38:10

    it, it, it at least from, from where I sit, determine the [laughs] future of a, of a summit like this, right? Uh, I think w- it has the potential to look very different in two, uh, radically opposed ways.

  212. 38:21

    Uh, as for bold predictions, I think that we will not be needing to go to an API for, uh, most of the tasks that we all do each day with AI.

  213. 38:32

    I think it's, uh, likely to assume that you'll be running a model that is sufficiently capable in let's call it day-to-day work, uh, on your, on your MacBook, uh, within the year.

  214. 38:43

    Uh, it's already extraordinarily close, so not, maybe not that bold of a prediction to be honest with you. Uh, I also think that we're going to continue to see, uh, models, uh, become the, the future of AI, so not model, right?

  215. 38:58

    Uh, uh, swarms of or s- uh, specialized, uh, systems of models I think are going to be, uh, uh, increasingly, uh, important. Uh, and, and lastly on the open model front, I think we're gonna see some very large architecture shifts, especially as we start to crack things like diffusion models for text a little bit more, uh, to,

  216. 39:21

    to get us models that are, uh, better suited for the, the hardware that we have in our houses. And then last me, me one, I think you're gonna buy, uh, computers with agent operating systems on them instead of traditional operating systems, uh, similar to like y- buying a Spark preloaded with, with Hermes or whatever.

  217. 39:41

    Uh, I, I, I think that's, that's likely to occur.

  218. 39:44

    I also predict that come September when the next iPhone comes out, you're gonna get a lot of texts from family members asking about this magical new Siri. [laughs] So a lot of people who have not engaged with AI are about to in a very real way, um, and the, the response to that's gonna be very, very cool.

  219. 39:59

    And, uh, so just like when DeepSeek came out, and I'm sure a lot of y'all got questions about, "What's this DeepSeek thing?" Uh, there, there will be another one in September.

  220. 40:06

    Get ready for it.

  221. 40:07

    Yeah, I think that this might be actually one of the consequential, it's like unlocks, right? It's like, uh, I think combination of like basically open models getting good enough, uh, as well as like the on-device compute getting strong enough to serve the equivalent of like today's frontier models, right?

  222. 40:20

    Yeah.

  223. 40:20

    Like in a year or two. Like, like basically if you can, like run OPOS like, uh, at decent speeds on your like phone or laptop, I think the majority of humanity will probably like run local models.

  224. 40:31

    Like, and, and I think this probably applies more to the consumer than to the heaviest like enterprise agents, but I think, uh, it, it seems pretty likely to me that like there will be this inflection point even then like almost like similar to a new platform shift where like, um, you can almost like tap into the local

  225. 40:46

    compute of a phone or laptop and then like start a next generation of almost like AI-enabled applications without like... That it can ultimately really like leverage the local compute of like, um, device, on-device compute.

  226. 40:57

    You can run a four billion parameter model on your, on your phone right now that is way more useful than GPT-4 was when it came out, and I think it's important for us to f- continue to, to, to focus on how do we best utilize that in the most meaningful way possible, uh, as well as chase, uh,

  227. 41:15

    the newer capabilities that will come from things like, you know, drug discovery and, and, uh, um, and, and scientific exploration. It's, it's gonna be a, a fun couple years.

  228. 41:25

    Absolutely. If I were to try and summarize, I think that, you know, we're gonna learn a lot more about how these local models are, are in- built incredibly and the systems and their relationships as they, uh, you know, interact with frontier models over the next panels.

  229. 41:38

    But, uh, this panel really shows that I think that we are at an inflection point to where if you think about how, uh, you know, not even a short six years ago it was...

  230. 41:46

    AI was really for the research crowd and not really many people cared about it. And then of course it came, uh, into the public consciousness with ChatGPT, um, but now there's this next thing which is that open source is now really starting to enter the public consciousness, but very few people have touched and played with it, um,

  231. 42:02

    and have had that aha moment. And it sounds like we have the potential to do that this year and sort of, uh, guide the, the future wisely. Um, but ultimately it's up to a lot of the builders in this room as well to, to, to leverage that and represent, um, you know, this important inflection point that we're

  232. 42:17

    in, uh, on the side that hopefully brings, uh, you know, intelligence, more intelligence to all of us, which is ultimately I think what everyone in this room would agree is, is sort of the direction of progress.

  233. 42:27

    And you know, uh, everyone has said it on a panel previously, so I'll just also say it, which is that, uh, a- a- and, and both of you have already said it in fact.

  234. 42:36

    Uh, like you, you guys are extraordinarily important to this goal. Uh, every one of you who is in this room and your friends and whoever, what- whatever communities you're part of, uh, without you guys, uh, we, we lo- we lose the fight, right?

  235. 42:48

    So thank you for showing up and, uh, I, I can't wait to see what we all build together.

  236. 42:53

    And with that, thanks Vincent and Lucas, Arcee and Prime Intellect, and of course Chris from NVIDIA. [audience cheering]

  237. 43:00

    Always, uh, always of course. [laughs] [outro music]