Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
Read the talk
Why Agents Need More Context: MiniMax M3’s Attention, Multimodal Training and Research Workflows
Thomas Wolf and Olive Song discuss how long context, selective attention and training with text and vision fit together—and how model feedback and agent harnesses feed into the next round of research.
From a talk by Thomas Wolf and Olive Song
At a glance
Ideas worth remembering
Long context serves an agent’s accumulating interactions and tool responses. MiniMax Sparse Attention makes that context more tractable by selecting blocks with an index branch and computing attention over the selection.
Native multimodal training addresses problems MiniMax observed when adding vision later, but introduces its own stability challenge. Song credits vision-component work, interleaved data, cleaning, masking and reward modeling with preventing collapse.
Multimodal understanding becomes especially useful when it feeds subsequent action: presentations, reports and long videos can supply information for tool-using agents. The interview proposes these workflows without demonstrating their reliability.
MiniMax describes several routes from observed behavior to better models: internal evaluations and research projects, community feedback and pull requests, and automated research harnesses that help accelerate iteration.
The parameter figures are not fully consistent: the opening gives approximately 400 billion total and 20 billion activated, while the later discussion gives 428 billion total and 23 billion active. The supported distinction is between total and active parameters; the exact counts remain uncertain.
Song’s closing emphasis is on multi-agent systems and model routing, both to handle more complex tasks and to expose the boundaries of individual models’ capabilities.
An open-model conversation
After the musical introduction, Hugging Face co-founder and chief science officer Thomas Wolf welcomes Olive Song to discuss MiniMax. He frames the conversation around competition among leading open-source model teams and introduces M3 as MiniMax’s recent release. His leaderboard comparisons establish the interview’s context, rather than providing a detailed evaluation of the models.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Coding, vision and a million tokens
Song introduces M3 with approximate figures of 400 billion total parameters and 20 billion activated parameters. These are distinct quantities: the model’s total parameter count is larger than the count activated during its computation. She emphasizes the combination of coding performance, image and video understanding, and a one-million-token context window, enabled by an architecture called MSA, or MiniMax Sparse Attention.
The design goal is to bring these capabilities together for future applications. Coding and agentic behavior address what the model can do; multimodal understanding broadens what it can receive; longer context expands how much information it can work with at once. Wolf stresses that the million-token window is intended to be functional, then asks how the attention architecture makes that length efficient. The exchange supplies a capability claim, but no benchmark results or measurement procedure for that claim.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From reading a book to retaining an agent’s history
Song traces MiniMax’s long-context work to earlier models that she says could perform tasks with ten million tokens of context. Her example is supplying a book and asking for a review. She distinguishes that use from an agentic model: accepting a large input does not by itself establish that a model can manage repeated actions and interactions.
An agent’s context grows through interaction. It receives user messages, acts in an environment, gets tool responses and proceeds through multiple rounds. Song argues that a short window cannot hold enough of this accumulating information for complex tasks. Restoring longer context therefore serves a concrete purpose: keeping more of the material generated during the task available while the agent continues working.
MiniMax Sparse Attention separates selection from attention computation. An index branch identifies which parts of the context matter more; a sparse-attention branch then computes over the selected blocks. The efficiency mechanism is selective computation over context blocks. That makes the quality of selection consequential, although Song does not specify the selection rule, block sizes or how selection errors affect performance. She describes the simple architecture as a foundation for scaling both context length and model size.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Efficiency remains an architecture problem
Wolf places sparse attention in the history of efforts to address attention’s quadratic, or N-squared, computation. He recalls work on linear attention and the subsequent prominence of FlashAttention, then welcomes renewed investigation into attention’s underlying design. His comparison with GPT-2’s 1,024-token window illustrates how expectations for context length have expanded.
When Wolf raises trillion-token attention, Song treats it as a research direction. She says extreme context lengths would require work on architecture together with hardware. She also sees room for further architecture and inference optimization, particularly where applications need strong capabilities and efficient operation. No concrete trillion-token implementation or quantified efficiency gain is offered.
The architecture’s origin turns the discussion toward research access. Song credits an intern with designing it and contrasts MiniMax’s openness to contributions with labs where interns may lack access to data or substantive work. The example connects technical invention to who can use the resources needed to investigate an idea.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turning observed weaknesses into research projects
Song describes an internal process built on shared infrastructure. Researchers can use the model, devise their own evaluations, identify weaknesses and propose improvements. Interested colleagues join a project, work on it for weeks or months, and incorporate successful results into final training and a subsequent release. Model access and evaluation make a proposed improvement concrete before it becomes part of training.
Architecture projects can require a longer investigation. Song lists research, experiments and even redoing pretraining evaluations as necessary work. This gives the process room for deeper changes whose consequences cannot be established through a quick trial.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Training text and vision from the first step
Song calls MiniMax’s approach native multimodality. She contrasts it with completing text pretraining, adding adapters and then training vision understanding. In MiniMax’s experiments, she says, the later addition harmed text performance while vision performance did not converge well. Her explanation is that the model had already converged toward text understanding, making the subsequent adjustment less effective.
Introducing vision partway through pretraining presents another difficulty: sensitivity to the training recipe. Song names architecture, data mixtures and learning rates as variables that change the result. A recipe that works in one experiment may therefore fail to carry its conclusions to a larger model. Her concern is both control of the experiment and scalability of what researchers learn from it.
Training multimodally from the first step avoids that later transition, but Song acknowledges a serious failure mode: text and vision understanding can collapse after only a few training steps. She says MiniMax addressed this through work on the vision component and the training data. The data remain interleaved, retaining images and videos in their natural context, alongside cleaning and masking. She also credits reward modeling. These are the ingredients she identifies; the interview does not provide the detailed recipe or isolate each ingredient’s contribution to stability.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Larger models and interfaces people can use
Wolf asks about scaling beyond a trillion parameters, giving figures of 428 billion total parameters and 23 billion active parameters for the model under discussion. Song says MiniMax intends to pursue larger models because some tasks perform poorly with smaller parameter counts. This is an ambition and a motivation, without a specified future architecture or evidence showing which tasks require that scale.
The conversation then moves from model scale to MiniMax’s applications. Wolf asks whether apps and their data came first. Song describes the company as model-led from its beginning: the founding ambition was a model that could understand and output across modalities. That broad ambition should be distinguished from the specific capabilities demonstrated or described for any individual release.
Apps followed as a way to make model capabilities accessible. Song argues that people cannot all be expected to use an API, so useful interfaces, interactions and application scenarios matter. She reports reach of more than 300 million people across around 200 countries, and over a million companies. The figures describe the scale she attributes to MiniMax’s products; the interview does not define activity levels or the measurement period.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Open releases as a source of improvements
Wolf raises the tension between open releases and the need for revenue. Song responds with the research team’s intention to keep open-sourcing models and the value it receives from the community. Performance feedback and pull requests inform later versions. Her answer explains a research benefit of openness, while leaving the revenue model and the relationship between app-specific models and public releases unresolved.
Asked what feedback would help most, Song welcomes reports of failures, especially in multimodal behavior. She acknowledges flaws in the current combination of capabilities and describes improvement as ongoing. Feature requests also matter: she gives thinking effort as an example of something users have requested. This is an invitation to influence future models, rather than a commitment to a particular implementation or delivery date.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Understanding presentations and videos before acting
Wolf and Song agree that multimodality in coding agents remains underexplored. Song’s examples show the broader opportunity: reading a PowerPoint presentation, interpreting a report that is not well structured, or understanding a long video and then using tools. The proposed workflow connects perception to action, with the content supplying information the agent needs for its subsequent task.
Wolf makes the idea concrete by asking whether an agent could watch his YouTube tutorial and learn how to use the coding tools he describes. Song says she thinks so. It is a plausible application in the discussion, but the exchange does not establish that a particular tutorial was tested or that the resulting tool use was successful.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Research harnesses and the move toward multiple agents
Song says MiniMax builds its own research harnesses and uses them to automate many workflows. She mentions kernel optimization and building data as examples of work models can support. The harness combines model capabilities with a workflow so they can assist daily research and accelerate iteration. She reports that M3 is already helping develop its next version, though the exchange does not establish how much of that development is automated or quantify the speedup.
For the closing question, Song identifies multi-agent systems and model routing as the direction she finds most exciting. She sees them as ways to support more complex tasks and reveal what models can and cannot do. This puts application design alongside model development: combining and routing models can expand what a system accomplishes while exposing capability limits. She offers a direction rather than a specific coordination architecture, and the interview ends with thanks and applause.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> Joining us on stage is the co-founder
- 0:14
and chief science officer at Hugging
- 0:16
Face, Thomas Wolf.
- 0:20
>> [music]
- 0:26
[music]
- 0:32
>> Hello everyone.
- 0:33
Hello Olive, nice to have you on stage.
- 0:36
>> Hi, nice to meet you. Thanks for having
- 0:37
me, yeah.
- 0:38
>> So I think you're on for a treat today
- 0:40
because you just saw a GLM
- 0:43
uh which is current number
- 0:46
two on the artificial intelligence
- 0:48
leaderboard. I take Fable out because
- 0:50
nobody can use it. And now we have
- 0:52
number four. So you basically you will
- 0:53
have all the top models, at least the
- 0:55
top open source model in a row. And
- 0:57
we're very lucky to have
- 0:59
Olive who has a
- 1:01
pretty amazing path in life.
- 1:04
Uh so she came to the US, Pennsylvania.
- 1:06
She was studying, doing PhD at uh
- 1:09
NYU uh in the lab of Jan LeCun
- 1:13
working on J Pa, but we decided we won't
- 1:15
talk about J Pa today, right?
- 1:17
Something for another day. Um and then
- 1:20
instead of joining Hugging Face, which
- 1:21
was in New York also at that time she
- 1:24
decided to go join MiniMax.
- 1:28
So for those who who maybe don't know
- 1:30
all the all the neo labs around the
- 1:33
world and you're you're forgiven because
- 1:35
I think there's like 64 neo labs right
- 1:37
now. MiniMax is one of the top of what
- 1:40
we call the AI dragons in China.
- 1:43
So these are the new there's there's
- 1:45
Deep Seek which is very well known now,
- 1:47
Moonshot who does Kimi Z and GLM that
- 1:50
you just saw and now we have a MiniMax.
- 1:53
They're all extremely good, extremely a
- 1:55
team
- 1:56
uh fighting for the first spot.
- 1:59
Uh so the the the latest uh release of
- 2:01
MiniMax was M3 uh just earlier earlier
- 2:05
in June, which was the the top model at
- 2:07
the time, top open-source model.
- 2:09
Uh very impressive. There's a lot of
- 2:11
very interesting things about this
- 2:12
model, so we'll quickly dive in them.
- 2:15
And then talk a little bit about uh
- 2:17
what's what's what's specific about
- 2:18
MiniMax, what's what's great there.
- 2:22
So uh maybe Olive to to start a little
- 2:24
bit. Can you Can you give us, you know,
- 2:26
a
- 2:27
a little bit of your your view of of M3,
- 2:30
what you like about this model, how was
- 2:32
the release?
- 2:33
>> Mhm.
- 2:34
Yeah, M3 we released M3 earlier this
- 2:36
month, and it is a smaller model with
- 2:40
400 around 400 billion total parameters
- 2:43
and 20 billion activated. Um
- 2:46
but it is very capable in terms of both
- 2:48
coding performances, and also it
- 2:51
understands vision. So um
- 2:54
that's uh what
- 2:56
open-source models don't usually have.
- 2:58
It's that they can
- 3:00
the model can only deal with coding, but
- 3:03
it can also understand videos, um
- 3:05
images, and it has a super uh long
- 3:07
context of 1 million.
- 3:09
Um with our new architecture called MSA,
- 3:12
MiniMax Sparse Attention. So we
- 3:15
we really put these three things
- 3:17
together uh because we know that they
- 3:21
are they will be very important in
- 3:23
future AI applications. Coding
- 3:25
capabilities, agentic capabilities,
- 3:28
longer context, and multimodal
- 3:30
understanding. Um yeah, I think that
- 3:33
would be very interesting about the
- 3:34
model.
- 3:35
>> Yeah, so so there's a lot to unpack
- 3:37
unpack in this model, and it's um
- 3:39
it's it's still, I think, the only top
- 3:41
five model open-source model that is
- 3:43
actually multimodal, so we need to talk
- 3:45
about that. But maybe first about the
- 3:47
long context, because that was also the
- 3:48
first one that really had this real 1
- 3:51
million token long context is actually
- 3:53
functional and you guys had also the the
- 3:56
Minimax pass attention
- 3:59
which is this one technique to to make
- 4:00
that efficient that you also published
- 4:03
and and share extensively. So can you
- 4:05
can you talk a little bit about this?
- 4:07
Maybe how the project went from
- 4:09
from the attention how to make this long
- 4:11
context.
- 4:12
>> Yeah, I would say the story about long
- 4:14
context went back to even Minimax M1 and
- 4:17
Minimax 01 where the model was actually
- 4:21
was able to perform tasks of 10 million
- 4:25
token context.
- 4:26
>> 10 million?
- 4:26
>> 10 million, yes.
- 4:28
But then it was not an agentic model,
- 4:31
right? It was just 10 for example, you
- 4:33
dump in a book, it would be able to give
- 4:35
reviews on it and stuff like that. So
- 4:38
what we realized was that, you know,
- 4:41
longer context actually unlocks a lot of
- 4:43
capabilities especially when interacting
- 4:47
with users and now when, you know, the
- 4:49
agent is interacting with the whole
- 4:51
environment and getting all the tool
- 4:53
responses,
- 4:54
getting multi rounds
- 4:57
the like shorter context wouldn't be
- 5:00
enough to perform the complex task. So
- 5:04
for this version we said, "Oh, we have
- 5:07
to have our longer context backs."
- 5:10
So what we pursued was with our Minimax
- 5:14
sparse attention.
- 5:15
Um which you know, was the architecture
- 5:20
that was scalable
- 5:22
and had a simple design. So I would say
- 5:26
from a higher level, right? It has an
- 5:28
index branch
- 5:30
that, you know, selects
- 5:33
on a higher level what is what matters
- 5:35
more in the context and then we have a
- 5:38
sparse attention branch that calculates
- 5:40
performs the calculation on the selected
- 5:42
blocks to actually performs the task.
- 5:45
Um and so
- 5:47
yeah, like that we really designed um an
- 5:51
elegant architecture so that we can
- 5:53
scale the length and then scale the
- 5:55
model size in the future with that.
- 5:57
>> That's beautiful. I like how for those
- 5:59
who've been in the field for quite some
- 6:01
time, we we had a lot of work on
- 6:02
attention, right? This N square and
- 6:04
there was a lot of linear attention.
- 6:06
>> Yeah.
- 6:06
>> And then some that somehow all of these
- 6:08
disappeared at some point when flash
- 6:10
attention came around. We discovered we
- 6:11
just needed more efficient camera. Now,
- 6:13
I like how we come back to thinking, you
- 6:15
know,
- 6:16
first principle, what is attention? How
- 6:18
can we make that more efficient? So, 1
- 6:21
million token is crazy, right? GPT-2 was
- 6:24
1,024 and and everyone was like, "Oh,
- 6:26
that's really big. We we we never need
- 6:28
more."
- 6:29
Where do you see this coming? Like,
- 6:30
going in the future? Like, Jeff Dean was
- 6:32
pitching me the other day a trillion
- 6:33
token attention. You think we should go
- 6:36
to
- 6:36
a trillion token attention?
- 6:38
>> So, that's definitely something we can
- 6:39
explore towards, right? Ultra lengths of
- 6:43
the context. Definitely, it's something
- 6:46
that's very exciting to explore with.
- 6:47
And something that architecture design
- 6:49
along with hardware um
- 6:52
would require a lot of research onto
- 6:53
that. Yeah.
- 6:54
>> You think there's still a lot of
- 6:55
low-hanging fruits? So, typically today
- 6:58
we saw Open AI really reducing I mean,
- 7:00
we don't know how as of firm, but like
- 7:02
reducing their their inference bill by
- 7:04
half, by probably having some more
- 7:06
efficient processing around tensions of
- 7:08
any type. You think there's still a lot
- 7:10
of low-hanging fruit that can be getting
- 7:12
um
- 7:13
how we can process that. So, so one one
- 7:15
thing we're still very interesting about
- 7:16
M3 is how cheap it is in particular
- 7:18
because of this part attention and part
- 7:20
because of its small one, but it's also
- 7:22
very efficient.
- 7:23
>> Right.
- 7:23
>> You think we can go even way further?
- 7:25
And maybe how did you guys invented uh
- 7:29
Min Max Fast Attention? Was it an agent
- 7:31
coming up with the idea? Was it a human
- 7:32
still coming up with the idea? Tell us a
- 7:34
little bit about
- 7:35
>> Yeah. Um, so we do think there's still a
- 7:39
lot of work that can get into
- 7:41
architecture and inference optimization
- 7:43
so that the model can be more efficient,
- 7:46
especially if there are tasks that are
- 7:48
very task sensitive but require very
- 7:51
strong capabilities, right? And for that
- 7:53
those kind of tasks we really want model
- 7:55
to be efficient.
- 7:57
Um and who came up with this part of
- 7:59
actually I think an intern from our team
- 8:02
worked on that. That yeah, an intern.
- 8:05
Uh that
- 8:06
doesn't usually happen in a lot of labs
- 8:09
because I think in some labs interns
- 8:11
don't have access to
- 8:13
the data, the work, and stuff. Uh but
- 8:16
yeah, we are open to anyone who would
- 8:19
like to contribute to our models. So, um
- 8:21
the architecture was actually designed
- 8:23
by an intern.
- 8:24
>> It's very good. Still some work for
- 8:25
interns here. Good good news. Um that's
- 8:29
also a good segue to also how Min Max is
- 8:31
working internally. So, so we were
- 8:33
discussing before coming on stage I was
- 8:35
saying everyone can propose a project.
- 8:38
Can you tell us a little bit about how
- 8:39
you are organized, how you do your
- 8:41
research?
- 8:42
>> Mhm. Mhm. I think that is very different
- 8:44
from uh
- 8:46
even in school or even in earlier
- 8:50
you know, the earlier tech companies is
- 8:52
pretty pretty different. It's that um
- 8:55
what we
- 8:56
what we make sure is that we have good
- 8:59
foundation and good um infrastructure so
- 9:02
that anyone can play with the model and
- 9:05
can think of what they can improve with
- 9:07
the model.
- 9:08
And then after model releases when they
- 9:10
are free, right? They can play with the
- 9:13
model. They can think of their own
- 9:15
evaluations. They can find their own
- 9:17
weaknesses and propose a thing that they
- 9:19
want to improve on the model. And then
- 9:22
other people who are interested in that
- 9:24
would, you know, propose to join the
- 9:25
project and they will work on for a
- 9:27
couple of weeks or even a couple of
- 9:29
months. And when they work out, the
- 9:31
final thing is shipped to our model. It
- 9:34
it is, you you we use that in our final
- 9:36
training and it's shipped out to the
- 9:38
audience.
- 9:39
>> Interesting. So, you can have people
- 9:40
working for a really long time on
- 9:42
project. When you say a couple of
- 9:43
months, it can be like really deep
- 9:45
exploration of
- 9:46
>> Yes.
- 9:46
Yes.
- 9:48
I would say, for example, architecture
- 9:50
might require longer time of
- 9:51
investigation, research, experiments,
- 9:54
even redoing the evaluations for
- 9:56
pre-training.
- 9:58
Yes, so it might require longer time.
- 10:00
>> Very nice, yeah. And I know you're also
- 10:02
very big on evaluation. I agree. We
- 10:03
could talk about that. I think what One
- 10:05
thing probably related to that is this
- 10:07
unique specificity that M3 and your team
- 10:09
has um around multimodality. So, not
- 10:12
just text, but the similar can also
- 10:14
understand image and video. And as I
- 10:17
understand, but but please explain
- 10:18
explain better,
- 10:20
when you read the model card on Hugging
- 10:21
Face, it say the model was trained from
- 10:23
the first step as a multimodal, not just
- 10:25
have a like user one as after solved,
- 10:27
right? Can you tell us a little bit more
- 10:29
about that and why you think it's
- 10:31
important and
- 10:32
and and why starting from the first step
- 10:34
on multimodal training and not just just
- 10:36
training this.
- 10:37
>> Um so, we call it native multimodality.
- 10:40
Um and so, it is somehow typical for
- 10:45
model labs to train the multimodal,
- 10:48
let's say, vision understanding
- 10:49
capabilities after the text pre-training
- 10:52
is done. Um they put adapters and then
- 10:55
train that part.
- 10:57
But what we found out was that that
- 10:59
would actually harm the text
- 11:01
performance.
- 11:03
And the vision vision understanding
- 11:05
performance wouldn't converge that well
- 11:07
because the model is kind of converges
- 11:09
towards the model the text
- 11:10
understanding. Um and it's just not the
- 11:14
most optimal. And also not the most
- 11:17
scalable, if you think about it. We want
- 11:18
to scale the data, right?
- 11:20
And also, we can also some labs um train
- 11:24
this capability from halfway through the
- 11:27
pre-training. For example, continued
- 11:29
pre-training. But what we found that
- 11:31
this would be very, you know, uh recipe
- 11:35
sensitive.
- 11:36
It is different for the recipe would be
- 11:38
different for different architectures,
- 11:40
different, you know, data mixtures,
- 11:42
different learning rates.
- 11:44
It's hard to control, hard to, you know,
- 11:46
scale to
- 11:49
you can't really scale your experiment
- 11:52
results and conclusions to a larger
- 11:54
model.
- 11:55
And so
- 11:57
you know, what we thought was why not
- 11:59
just training from the very first step?
- 12:01
That comes to the most natural. We know
- 12:04
that a lot of labs run into problems
- 12:06
doing that.
- 12:08
The model would collapse after a couple
- 12:10
of steps of training, you know, both
- 12:13
text and vision understanding, but we
- 12:16
managed to solve that problem.
- 12:18
We did a lot of work on
- 12:21
VIT and we did a lot of work on the data
- 12:24
that we actually training. For example,
- 12:27
we do interleave the data, what we call
- 12:29
interleave the data.
- 12:31
It's actually natural data, but we keep
- 12:33
the
- 12:35
images and videos in instead of masking
- 12:37
it out and we do some pretty
- 12:40
good cleaning and masking on the data
- 12:43
and we do very good reward modeling so
- 12:45
that we train it from the first step and
- 12:47
scales up
- 12:50
a lot. Yeah, it does does not collapse.
- 12:53
>> That's really impressive. Impressive.
- 12:55
Should we Should we expect much larger
- 12:57
model in the future? So this one is
- 12:58
still fairly small, right?
- 13:00
>> It's It's 428 billion parameters, 23
- 13:03
active billion.
- 13:05
Well, do you think you will go past the
- 13:07
trillion?
- 13:09
>> Definitely. Yeah, definitely in the
- 13:11
future.
- 13:13
There are many tasks that wouldn't be
- 13:15
able to the more model
- 13:17
wouldn't be able to perform very good at
- 13:19
with smaller parameters. We are
- 13:22
definitely going more ambitious than
- 13:24
this.
- 13:25
>> It's great. Looking forward. Um another
- 13:28
interesting thing I
- 13:30
I always been find fascinating about Min
- 13:32
Max is how how you also have this whole
- 13:34
range of of apps and product, right? So,
- 13:36
I remember already So, so Min Max
- 13:39
started to open source things on the on
- 13:41
the hugging face platform in in January
- 13:43
last year. So, that 18 month ago and and
- 13:45
we were chatting a little bit about the
- 13:47
team to understand what you were doing
- 13:48
and I remember you So, you were already
- 13:50
having a huge usage on some of these of
- 13:53
some of these apps. Um can you tell us a
- 13:56
little bit how how this started, right?
- 13:58
So, was it basically you had a lot of
- 14:00
apps and then you thought we have we
- 14:01
have all this data, why not training a
- 14:03
model and then they build up research
- 14:05
team. How is how is the story there?
- 14:07
>> Um our our story is modeled from the
- 14:11
first day. So, um I believe that
- 14:14
multi-modality model a model that can
- 14:16
understand all visions and outputs all
- 14:19
modalities was the first thing that our
- 14:22
um CEO planned on the first day even
- 14:24
before the company even started. So,
- 14:26
that was the dream of AGI. I think that
- 14:28
was very very early even before ChatGPT
- 14:30
came out.
- 14:31
>> Wow.
- 14:31
>> Um yeah.
- 14:33
And then apps were something that comes
- 14:35
along because you have some model
- 14:37
capabilities you want people to
- 14:38
experience it well.
- 14:40
Not many people can use it with API,
- 14:43
right? We can't expect everyone to
- 14:44
experience with API. So, we need good um
- 14:48
user interaction,
- 14:50
you know, interfaces, good apps, good
- 14:52
scenarios that people can
- 14:55
can you know, experience experience
- 14:57
model with. I think actually those apps
- 15:00
covered more than 300 million people
- 15:03
around 200 countries globally.
- 15:07
And I think over a million companies as
- 15:09
well.
- 15:10
>> Yeah, this was a mind-blowing when I
- 15:11
heard about the the size and we we don't
- 15:13
often realize the size of of of this
- 15:16
type of usage already. And and that I
- 15:18
kind of brings me to the question around
- 15:20
um
- 15:22
open source business model, and and all
- 15:24
of that, which is the
- 15:26
always existing question, which is right
- 15:28
now it's nice to open source model, but
- 15:30
you you also need to have some revenue
- 15:32
stream, right?
- 15:34
So, I guess M3 is something you you
- 15:36
decided, for instance, to be for free,
- 15:38
and I'm I think it's it's it's great for
- 15:40
the world. Um how do you see this? Do
- 15:43
you also have some specific models you
- 15:44
use for the app? Do you think about Do
- 15:47
you think in the future you'll keep It's
- 15:49
probably hard to say for sure, but do
- 15:50
you think you'll keep open sourcing
- 15:52
models? How is the culture around open
- 15:54
sourcing right now?
- 15:55
>> Personally, and also for the model
- 15:57
research team, we always hope to open
- 16:00
source the models. Um that is our plan.
- 16:03
Because we
- 16:04
really see how the open source community
- 16:06
together can help the model build
- 16:09
better. For example, we receive a lot of
- 16:12
um feedbacks on the model performance
- 16:14
from the great community, and we receive
- 16:16
PRs on uh wh- whatever we open source.
- 16:20
And those are very very valuable and
- 16:22
come comes to our later versions. So,
- 16:24
definitely open sourcing is great.
- 16:27
>> That's great. And actually, do you have
- 16:29
some ask for the audience, people who
- 16:31
are using uh M3 or MiniMax? Is there
- 16:33
something you would love them to send
- 16:36
back to you as feedback? Do you Do you,
- 16:38
for instance, do you read when people
- 16:40
try to modify the models or play around,
- 16:42
you know, tweaks? Or what is the best
- 16:45
thing you you think you can take from
- 16:47
the community uh for future models, for
- 16:50
instance?
- 16:50
>> Mhm.
- 16:51
I would say whatever um issues that
- 16:54
people are running into, especially with
- 16:56
multimodality, right? This is the first
- 16:58
time that we're combining it together.
- 17:00
We are definitely going more ambitious
- 17:02
on that in the future. It might have
- 17:04
some flaws right now, but we are
- 17:05
improving on that. So, whatever that's
- 17:08
uh feedback that model is not good doing
- 17:10
that great, we will definitely improve
- 17:12
that in future versions. And also,
- 17:14
whatever features that uh people want.
- 17:17
Say,
- 17:18
you know, for example,
- 17:20
thinking effort. Right? Some people ask
- 17:22
for that. Um
- 17:24
like everyone can ask, and we will try
- 17:26
to accomplish that in the future models.
- 17:29
>> Yeah. Do you see a lot of usage right
- 17:30
now already in multimodality in terms of
- 17:33
coding agents? I feel like it's it's a
- 17:35
little bit un- underexplored.
- 17:38
>> It is It is, but um it can actually
- 17:42
unlocks a lot of capabilities and a lot
- 17:45
of
- 17:45
uh agent applications.
- 17:47
Say that for example, you want a model
- 17:49
to read a PP- PowerPoint, uh or to read
- 17:52
some report that is not very structured.
- 17:55
Um and you want it to understand a very
- 17:58
long video, say that you dump in a
- 18:00
long-playing video, and then you want
- 18:02
the model to act uh using some tools uh
- 18:05
after understanding it. And it unlocks a
- 18:08
wide variety of um agent use cases.
- 18:12
>> So, like the agent could finally watch
- 18:14
my YouTube tutorial and understand how
- 18:16
to use my coding tools, how I described
- 18:18
it? Is it something like that?
- 18:20
>> Uh
- 18:21
>> Could the agent finally watch YouTube
- 18:23
tutorials and understand things from
- 18:24
them?
- 18:25
>> Yeah, yeah, yeah.
- 18:26
I think so.
- 18:27
>> Do you use a lot of uh agent coding
- 18:29
tools internally? Is it like I mean,
- 18:31
coding for sure, but like is it also
- 18:33
already in terms of research? Is it
- 18:35
automated part or not? How How does this
- 18:38
uh
- 18:38
>> Yes.
- 18:39
>> work?
- 18:39
>> Um we have our own research harnesses.
- 18:42
Um we build our own research harnesses
- 18:44
that automate our workflows. I would say
- 18:47
a lot of our workflows are automated.
- 18:49
You can see how um the latest frontier
- 18:52
models all pursues capability less
- 18:56
kernel optimization, right? Like let the
- 18:58
model post string other models.
- 19:01
Um let the model build data.
- 19:03
Auto data, stuff like that. Um you can
- 19:06
see how more and more models are capable
- 19:09
of doing those, including M3. Actually,
- 19:11
we were very good at those cases, longer
- 19:14
horizons and coronal organizations.
- 19:16
Um and so we can use that model
- 19:19
capability, harness it together, and
- 19:22
help with our um daily routine, and make
- 19:25
our iterations even faster.
- 19:27
>> Is M3 building M4 already?
- 19:29
>> Um building M3.1.
- 19:32
>> M3.1, okay.
- 19:33
>> [laughter]
- 19:34
>> Let's hit the gym.
- 19:34
>> Already.
- 19:36
>> Um I I would love to finish on what what
- 19:38
you find exciting in the coming month,
- 19:41
what what do you think it it can be Asia
- 19:43
in terms of feature or or things you
- 19:45
want to see happening in in AI or more
- 19:48
generally in terms of whatever whatever
- 19:51
really is top of your mind I would say
- 19:52
and it's going to
- 19:54
happen.
- 19:55
>> A lot of things are very exciting. Um
- 19:58
but what I recently find the most
- 20:00
exciting would be a multi-agents that I
- 20:03
think a lot of AI applications are
- 20:05
using, model routing, multi-agents um
- 20:09
that allows even more capabilities, even
- 20:12
more complex tasks, and also it tells us
- 20:16
what the models are capable and not
- 20:18
capable of, and you can, you know,
- 20:21
do a lot of things with that. It's
- 20:23
pretty exciting.
- 20:24
>> Thanks a lot, Alif. Pleasure to have
- 20:26
you.
- 20:27
>> Thanks for having me.
- 20:28
>> Thanks, everyone.
- 20:30
>> [applause]