How MiniMax M3 Was Built: Sparse Attention and Native Multimodality — Olive Song
Read the talk
How MiniMax M3 Was Built: Sparse Attention and Native Multimodality
Olive Song explains how MiniMax paired GPU-aware sparse attention with text-and-vision training from step zero, then improved visual representations through cross-frame attention.
From a talk by Olive Song
At a glance
Ideas worth remembering
MSA trains a lightweight index to match main-branch attention scores, then uses its selections to restrict the keys and values the sparse branch reads.
Head-specific selection preserves GQA diversity, while block retrieval reduces top K overhead and improves GPU memory access. Efficient kernels complete the design.
In MiniMax's experiments, introducing vision from step zero almost preserved text capability under equal text-token budgets; later introduction caused degradation or sensitivity to training settings.
3D attention lets video patches interact across frames inside the ViT. Four frames gave the team's best performance–efficiency tradeoff, with reported benefits extending to some image benchmarks.
A million-token context needs an efficient attention design
MiniMax M3 brings together coding and agentic capabilities, a one million token context window, and visual understanding learned alongside language generation from the beginning of training. Olive Song, who introduces herself as a reinforcement learning researcher at MiniMax, connects those choices to applications: task decomposition for coding agents, full codebases in context, and computer-use agents that need to understand what they see.
The long context window makes attention efficiency central. MiniMax Sparse Attention (MSA) is reported to deliver nine times faster prefill—the initial processing of the input—and fifteen times faster decoding than full attention. Those are the comparison figures presented in the talk; their hardware and workload conditions are not specified, so they should not be treated as speedups for every deployment. The useful engineering question is how the model reduces attention work while preserving access to relevant context.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A lightweight index chooses what the main attention branch reads
MSA separates finding relevant context from performing the main attention computation. Its design follows a preference for simpler architectures that scale cleanly. Two branches divide the work:
- Index branch. A lightweight full-attention computation uses four query heads and one key head. A KL loss trains its attention scores to match those of the main branch, teaching the cheaper computation to identify where the main branch would place attention.
- Sparse branch. The main branch attends to a selected subset of context. The index branch chooses blocks, and those blocks determine which keys and values the sparse branch needs to read.
This division explains where the savings come from. The index branch still examines the context through a cheaper computation, but the main branch performs attention only over what the index selects. The index therefore has a consequential job: its scores must be useful enough to guide the more expensive branch. Matching the main branch's attention scores gives that selection process a training target.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The same amount of data can produce different GPU work
An attention algorithm also has to fit the architecture and its memory access pattern. DeepSeek's DSA provides the comparison: Song describes it as a strong design optimized for multi-query attention (MQA) and a large head dimension. Moving that design directly into grouped-query attention (GQA), which has multiple key/value heads, introduces both a modeling problem and a systems problem.
The modeling problem appears when index outputs are merged into one selection. Multiple key/value heads then attend to exactly the same selected tokens. GQA has several groups that could draw on different parts of the context, but a shared selection removes that diversity at the retrieval stage.
The memory-layout example makes the systems problem concrete. Start with one key/value head whose dimension is 512: the illustrative layout is 1 × 512, which can be read relatively contiguously. Split the same total dimension across four heads and the layout becomes 4 × 128. In the access pattern described here, the GPU now makes multiple independent reads across heads. The total dimension has stayed the same, yet memory access overhead rises and the amount of computation per memory access falls. Equal arithmetic size does not imply equal execution cost.
Selection adds another cost. Token-level retrieval requires finding the top K scores among a large number of candidates. Those comparisons have data dependencies, which make the operation difficult to parallelize fully on a GPU. Computing and merging the selection can therefore consume a meaningful part of the work that sparse attention was supposed to save.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Preserve head diversity and retrieve blocks
MSA changes both the selection rule and the unit of retrieval:
- Keep selections separate across key/value heads. Removing the aggregation of index outputs allows different heads to attend to different token sets. The retrieval process preserves the diversity that motivated using GQA.
- Select blocks instead of individual tokens. Block-level retrieval reduces the number of candidates entering top K and lowers selection and merging overhead. It also improves memory access efficiency and raises the compute-to-memory-access ratio.
- Build kernels for training and inference. Dedicated kernels make the design practical in both phases. The talk leaves their implementation details outside its scope.
Return to the four-head example. A merged index selection would send all four heads to the same subset of context, even though the heads are separate. With MSA, those heads can choose different subsets; block retrieval then supplies their keys and values with less selection overhead and more efficient memory access. The change addresses two separate losses: shared selection restricts what the groups can examine, while scattered token retrieval makes that examination expensive.
Where does the index's output become a reduction in main-branch work? The flow below follows that handoff. Head-specific selection retains different routes into the context; block retrieval determines the keys and values the sparse branch actually uses. The expensive attention operation sits after the selection.
The reported result is an efficient extension of M3 training to one million tokens of context, without a significant performance drop from block retrieval in the team's experiments. The tradeoff is a coarser retrieval unit in exchange for fewer candidates and better hardware access. Its success depends on the selection algorithm and the kernels working together.
Four query heads and one key head; trained with KL loss to match main-branch attention scores.
A lightweight index selects blocks separately across key/value heads; the sparse branch reads and attends to the selected context.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
When vision enters training changes what it disrupts
The next problem is preserving text capability while introducing vision. A model intended for coding and agentic work cannot gain visual understanding at the expense of its core language performance. MiniMax explored three ways to introduce multimodal data during large-scale training.
- After pretraining decay. In this continued-pretraining (CPT) approach, the text model has already largely converged. Adding multimodal data tended to interfere with learned text capabilities and degrade performance.
- Before pretraining decay, after some text training. Introducing vision earlier still produced results that were highly sensitive to the data mixture and hyperparameters. A smaller learning rate could reduce the drop in one setting without generalizing to another run. Large-scale pretraining also constrained how freely those settings could be changed for multimodal training.
- From step zero. Native multimodal training introduced text and vision together from scratch. Under the same text-token budgets, the team found that this approach almost did not hurt text capability.
The timing changes the learning problem. In the first two approaches, visual understanding enters after the model has already learned from text alone. Training must accommodate the new modality within that existing structure, and the experiments exposed interference or sensitivity. Starting both modalities together lets the model develop their relationship from the beginning. The equal text-token-budget comparison matters because it tests preservation of text capability under that stated constraint.
Attention maps offered a view into that relationship. In the CPT-style models described in the talk, text tokens often paid little attention to visual tokens: the visual-token regions were mostly dark. With native multimodal training, the maps showed less separation between text and visual tokens. That pattern suggests more integrated use of the modalities, though attention patterns alone do not establish the quality of visual reasoning. It supports the training goal: make vision participate in the language model's processing while preserving strong text capabilities.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Let video frames interact before they reach the language model
The multimodal architecture follows a straightforward data path. A vision transformer, or ViT, processes an image or video and produces visual tokens. Those tokens replace image placeholders in the text sequence. Visual embeddings and text embeddings then enter the language model together. Training from step zero changes how the model learns to use these inputs; it does not require an unusual input pipeline.
A consequential architecture choice sits inside the ViT. In the standard setting described here, attention stays within each image or video frame. Patches in one frame cannot directly attend to patches in another at this stage. With 3D attention, patches across frames can interact directly, giving the visual encoder a way to model temporal and cross-frame relationships before it produces the tokens passed to the language model.
Where does temporal interaction enter the pipeline? The comparison below places it inside visual encoding. The important difference is which patches can exchange information before visual tokens join the text sequence. Keeping frames separate postpones that interaction; cross-frame attention allows it while visual representations are being formed.
The team found improved visual understanding, with benefits that also transferred to some image benchmarks. That transfer suggests stronger visual representations beyond video-specific tasks. But expanding cross-frame attention indefinitely introduces systems and load-balancing challenges. In MiniMax's exploration, four frames gave the best tradeoff between performance and efficiency. The ending lands on a concrete design choice: let frames interact early, and bound that interaction to a size the training system can handle well.
Patches attend within each individual frame; different frames do not directly interact at this stage.
Both paths pass visual tokens into the language model. 3D attention changes which patches can interact inside the ViT.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
Um, all right. Um, so hi, everyone. I-- My name is Olive. I work at MiniMax on reinforcement learning research. And today, I'm gonna talk a little bit about, um, our new model, MiniMax M3, and how we trained it natively multimodal from step zero, and how we trained it with MSA, which is called, uh, which is MiniMax Sparse Attention. Um, a little bit introduction of our team. Uh, we are one of the few
- 0:42
labs over the world that, uh, works on multi-modalities, which means that we have published our, um, of course, LLM models before. Um, we worked on our speech models, we have our own video generation models, and we have our coding agentic models.
- 1:03
And lately, we just released our MiniMax M3, which is open-weight, uh, with frontier performance. So it is actually the first open-weight model to combine the following three, um, frontier capabilities. The first one is what we call SODA, uh, coding and agentic capabilities. The model is able to do a lot of the, uh, agentic tasks that you imagine. Um, it, it can do automatic task decomposition, it can work in your coding
- 1:34
agents, it can drop in for ninety percent of the coding works that you are currently doing. And the second pillar is that the model has one million context window, um, so which is very long, that enable, like, unlocks pretty much base, uh, all basic user, long user interactions, um, of like it handles full code bases, and we trained it with MSA, MiniMax Sparse Attention, so it is nine times faster prefill and fifteen times
- 2:04
faster decode versus full attention. And the third one is that the n- the model was trained native multimodal from step zero, which means that from step, from scratch, we combined visual understanding with, um, language generation. So the model would understand both capabilities very well from the start and, uh, so that we can build many more applications on that, say, computer use agents, right?
- 2:35
So first, I'll briefly explain what sparse attention does. Our core design principle is that a simpler architecture can always lead to better scaling properties. This is the principle we have followed throughout the entire design. So our goal is to always pursue a cleaner and simpler architecture. The sparse attention consists of two branches. The first one is what we call the index branch, which performs a lightweight full attention
- 3:04
computation. It uses four query heads and one key head, and it is trained with a KL loss to match the, the attention scores of the main branch. The index branch is used to identify the tokens with the highest attention scores. These selected tokens are then attended s- by, by the main branch. In other words, the sparse branch does not attend to all of the tokens, but only to a subset of them.
- 3:35
This subset is selected through the lightweight computation of the index branch.
- 3:43
Then, the sparse branch would use the blocks selected by the index branch to de- to determine which keys and values are needed for attention computation.
- 4:02
Uh. All right. And what I wanna emphasize today is that our design is not purely algorithm, like algorithm-driven. Instead, we took an algorithm, algorithm and infrastructure co-design approach. Um, a good e- good example to start with is DSA. DSA was a very strong design, but it is also highly optimized for DeepSeek's own architecture, especially the combination of MQA and the
- 4:32
large head dimension. However, when we move to a more commonly used GQA architecture, directly applying the same design becomes less straightforward. It introduces several practical issues, es- especially from the infrastructure and implementation perspective. So this motivates our design choice. Instead of simply borrowing an existing sparse attention design, we rethink the architecture under the constraints of GQA and real system efficiency.
- 5:07
So what are the issues, um, if we directly apply the existing sparse attention architecture? First, from the algorithm perspective, GQA has multiple KV heads, which should ideally provide more diversity across different KV groups. However, if we directly reuse the same token selection strategy, multiple KV heads will end up attending to exactly the same selected tokens. This essentially cancels out the diversity that GQA is,
- 5:37
is supposed to provide. Second, from the instr- infrastructure perspective, this design is not very friendly to GPU memory access. GPUs are much more efficient when reading contiguous data. In the MQA setting, we can think of the KV, uh, layouts as one KV head with a large head dimension. For example, one by five hundred and twelve. This is relatively contigu- contiguous and efficient to read. However, in GQA,
- 6:07
the same total dimension may be split across multiple KV heads. For example, it's four by one twenty-eight. These are no longer one single contiguous read. Instead, the GPU needs to perform multiple independent reads across different KV heads. This increases memory access overhead and reduces the com-compute to memory access ratio. And another issue here is the top K computation itself. To select the most important
- 6:37
tokens, the indexer needs to compute top K scores over a large number of token-level candidates. However, top K has inherent data dependencies, so it is difficult to fully be parallel on GPU. As a result, token-level selection introduces non-trivial computation and merging overhead. So how did we address these issues in, um, MSA? From the algor- uh, from the algorithm side, we remove the
- 7:07
aggregation of the index branch outputs. In the original design, the results from all index branch ha- uh, in-index branch heads are merged into a single set of selected tokens. But as we discussed earlier, this would remove the diversity brought by GQA. Instead, we keep the selection different, different across KV heads. Under our sparse attention design, when using GQA, different KV heads can attend to different sets of tokens. This
- 7:37
helps preserve the diversity of GQA rather than forcing all heads to share the same sparse attention. For the infrastructure side, we move from token-level selection to block-level retrieval. Compared with token-level selection, block, block-level retrieval significantly reduces the number of candidates and lowers the overhead of top K computation and result merging. At the same time, block-level retrieval is more hardware friendly,
- 8:07
friendly. It improves memory access efficiency and increases the compute-to-memory access ratio. Importantly, we also show that the design does not introduce a significant performance drop.
- 8:21
Finally, we design efficient kernels for both training and inference. I, I w- I won't go into the kernel details here today, but this is the system-level component that makes the sparse attention design practical and efficient. So with this design, we were able to extend the training of our M3 model to one million context very efficiently.
- 8:56
Um, and as we talked about earlier, the third pillar is the native multimodality. Our model, MiniMax M3, was trained from scratch with multimodal data. What does that mean? Our model was trained with both text and vision understand-standing from step zero, which is very different approach from current open weight models. Our starting point is very clear. At the current stage, a standalone VL model has relatively limited application
- 9:26
scenarios. But co-combining it with, with, you know, text understanding during the training can be challenging. So the key question and challenge there was, how can we introduce multimodal capability without hurting the core text performance? To answer this, we actually explored several training strategies at scale. The first, and most common, approach is to add multimodal training after the
- 9:56
pre-training decay stage, which is essentially a CPT style approach. However, at that point, the model has already largely converged. Adding multimodal data at this stage tends to interfere with the learned text capability and causes degradation. Another approach is to ensure multimodal training before pre-training deca-decay stage. For example, after the model has already seen a number of amount of
- 10:26
text tokens. We also tried this direction but found that the results were highly sensitive to the data mixture and training hyperparameters, such as learning rate, for example. In one setting, a smaller learning rate might reduce the performance drop, but this conclusion does not necessarily generalize to another run. More importantly, during a large-scale pre-training, these parameters cannot be arbitrarily changed just for multimodal
- 10:55
training.
- 10:58
And this led us to our final strategy, native multimodal training from scratch. In this setting, multimodal capability is introduced from the very beginning rather than being added after the text model has already formed its structure. Under the same text token budgets, we find that the pro-- uh, this approach almost does not hurt text capability.
- 11:26
We actually did many fun an-analysis during these expe-experiments. Uh, one of them is this interesting observation, uh, from the attention maps.
- 11:38
You can see from this picture that for CPT style multimodal models, the text tokens often pay very little attention to visual tokens. The visual token regions are mostly dark in the attention map. This suggests that the multimodalities are not fully integrated. While in contrast, with native multimodal training, the attention map looks much more integrated. We do not see here a clear
- 12:07
separation between text tokens and visual tokens. This suggests that the model learns to fuse visual and textual information more naturally from the beginning. So overall, our training strategy is not just about adding vision capability. It is about integrating multimodality into the model while prese- preserving its strongest text capabilities.
- 12:36
And of course, uh, we did many experiments on multimodal architecture choices. Our overall architecture is relatively standard. The image or video is being processed by a ViT, which produces visual tokens. These visual tokens are then inserted into the language model input by re- replacing the original image placeholders in the text sequence. After that, the visual embeddings and text embeddings are fed together
- 13:06
into the language model. One important finding is that applying 3D attention inside the ViT is very helpful for vision understanding. In a standard ViT setting, attention is, is applied within each individual image or frame. For video inputs, this means that visual patches from different frames do not directly interact at the ViT stage. But with 3D attention, patches across different frames can attend to each other
- 13:36
directly. This gives the model stronger temporal and cross-frame visual modeling capability. And we found that this improves visual understanding and the benefit also transfers to some image benchmarks, suggesting that it strength the overall visual representation rather than only improving v- video specific, uh, tasks.
- 14:03
And one practical detail is that we cannot simply extend 3D attention to an unlimited number of frames because that would introduce additional system and load balancing challenges. In our exploration, using four frames gave us the best, best trade-off between performance and efficiency. So overall, with our architecture design, with our training strategies that we explored, uh, throughout the stages, and with our
- 14:33
analysis and ViT explorations, we were able to train the model from scratch, uh, with multimodality understanding. And that will be the end of the presentation today. Thank you.