← All AI Engineer talks

AI Engineer World's Fair 2025

Latent Space Paper Club: AIEWF Special Edition (Test of Time, DeepSeek-R1/V3) — Vibhu Sapra

About this talk

Vibhu Sapra reviews Latent Space Paper Club’s community growth and launches Test of Time Paper Club, a curriculum-oriented series on foundational AI papers and engineering concepts. He then examines DeepSeek-V3 and DeepSeek-R1, discussing reasoning benchmarks, test-time compute, mixture-of-experts models, knowledge distillation, reinforcement learning, GRPO versus PPO, rejection sampling and the broader adoption of open reasoning models.

Chapters

  1. 0:00Paper Club retrospective and Test of Time launch
  2. 10:04DeepSeek benchmarks and model distillation
  3. 16:46Test-time compute and reinforcement-learning reasoning
  4. 26:01Open-model adoption and reinforcement-learning algorithms
  5. 39:52Rejection sampling and final-stage reinforcement learning
  6. 49:00Community participation and closing acknowledgments

Talk transcript

  1. 0:00

    [upbeat music] Okay. So Paper Club year in review.

  2. 0:16

    We've gone over a year, like a year and a half, no missed weeks. We've always done Paper Club, and it's pretty interesting, you know? I don't think any of us expected it to get this far, but every week for the past year and a half, we have always done a Paper Club Wednesday at noon.

  3. 0:31

    We've had a bunch of authors come share their work, people from NVIDIA, Meta, uh, Allen AI, Amazon, Together, Writer. Bunch of authors come, we get direct one-hour sessions with them.

  4. 0:41

    They share their work, we get nice feedback and, you know, some of these have started to do pretty well. On average, we get about 100 people in here every, every session, and just for, you know, um, for, for information, that's, like, on a Wednesday workday at noon, we have 100 people join to discuss a random paper.

  5. 1:00

    Um, and some of the big ones, DeepSeek V3 had 300 live people sitting, listening to me yap about, uh, DeepSeek. Of course, all the other speakers do great. This is all built on volunteers, right?

  6. 1:12

    And along the way, we've made many, many friends, and yeah, Paper Club went much further than we ex- uh, expected. Now, the, the launch. You know, we have to ship something.

  7. 1:21

    We basically have to launch at World's Fair, so we're launching our Test of Time Paper Club. This is gonna be a V2 second Paper Club. So the one that we do now are Wednesday at noon.

  8. 1:31

    This is still sticking around. It will still be the same thing. Every week we'll kind of have a paper, something that's trending, cover it, have an author, have our Q&A session, you know, 30 minutes of paper presentation, whether that's highlights or that's slides, and then we'll continue that with some discussion.

  9. 1:46

    But Test of Time Paper Club is gonna be a little bit different, you know? So it's more curriculum-based. Basically, we take this idea of what's everything that you would need to know to be a good AI engineer?

  10. 1:56

    What are the themes? What are the core papers that don't change? So stuff like, you know, what is attention? How does, uh, sequential text generation work? So what's going on in GPT-2?

  11. 2:06

    How do, like, optimizers work? What about, like, the key inference techniques, right? So stuff like speculative decoding, flash attention, uh, Stable Diffusion, Whisper, the key papers that, you know, are the foundation of what's being built.

  12. 2:19

    We're gonna kind of group those together and session by session we'll go over them. So we're gonna kick off in July and run till December. That leaves us six months.

  13. 2:27

    Six months is about four weeks a month, you know, 24 weeks. Every week we plan to go through two to four pre-presented papers, so you kind of have a bit more structure to it.

  14. 2:36

    There'll be presentations but, you know, in six months we can get through, like, 50 to 100 papers, and we can cover the core concepts. And every week will be kind of different, right?

  15. 2:47

    So, like, one week we might be talking about Whisper and you're interested, right? In one week you can get the fundamentals of speech, uh, speech-to-text, text-to-speech, all that stuff.

  16. 2:55

    Another day it might be something like image generation, so, you know, we'll go over CLIP, Stable Diffusion, how does all this stuff work, extend it out to video. Basically, we'll segment the core topics that we should need to know, and then we'll cover it every week.

  17. 3:08

    Uh, the exciting announcement here is we're also for the first time having a SF section. So we're gonna do, uh, in-person SF Paper Clubs and a remote section. I am gonna mute someone real quick.

  18. 3:23

    So we've got 40 of you. Someone needs to be muted.

  19. 3:28

    We've got crazy echo. Oh, crazy. I'll just mute myself. Easy. I'll mute my own laptop. Okay, continuing on, um, if, if someone needs to interrupt, just, just start in Zoom chat 'cause I don't have my speakers going.

  20. 3:38

    But yeah, um, we'll have a in-person session in San Francisco every week, and of course, we'll keep the remote thing, you know, so we'll keep it going. And original Paper Club will still be its own, you know.

  21. 3:48

    Every week something comes out, we'll cover that, and we won't really deviate too much on the schedule for Test of Time. It's gonna be foundational papers plus some blogs.

  22. 3:56

    Um, so what are the topics, you know? This is still up for us to all decide. So just listing out some of the few ones here, you know, we have foundations of deep learning, something like attention, optimization, ReLU, gradient descent.

  23. 4:11

    What is basic RL? How do these things work? So, you know, we'll have a section on foundations of deep learning. Another one, LLM foundations, right? So pre-LLMs, we have, like, RNNs, uh, LSTMs, bidirectional RNNs, all this stuff.

  24. 4:27

    Then we have, like, the very, very foundational models, right? So, like, BERT from Google, GPT-2, stuff like that. We'll have a day of everything kind of pre-LLM and go over those.

  25. 4:37

    Then after that we'll have, you know, the, the actual, like, generative LLMs. I'm sure that's missing here, but you know. Like Llama 3, DeepSeek, the, the core LLMs that you would expect to hear about, we'll go over those.

  26. 4:48

    We'll have a day of pre-made post-training, right? So what are the scaling law papers? What is Chinchilla? How does distillation work? What are the kind of key papers that you would wanna know for, um, training?

  27. 4:59

    And once again, this will be like a one to two session. So in one week we'll go over three to four papers. We'll have someone present, so come prepared.

  28. 5:08

    We'll go over, you know, what are the fundamentals in scaling laws, distillation? What is Chinchilla scaling? What is, uh, over-training? What is Llama scaling laws? What are small model scaling laws?

  29. 5:18

    So, like, the Phi team really has, you know, for small language models, how should we do proper RL? What is post-training for small models versus big models? Well, we'll have someone cover all these, and that'll be a one to two-week section.

  30. 5:29

    We'll have generative models, so you know, CLIP, Sora, segment anything, diffusion, some of the key agent papers, fine-tuning. We'll go over LoRA, QLoRA, DPO, RL, GRPO. Uh, voice, we'll have Whisper, optimization stage, so you know, speculative decoding, flash attention.

  31. 5:46

    We'll have eval tracks. Rexus, Eugene Yan, you know, he hosts our original Paper Club. He'll fill in good ones there. Much, much more. Um, so yeah. You know, these categories are still up for, uh, debate, so I have a form here later.

  32. 5:59

    Fill it out on stuff that you would wanna add, any papers that you recommend, any topics. It's, it's very straightforward, you know. Uh, also join our Discord. We have the Paper Club channel on Discord, so join that.

  33. 6:11

    Add in topics, anything if you wanna cover, you know. If you wanna volunteer to be a speaker. Over the next, like, week I'll, I'll flesh out a rough schedule, and then from there we'll, we'll kind of take it, you know, start covering it, make sure that every session we have speakers.

  34. 6:26

    We'll figure out logistics, so SF will have a venue. We have a few in mind. Remote it'll be the same, uh, Zoom thing. But yeah, we wanna, we wanna fragment this out, get people, get people interested.

  35. 6:37

    We'll obviously still bring in speakers, you know. So some of these key authors, we know them, so we'll, we'll invite them to speak on these papers and it'll be good.

  36. 6:45

    We'll have discussion sessions. These, these sessions might be a little bit more than an hour since we now have three to four papers. We still wanna just go deep, right?

  37. 6:53

    We don't wanna just, like, do a TLDR of the paper. We wanna do what are the founda- what are the fundamentals and then go a little bit deeper. So, um, stay in touch on Discord.

  38. 7:02

    Very, very basic Google Form. But yeah, now we'll have curriculum-based stuff. Um, you know, like Flo here might talk about music generation, uh, Eugene will talk about state space stuff, so key same members, but just, yeah, second Paper Club.

  39. 7:16

    Now, the, the other part of this is, you know, if you kind of have an area of interest, this will be like your go-to session of, you know, here are the five to 10 papers, here's a presentation on them, here's the discussion, and it will stay live on YouTube, you know?

  40. 7:32

    It's just you can kind of sit in at any time and be caught up to date. You'll, you'll kind of go fundamentals to what you need to know for every session.

  41. 7:40

    And of course, we'll have the same lively discussion. Okay. Test of Time Paper Club, very, very hype. But today is still Paper Club, we can't not do a paper, so you know.

  42. 7:49

    This is just mochi picture 'cause everyone at the conference is having dogs in their slides, so I need them, too. So we can't just have a old Paper Club back here.

  43. 7:57

    We need our OG. So today's paper is gonna be, uh, DeepSeek. So DeepSeek was obviously a popular paper. I know a lot of people haven't had a chance to actually go through the paper, and frankly, I didn't have much time to prep, so you know, I get to reuse my slides.

  44. 8:13

    I'm smart like that. Uh, this is also being recorded for the broader AI engineer conference workshops and speakers, so it's another point, you know, this is why we do Paper Club.

  45. 8:23

    We get these discussions on papers out and then, you know, people can find them later. Like our original DeepSeek paper reading, basically like, you know, that was just last minute, let's highlight the paper, make some basic slides, have some discussion.

  46. 8:37

    But guess what? 300 people joined live. There's over 1,000 views on a YouTube video of us just reading through a paper, so it's a pretty key paper. It's like one of the big open source papers, uh, models that kind of had a big transition, so let's just go over it again.

  47. 8:50

    And of course, there is new stuff. So as of this week, um, or last week, we, we have DeepSeek-R1-Zero-Five-28, so basically the May 28th update. Now, um, you know, there have been rumors that, okay, DeepSeek, uh, V two is coming out, it's coming out, and then people start launching stuff.

  48. 9:08

    Guess what? They didn't do it. They just did the same model. They called it a minor update, but it's actually not that small of an update. So let's, let's dig into what it basically is.

  49. 9:16

    Uh, it's not V three, like revision two. It's, it's the same naming scheme, but it's, it's significantly better actually. So, um, yep, Simon Willison in his keynote mentioned, um, how we don't have good naming for models.

  50. 9:30

    This is basically quite a step up, but they've kept the name. Um, so also plugging Simon's talk, uh, he gave a keynote for the past six months, what were the 30 models.

  51. 9:41

    He launched PelicanBench, it's his own benchmark of how he judges how good foundation models are. Uh, he needs to check the new DeepSeek model. I don't think he did.

  52. 9:49

    But basically, here's what they did. Um, they did better post-training on DeepSeek V3, and now guess what? It got better. So, um, some of this stuff, they put out very little information, but when you dig in, you see that one of the key things is it's much better at reasoning.

  53. 10:04

    So the AIME 2024 score went from 70% to 87.5%. That basically means, you know, this is a good reasoning benchmark from before. Now we have V3 matching the performance of o3 and 2.5 level, um, on math coding and reasoning, which is quite a step up, you know?

  54. 10:23

    It used to be like, okay, we have o1 level intelligence in open source. Yeah, all we needed was a little bit more training, and now we have o3 and 2.5 level intelligence.

  55. 10:33

    And this isn't even like a new model from them. This is just like, let's do a little bit better post-training. Let's do a little bit more and we can get significant, significant, uh, performance increases.

  56. 10:44

    Pretty wild, right? Like 18%, um, improvement on benchmarks, and no one's really talking about this. DeepSeek got a lot better. Uh, one of the quotes that they say is, originally, the original DeepSeek V3, it would take 12,000 tokens to reason through on average on the benchmark.

  57. 11:02

    For an AIME, it would take 12,000 tokens of reasoning. They basically did more RL, now it reasons more. On, on average, it reasons for 25,000 tokens. So they got the model to do double the reasoning.

  58. 11:15

    One thing that we talk a lot about is scaling laws, right? So before we would do optimal scaling for base models, now in Llama we kind of did overfit our training for inference time, right?

  59. 11:28

    Let's really overtrain so we can do fast basic inference. And then now in this world of test time compute where we do, um, we do more inference time compute, we can also scale even more in that dimension.

  60. 11:39

    So original DeepSeek on average would reason for 12,000 tokens. The new model can double that. So in this domain, we've doubled the amount of reasoning it could do, and we have a lot of benchmarks to increase, you know.

  61. 11:52

    18% on AIME, and it's a lot better at coding. So in this, they intentionally wanted to do better JSON output, function calling, and more reasoning. And yeah, they just dropped it like that.

  62. 12:03

    Here's kind of our, you know, benchmark chart. On most Paper Clubs we don't do benchmarks, but yeah, it's, um, you know, it's actually kind of up there with o3 and Gemini 2.5.

  63. 12:13

    So, uh, you know- The darkest color is our new revision of DeepSeek-R1. And yeah, it's, it's actually very good. On most benchmarks, it's, like, significantly better than the original DeepSeek-R1.

  64. 12:26

    Um, and you can see this, you know, humanity, humanity's, uh, last exam. It basically went from not being able to do anything to, "Okay, now this thing can do well."

  65. 12:33

    What they did is now it can basically reason for twice as long. Okay. That's not the only drop. They also launched another distillation. This is kinda the interesting one.

  66. 12:43

    Not, not many people really talked about this at all on Twitter, but, um, if we remember in the original DeepSeek paper, what they did was outside of their, um, original DeepSeek model, they did three distillation models.

  67. 12:56

    They distilled the Qwen series and the Llama series, and they showed how distilling from the big model, distilling on these reasoning traces, we can get really, really good performance on small models.

  68. 13:07

    Well, they did it again. They took Qwen3-8B, and they did another distillation with their new reasoning model. And they show that, you know, the new model, this new basic distillation basically, like, kills the old one.

  69. 13:20

    So, um, basically when you look at their old distill versus their new distill, they s- they get another ten percent performance boost by just doing distillation from a better reasoning model.

  70. 13:30

    So, in a few months' time, they were able to get, uh, DeepSeek to do more chain of thought, more reasoning, use that to distill down to an 8B, and now we have an even better 8B.

  71. 13:43

    So very, very interesting little note, right? Not just do we get a ten percent improvement from the last base model, um, their 8B distillation is actually matching performance of Qwen3's two thirty-five billion/twenty B active thinking model.

  72. 13:58

    Horrible naming, I know, but let's take a second to think about that, right? Uh, uh, Qwen3-8B dense model, so a small 8B model, is matching the performance of the Qwen3-two thirty-five B thinking model, and this is not a native thinking model, right?

  73. 14:16

    This is a base distillation model of an eight billion parameter model. Just on distillation, so logic matching distillation from a big model, we're matching performance of their two hundred thirty-five billion MoE thinking model.

  74. 14:30

    That's pretty wild. They didn't do this to the 32B, the 70B, but yeah, it's pretty crazy, right? Untalked about release that the new 8B does very, very well. Um, how do we see this?

  75. 14:41

    We can see, um, you know, the chain-of-thought improvements distilled down really, really hard. This was one of the key findings in the original paper that, um, a better recipe for training small models is to distill down from big models, and reasoning models make this even more efficient.

  76. 14:58

    And then this is just kind of their follow-up, right? We don't have a paper. We don't have too much on this, but these are benchmarks that show it. And of course, model is open source, open weight, everything.

  77. 15:06

    So those are kind of the, um, those are kind of the overviews. Uh, let's see. Let's see. Now let's go over the actual, um, the original DeepSeek paper. So that's kind of ending where we had, um, the new releases.

  78. 15:23

    So we have two mo- two models to recap. We have a new DeepSeek version. So for May twenty twenty-- uh, for May twenty-eighth, we have a new DeepSeek. It's now at on par with, uh, OpenAI's o3 and Gemini two point five.

  79. 15:37

    Uh, it's significantly better, and it reasons for twice as long. We also took that model, we distilled it down to Qwen3-8B, and we have a much, much better small 8B reasoning model.

  80. 15:48

    Um, this shows that, you know, reasoning models distill down very, very efficiently, and there's still a lot of juice to be squeezed out there. Now, um, from here, for those that haven't seen it, we're gonna take a two-second pause, see if there's any interesting questions.

  81. 16:03

    Um, there's a Hugging Face link. "If these benchmarks are public, won't models be trained to score better on these benchmarks?" Yeah. Benchmarks obviously have their, uh, cons, their flaws, but you know, there's ways to see what models are overfit on them.

  82. 16:15

    How do they do in general performance? In general, the DeepSeek models are actually doing very, very well. So, um, from there, let's go into the original DeepSeek model. So, uh, okay, DeepSeek-V3, hypest paper of the year.

  83. 16:28

    Three hundred people joined us live. Let's do a quick recap. This is basically me using my old slides because I can, but let's talk about what happened in DeepScre- DeepSeek.

  84. 16:37

    So we're gonna kick off with a high-level model overview. So what are the models they release? When they release this, it's not just DeepSeek-V3, right? They also have R1.

  85. 16:46

    Um, what is inference time taini- training? What is test-time compute? What makes reasoning models different? So if you guys don't remember, this was the first test-time scaling open, um, you know, open model, right?

  86. 16:58

    This was the first one that got good. OpenAI released o1. Uh, Claude released Claude Thinking much later. Gemini Thinking came much later. We didn't really understand what was happening.

  87. 17:08

    We thought there was, like, MCTS. Uh, there was a lot of, you know, Monte Carlo tree search. What's going on? There was internal, let's generate chain-of-thought. Let's train it to do chain-of-thought.

  88. 17:18

    But turns out, uh, DeepSeek comes out with this paper. They do a great model, and they're like, "Yo, RL, RL works." So two models were released. Uh, DeepSeek-R1-Zero, basically they, they take a base model.

  89. 17:30

    They do a lot of GRPO RL. They have training templates, reward models. Um, they have this emergence capability, reflection, aha moments. Then they do, um, R1, which is basically a four-stage pipeline.

  90. 17:44

    They have cold start, reasoning RL stage, rejection sampling and SFT, and then, you know, of course, the little, um, RL round two to get it to really, really reason.

  91. 17:54

    From there, we'll talk about performance and evals of how does original R1 do, um, how is DeepSeek. Then the original distillation models, future work, reproductions, and whatnot. We'll kinda skip over the base DeepSeek, um, evals and performance because, you know, we already covered the new one.

  92. 18:11

    But okay, continuing on. So high-level overview. For those that understand... that, uh, don't really follow, you know, we keep hearing this term of test-time scaling. What is test-time scaling?

  93. 18:23

    What are thinking models, and what does this mean? So basically we got to a point where we started to over-train our models. We basically hit a scaling limit on how much we can train models.

  94. 18:34

    Originally, back in the day, we used to do sort of these Chinchilla scaling laws, right? We had a fixed compute budget. We had a fixed amount of dataset. We would design a model around that.

  95. 18:46

    So how many parameters should it be based on how much data we have? Let's fit a model to our data. Let's fit how many GPUs u- GPUs we have, and then let's train it to be kind of Chinchilla optimal.

  96. 18:57

    From there, we kind of realized that, okay, this isn't, this isn't really what we want, and we started really, really scaling up our training. So we had stuff like, you know, Mistral, Llama 1, Llama 2, Llama 3.

  97. 19:09

    We started training these models from billions of tokens to trillions of tokens. So, you know, we had originally like a 1B, 1 trillion token, then Llama 3 was 15 trillion tokens.

  98. 19:19

    Now they're up to like 45 trillion tokens. So what we shifted was instead of training for, uh, you know, model Chinchilla optimal, let's start training for this sort of inference optimal, uh, training regime.

  99. 19:32

    So instead of, you know, thinking about what we have now, let's think about inference time. As we scale this, we want a model to be as densely packed, as smart as possible.

  100. 19:42

    So Llama is like, okay, basically, if we continue training, we don't really see degradation. But the problem in this is it gets very, very expensive, right? As you train more and more, yeah, it's very, very heavily compute extensive and, you know, how many times can you scale this up?

  101. 19:58

    Like, how much data do we have? How much can we really fit in? If we're at 45 trillion parameters for like a 70B, can we scale that up 10X again, you know?

  102. 20:07

    Are we gonna do 450 trillion parameters? What if we wanna 10X it again? We're basically hitting the compute scale for training, right? We can't continually just keep scaling up our train runs because it's no longer like cost efficient, right?

  103. 20:21

    We're spending millions and millions and hundreds of millions of dollars on these train runs. You can only scale so much. So a lot of hypothesis was, you know, there's gonna be a sort of plateau and open models will start to catch up because, you know, we're already all scaling so much.

  104. 20:37

    But that's where, you know, we need to unlock another dimension, which in this case was reasoning or test-time training. So this is where we basically started to do, um, reasoning capabilities without any supervised data, right?

  105. 20:51

    Um, there were approaches to try to do, you know, let's generate a bunch of chain-of-thought reasoning style data. Let's do post-training on it, and yeah, our models do a little better, but this didn't scale.

  106. 21:01

    What we needed was, um, we wanted to do pure RL. Can we do pure RL to do re- uh, reasoning data? So, uh, this is a quote from the paper.

  107. 21:12

    Basically, the DeepSeek team says, "Our goal is to explore the potential of LLMs to d- to develop reasoning capabilities without any supervised data, fo-focusing on their self-evolution through a pure RL process."

  108. 21:25

    So what they're gonna do is they're gonna post-train the DeepSeek-V3-Base with GRPO, which is, you know, basically pure RL, and then as they do this, they start to notice emergence of great reasoning, reflection, aha moments, and they start to match o1.

  109. 21:41

    Now, fast-forward a few months, as we can see today, um, they can basically continue this. They do a little bit more RL. They, they do RL on longer traces.

  110. 21:52

    They can now get the model to not match o1, but match o3 and double its amount of reasoning tokens. Okay. Here's kind of the four-step approach to how they train this R1 model.

  111. 22:02

    They start with this sort of cold start where, you know, they, they jumpstart with some SFT, then they do RL for reasoning, then they have this key step of rejection sampling for generation purposes, you know.

  112. 22:13

    Rejection sampling is something we'll talk about later. Then once again, they do the fourth stage of basic RL polishing. Okay. Um, that's a very high-level, uh, overview of what DeepSeek did to make RL.

  113. 22:25

    So to recap, you know, we needed to shift from next token predictors to scaling in another axis. To do this, we needed to scale on instead of spending the same compute for every token generated, we wanna dynamically generate-- We want to dynamically spend more compute on different queries.

  114. 22:43

    So we train the models now with pure RL to reason through their questions. So instead, you know, they're now trained to do RL with verifiable outputs on a lot of code and math data that can be verified to be, you know, co- uh, whether it compiles, whether it's factually correct, whether the math is logically, um, correct.

  115. 23:04

    And then now we can basically do native RL. Doing this, we notice emergence, reflection, aha moments, and now we basically have another domain in which to scale models. Instead of scaling up a 10X order of magnitude of, you know, instead of 45 trillion tokens, let's train on 450 trillion tokens and then, you know, scale that up and

  116. 23:25

    up. We're kind of at the limit there. We start with a really good base model that's a good general next token predictor. We do RL, and now we can scale in the reasoning domain.

  117. 23:33

    So a few months ago they showed that, you know, they can get o1 level performance, and then fast-forward to now, we have o3 level performance with just more RL.

  118. 23:42

    Tangentially, this was kind of the paper to kick it all off, right? Um, DeepSeek showed that two things: One, you can do RL from base models and we can get reasoning.

  119. 23:54

    Two, distillation really works and it's a much better approach to small models. Following this, we've now had a lot more papers, so, uh, shout out to like the Phi-4 models, right?

  120. 24:03

    Phi-4 showed how effective RL can be. They took Phi-4-Mini and, uh, they made it a reasoning variant. Basically, they took about six thousand samples, and in six thousand samples of RL, they could take a small base model, a small base, uh, next token predictor, and do really, really good reasoning inference on it.

  121. 24:22

    So this was the one to kick it off, and now we have, um, you know, the Qwen models have done this, so there's Qwen thinking models. There are DeepSeek thinking models.

  122. 24:30

    And then there's formulas like, um- The, the, the five models that show how to do this in small models. Okay. Let's, let's continue on from my high-level overview. So what did they release?

  123. 24:41

    They released two models early on. There's R1-Zero, which is a great reasoning model only trained on unraveled chain-of-thought with RL, but it's not a great general model. R1-Zero is when you only do RL on chain-of-thought, you know, it doesn't, it doesn't really emerge to general performance.

  124. 24:59

    So R1 is trained, um, as the second model. It uses outputs of R1-Zero using these four stages of training and then, you know, now we start to do RL back on human tasks, back on chat, and we have a good, good RL...

  125. 25:15

    We have a good, good, um, reasoning model. Second thing they release is their distillations, right? So they take these, uh, thinking traces, and they take models in the same family, and then they distill them down.

  126. 25:26

    So these are not natively trained with RL. They're, um, distillations from their base models in Qwen and Llama f-families, and then they show how this works really well. Okay.

  127. 25:36

    Now, of course, you know, it's 2025. We don't get real papers anymore, so they don't talk about data, how many tokens, where the data comes from. But, you know, they still do share a lot about, about how this stuff works.

  128. 25:49

    Um, models, of course, fully open source, MIT license, no training data, no code. They have a DeepSeek API, which at the time, you know, it was much faster than anyone, it was much cheaper than anyone.

  129. 26:01

    Turns out this was fake news. This API is very unreliable. It's never... It's barely ever up. But, um, you know, now, now a few months later, after this is coming out...

  130. 26:10

    after this has come out, we can see stuff like from OpenRouter, you know. DeepSeek actually takes a significant chunk of the pie. They take about 10% of, uh, API usage that goes through them.

  131. 26:20

    So this model actually sparked a lot of adoption. It reopened up the race of, you know, okay, we're not done scaling. Models are still getting a lot better. And, um, yeah, you know, once again, it's been like a few months, and as of like last week, we have another update to this, and then the Qwen team followed

  132. 26:36

    along. Uh, what was fake news? Fake news was the DeepSeek API. When it launched, it was, uh, 10X faster and 10X cheaper than other inference providers, but turns out it was super unreliable.

  133. 26:50

    No one can make API keys. It was almost always down, but it's okay, you know. It exists if you still wanna try it. It's a really good model. A lot of people use it.

  134. 26:58

    Okay. Um, let's dig deep into these topics. Inference time scaling. So what is inference time scaling? We have o1 versus GPT-4o. Now there's o3, o4, o4 mini, all these things.

  135. 27:10

    But, um, basically what these do is you increase the chain-of-thought reasoning process. You allow models to spend more time thinking before they respond. Now it's very common. We've used a lot of these models, right?

  136. 27:22

    They're starting to become even a little bit more agentic with RL. That's kind of what's progressed since we last covered this. So, um, inste-instead of spending hundreds of millions of dollars pre-training LLMs, uh, that starts to get exponentially more expensive, right?

  137. 27:37

    Instead of changing from hundreds of millions of dollars to pre-train a Llama, we don't wanna spend billions on train runs, right? So we needed another access, so instead we shifted to this paradigm of inference time scaling, where you take really hard questions, you do, um, you know, inference time chain-of-thought RL, and we have a new dimension to...

  138. 27:58

    on which we can scale. So previously, people tried to, um, do process ba... uh, process-based reward modeling, so we would do RL basically on, um... We would try RL.

  139. 28:10

    We would do beam search. We would do MCT... uh, MCTS. We would do all these inference time, you know, let's predict multiple tokens, go down these processes. They were all hacks, but nothing was really close to o1, right?

  140. 28:22

    We would see Twitter demos. We would see, like, some fundraising that even came out of, you know, okay, I have much better performance than Llama 3 because I go down 10 trees of thought and, you know, I'm doing all this stuff on the back end using a bunch of token, tokens and gluing it together.

  141. 28:37

    But this is really, um, not the right approach as we see. What really worked is just native, pure, beautiful, scaled-up RL. So, um,

  142. 28:47

    once again now, this is what, uh, DeepSeek did to make V3. Here's the one slide. If you wanna know what, uh, DeepSeek V3 is, uh, it's open source GPT-4o slash, uh...

  143. 29:00

    Oh, sorry. This is DeepSeek V3. This is the precursor to R1. So, um, 4o quality, 37 billion active parameters. This is the, you know, regular MoE non-reasoning model. This is what they build R1 off of.

  144. 29:12

    So it's a MoE, 671 billion parameters, 37 billion active. Uh, they launched this. It was a good model. It's just really chunky, right? You can't run it on your laptop.

  145. 29:22

    No one can run a 700 B model. Um, they made this whole, you know, "We're better than everyone else," where, um, you know, "We could train this model in $5 million."

  146. 29:31

    And I think that this is actually true, right? Looking back at it, um, a lot of what we see is DeepSeek and [REDACTED:origin] labs were really able to catch up to the United States because of the constraints, right?

  147. 29:43

    We put a lot of, uh, trade restrictions. We couldn't give them GPUs, so they had to get clever and smart with what they got, and basically, you know, they, they, they did very, very strong inference optimization.

  148. 29:54

    They made the most of what they had, right? We could have continued scaling, right? We could have thrown this 14.8 trillion tokens into 150 trillion tokens, but China didn't have GPUs for that, right?

  149. 30:06

    The DeepSeek labs were like, "We realize that we can't scale this in the same dimension," so they had to get creative and think about, "Okay, what if we do RL?"

  150. 30:15

    And that's basically what they did. So, um, V3 was, you know, it was an MoE 37 billion active. They introduced this concept of multi-headed latent attention, 15 trillion tokens.

  151. 30:27

    They did SFT and then, you know, traditional RL, so RLHF to make it a chat model. Um, multi-token prediction. This came out of Meta. They needed it to be sample efficient with their 15 trillion tokens, so multi-token prediction for a little bit more sample effic- um, efficiency.

  152. 30:43

    Trained in FP8. Did some long context, uh, extension. Basically, first trained it at 32K, then they extended this down to 128K. Came out a month ago from, uh, R1.

  153. 30:55

    After that R1 came out, people got mad hyped. Now we have R1-V2 basically. So these are kind of, you know, they're fancy diagrams. We have DeepSeek-V3-Base. We do SFT and RL.

  154. 31:09

    You have a SFT checkpoint RL with, uh, RLHF to get-- Uh, sorry, fine-tune with RL to get DeepSeek-R1. Okay. What is DeepSeek-R1-Zero? R1-Zero is where you don't do any S- any SFT.

  155. 31:24

    You take a pure base model. For those that don't remember, base models are models that come when you do your pre-training. So you potay-- We train models to predict the next token, right?

  156. 31:33

    So these are not models like GPT-4o or o1. These are pure base models. All they do is they predict the next token. So they're kind of completionist models, right?

  157. 31:45

    You can't normally chat with these. All they do is complete your sentence, complete your word. So they take the DeepSeek-V3-Base model. They apply pure RL. They don't do any SFT.

  158. 31:55

    They don't train it as a user-assistant chat model. They don't do any of that. It uses GRPO for RL, which they, you know, they actually introduced quite a while in DeepSeek Math.

  159. 32:04

    The reward is based on both accuracy and format. Responses must be verifiably correct, right? So what are they doing here? They take a base non-chat model. They use GRPO style RL on math and code, and they have a verifiable output, right?

  160. 32:19

    So the models need to be verifiably output. They need their output to be verifiably correct, right? So for math, that's-- you have a correct output to your math question, right?

  161. 32:29

    Is the answer correct or not correct? If it's correct, that's good. We can RL on that. For code, does your code compile? If it compiles, that's good. We can do RL on that.

  162. 32:37

    And, you know, this is Leet- LeetCode-style questions. So LeetCode-style questions, we know the answer. We know if the answer is correct. Uh, then we format the rep-- Uh, basically, you know, in the little minute details, they format the rewards.

  163. 32:50

    They kind of output the thinking between think tags. So we take our base model. We do RL to do a bunch of thinking to get its chain-of-thought thinking process, and from there, we kind of have DeepSeek-R1-Zero.

  164. 33:02

    It's a, it's a reasoning thinking model that's good at outputting thinking and answers, but it hasn't really been trained to be a useful assistant. It's not o1 yet. This is just a good thinking model that can unravel its thought process to generate out answers.

  165. 33:18

    We'll probably skip this guide. This is a GRPO. This is kind of the RL-based algorithm that they use. It's, you know, comparing PPO, GRPO. This is how they do the RL.

  166. 33:30

    We'll kind of skip it. The key, the key things to note, there's no critique model. Uh, you have group-based rewards that are scored. Uh, it has stability updates. There's a, you know, KL divergence.

  167. 33:40

    But yeah, they, they, they do GRPO. We're gonna skip it in this talk for now. So okay, DeepSeek-R1-Zero. We took a base model. We did RL. We now have a thinking model.

  168. 33:52

    How does it perform? Pretty good. Uh, AIME, you know, it passes o1 mini for the time. Uh, math, it passes Math 500. It passes o1 mini. It's not-- It's, like, on par with, but slightly worse than o1.

  169. 34:05

    And then, you know, the charts show that, um, we're able to do this inference-time scaling by doing RL on really hard questions, and it kinda works. It works pretty well.

  170. 34:16

    So, um, yes, chart number goes up. Number goes up even more. The more you train, this thing is starting to stably learn. Um, what else do we learn? So it naturally has the ability to solve these complex tasks by extending test-time compute, right?

  171. 34:31

    Here we know that the original R1-Zero, it ranges from hundreds to thousands of reasoning tokens. Turns out that this was a pretty key, uh, factor. In the update that they released-- that DeepSeek released last week, we scaled from training on, uh, from taking thousands of tokens to now taking-- to doubling it.

  172. 34:49

    So now we take from twelve thousand tokens on average to twenty-four thousand tokens. And once again, we shift from getting o1 level to o3 level performance in the new DeepSeek, um, model.

  173. 35:00

    So other stuff, um, you know, there's this emergence of interesting behaviors as test-time compute increases. So as we're able to increase our test-time training, as we can reason for more and more time, we start to see this emergence of interesting behaviors.

  174. 35:18

    Basically, as models learn to reason for longer and longer, as you get to thousands of steps of reasoning, models are able to start to have these reflection moments. So, you know, the more reasoning you do, models start to learn, "Okay, I'm actually not forced to just output my next token.

  175. 35:34

    I'm not forced to do a hundred tokens of thinking. I'm not forced to do a thousand. I can continue down this thinking path." And we notice that, you know, as they start to reason for lo-longer, models start to do this sort of reflection phase.

  176. 35:47

    Models start to re-revisit, reevaluate their previous steps. They start to go down alternative paths as... And, you know, this kind of arises this spontaneity, um, of-- this spontaneity emergence, right?

  177. 36:00

    So spontaneously, you know, models will be like, "Okay, I tried this, this, this. It's not working. I don't have to answer right now. Let me try this new thing, and guess what?

  178. 36:08

    It works." And then we also have these aha moments. So very core takeaway of the, um, paper. You know, this is kind of what got the DeepSeek authors to realize, "Okay, RL actually works."

  179. 36:20

    Um, the more time that we think for, the, the further along the thinking traces, we start to get these models to do these aha moments. So here's a quote again from the paper: "This moment is not only an aha moment for the model but also for the researchers uh, observing its behavior.

  180. 36:37

    It underscores the power and beauty of reinforcement learning. Rather than explicitly teaching the model on how to solve a problem, we simply provide it with the right incentives, and it autonomously develops advanced problem-solving strategies.

  181. 36:51

    The aha moment serves as a powerful reminder of the potential of reinforcement-- that reinforcement learning unvo-unlocks new levels of intelligence in AI systems." So- Basically, instead of telling models how to reason, what traces, instead of doing SFT on chain-of-thought traces, if we just give it the sparse idea of, you know, reason as much as you want, um,

  182. 37:13

    the more time it takes, we start to notice emergence of, "Aha, I see what it should finally be. I tried these six steps, and this seventh step is working."

  183. 37:22

    And you know, this shows the power of RL, and this has been kind of passed down into other papers. Microsoft has put out really good scaling laws on doing RL on small models.

  184. 37:31

    Uh, Qwen team has done really good thinking models. They have an online talk here, so please, please watch out the Qwen reasoning talk. Uh, a speaker speaks about how they do that.

  185. 37:40

    But okay, back to DeepSeek, back to DeepSeek. Um, this is an example of an aha moment. So question, uh, a basic math question. You know, not, not that basic, actually.

  186. 37:51

    I can't answer this. If A is greater than one, then the sum of this square root of a square root is equal to something. Well, as the model starts to think, you know, it realizes, "Oh, there's a X squared," right?

  187. 38:02

    "If I square this, I can get X squared, then I can isolate this out," right? It's doing reasoning. It's thinking through this math problem step by step. Then it says this in its own generation, "Wait, wait, wait.

  188. 38:13

    That's an aha moment I can flag here. Let's reevaluate this. Oh, okay, instead of, you know, actually solving a square root, I can do this square both sides." And like, you know, very interesting aha moments start to pop up.

  189. 38:27

    So yeah, um, that's kind of overview of what R1 was. Uh, that's kind of an overview of R... what R1-Zero was. R1-Zero is basically where we take a base model, we train it on pure RL for math and code, we do RL on thinking steps, and now we have a reasoning model.

  190. 38:45

    It works, but, um, it's, it's not great. Uh, it doesn't actually have great re-readability. It starts to reason in multiple languages. Whoa, it reasoned in [REDACTED:origin]. Who would have guessed?

  191. 38:56

    So, um, we wanna make it, you know, we can't just skip RLHF, right? We wanna make this thing a good chat model. So next section. Instead of R1-Zero, how do we make DeepSeek-R1?

  192. 39:08

    How do we take DeepSeek-R1 from being just a reasoning model to a reasoning useful assistant that's good at, you know, actually being a chatful assistant? So

  193. 39:19

    key solution, giving it to you straight. Cold start. Instead of taking base model to just RL, take a base model, do some regular SFT. SFT is kind of, you know, here's some prompt answer.

  194. 39:31

    Here's chat assistant. Get it to, you know, do a cold start. Get it to understand you're still a useful assistant. After that, do some RL. Do this core RL on very hard code math.

  195. 39:42

    Get it to start to understand how to think, how to reason. Scale it out until it starts to see these aha moments, until it can do proper reasoning. From there, let's do some rejection sampling.

  196. 39:52

    We wanna take examples that don't work, right? Stuff where it is going wrong, where it has negative behavior. We do rejection sampling on it, and then once again, we do it once again, our last section of stage four training, which is another round of RL.

  197. 40:06

    So going through these. Stage one, cold start with strong SFT prevents the model from getting unstable. So basically, you take a base model. You have a long chain-of-thought style, um, dataset.

  198. 40:18

    So, you know, prompt-answer pair. So user, assistant, think through your process of how to solve this. Give your chain of thought. We take a base model, train it on this chain of thought.

  199. 40:29

    Then from there, um, you know, we do our cold start with strong, strong SFT. They don't just do, uh, synthetic data. They have human annotators. We want better reasonab-- we, uh, we want better readability, right?

  200. 40:41

    So do your thinking in think tags. Um, and then this is on the scale of a couple thousand examples. So base model, thousand examples, a couple thousand examples of SFT.

  201. 40:51

    Then our main RL stage, you know. So same RL process as R1-Zero. Do a lot of RL on verifiable hard questions. So, um, math, LeetCode style coding, stuff that we can verify has the right answer.

  202. 41:06

    Do a lot of RL. And then, you know, stage three, rejection sampling. So generate completions, rank them with a reward model, fine-tune the original model. So this was, uh, standard.

  203. 41:17

    So Llama 3 showed us this concept of rejection sampling, and you know, we take it to DeepSeek. We do rejection sampling on samples that we don't like. We had eight hundred thousand samples that were generated, six hundred thousand reasoning, two hundred thousand general chat.

  204. 41:31

    We rank them, and then we do our rejection sampling training. Then final RL stage for general use. This is very similar to, um, you know, similar to RLHF. You wanna make the model helpful, harmless, and reason good.

  205. 41:45

    So for reasoning, we use, you know, we wanna keep it a good reasoner. Add in that reasoning for hard math questions, code questions. For general chat, capture human preference and nuanced situations.

  206. 41:57

    So, you know, this question should have a very verbose, detailed answer. This is a basic summary. Keep it short, but keep your reasoning there. So final stage is, um, you know, this final stage of RL, and now guess what?

  207. 42:10

    Model is good at being a chat model. Model is good at thinking. Model has emergence of aha behaviors, and, um, the model is just a good chat model. Performance and evals, we're gonna skip this.

  208. 42:21

    Basically, um, when we launched DeepSeek, it was o1 level, but forget that. We launched an update, um, two weeks ago. So new model, DeepSeek-R1, launched May 28th, and instead of being o1 level, the new DeepSeek-R1 model is now reasoning for twice as long.

  209. 42:40

    It's as good as o3 and Gemini 2.5 Pro. Much better at math, much better at coding, much better at reasoning. It now has support for native function calling, JSON outputs, no longer hallucinates as much.

  210. 42:53

    Second model drop, and of course, you know, performance charts, all the regular benchmarks you would expect. Model is performing as good as o3 and Gemini 2.5 now. All that was done, do more RL.

  211. 43:06

    Uh, AIME jumped, you know, 17.5%, doubled the reasoning tokens. Basically, we, we doubled our reasoning access. We now reason for twice as long, so double the reasoning effort. And, um, yeah, now we're o3 and Gemini 2.5 level.

  212. 43:22

    The other drop, we have, uh, a new distillation. We take our new model that reasons for twice as long, we distill this down to Qwen3-8B, and we do a distillation loss on this reasoning, and we get performance matching the Qwen 235 billion reasoning model.

  213. 43:39

    So our dense 8B non-reasoning model that was distilled from our new DeepSeek model is as good as, um, Qwen3-235B, which is an MoE reasoning model, which is pretty crazy, you know?

  214. 43:55

    The, the implications of this show that, you know, long, detailed, good reasoning really has a deep impact. Once again, check out the Microsoft work for good distillation scaling laws on this.

  215. 44:07

    Okay, okay, back to our paper. Instead of looking at our original, um, DeepSeek, that's kind of the performance of where we're at. Distillation, let's talk about these distillation models.

  216. 44:18

    So what we did was distill R1 down into Llama and Qwen models. This is not RL. This is basic SFT. We have these models that reason for 25,000, 30,000 steps.

  217. 44:29

    We take these traces, do SFT style distillation. So proper distillation, match your logits, and guess what? Um, this showed so much performance, but we know RL can do better.

  218. 44:42

    So you know, all open source, all traces are open. Someone do RL-based distillation from the big models. No one has done this as far as I know, but you know, we're able to get such good performance and there's still so much, uh, left to be done.

  219. 44:56

    But let's go over what we, um, distilled out. So these are the family of models. This is, once again, pure SFT style distillation. Oh shit, my slides are gone.

  220. 45:07

    They're back. Okay. So we distilled Qwen 1.5B, Qwen 7B, 14B, 32B, Llama 8B, 70B. Performance killed all the models they are themselves. So you take the model itself, you look at our distillation, it worked.

  221. 45:22

    Number went up like crazy. Really, really good performance. All our distills are now basically on par with GPT-4o, and for our new one, our new 8B distill is much, much better.

  222. 45:34

    It's way better than 4o. Um, and this is just, once again, RL on long chain of thought, take that, do SFT style distillation. Um, question, what if we just did RL on the base model, right?

  223. 45:47

    We tried it. We tried RL on Qwen 32B for 10K steps. It's actually worse than distillation. So you know, for small models, we don't wanna just start with native RL.

  224. 45:57

    We saw this in our own model, right? We needed this cold start. We need to kick off something from base models. R1 Zero to actual R1, we had to do this SFT cold start, so it actually performed significantly worse than distillation.

  225. 46:11

    Um, and of course, the Qwen team at the time, they had their own reasoning model, right? They had QWQ 32B and, um, you know, the R1 distill, so the base model distill, it did worse than what Qwen did, but kinda on par, you know.

  226. 46:24

    It's probably what they did. Our distillation on the, uh, chat model, on the base chat model actually performed so much better. We were b- we were able to do so much better than the Qwen reasoning model.

  227. 46:35

    And of course, now we've taken this a step further with our new model. Okay. Future work. R1 is worse than V3 at function calling, multi-tap, uh, multi-turn, and JSON mode.

  228. 46:47

    Guess what? Two weeks ago, there's a new DeepSeek model. We now do native function calling. We have JSON mode. We fixed it. It works. Um, R1 struggles with language mixing.

  229. 46:57

    We don't know how we do on the new one. Uh, it's sensitive to prompting. I think this is something that the industry needs to figure out, right? Um, spoiler alert for some, uh, Latent Space Podcast fans, you know, we talked to some reasoning experts.

  230. 47:10

    So like some of the stuff that we're seeing, you know, um, researchers at OpenAI, they're saying that, you know, "If you're still doing scaffolding with reasoning models, we're failing as labs."

  231. 47:20

    So you know, we need to learn how to prompt these things better. Um, and it's not much better at engineering tests than V3. That's all fake. This is old news.

  232. 47:28

    This was our old DeepSeek. The new DeepSeek is a lot better. Um, open recreations, we wanna promote open research, right? So there were people trying to recreate this. Uh, Hugging Face has a version of this.

  233. 47:40

    Um, Bespoke Labs was doing this. I don't know how they're doing. Uh, there's quite a few people now that have done this. But yeah, that's kind of an overview.

  234. 47:48

    I know a lot more people have joined. We now have like 100, 200 people in the audience, so I'm gonna do a quick recap in our last 10 minutes.

  235. 47:54

    So, um, first, we are launching a second paper club. Um, every y- uh, every week we do our normal paper club where we take the latest paper. We have 100 people that join every week, 300 for DeepSeek.

  236. 48:09

    We're t- we're turning this into a Test of Time Paper Club. If you're interested, sign up. Uh, we're gonna run this in SF, and we're gonna do it remotely.

  237. 48:18

    Over the next six months, we're gonna cover 50 to 100 papers. We're gonna break up what you would need to know as an AI engineer into different buckets. So stuff like, you know, what are foundations of deep learning?

  238. 48:29

    Attention, RL, um, optimizers, Adam, gradient descent, foundation models that you should know about, GPT-2, BERT, RNNs, LSTMs, pre-training, post-training, mid-training, so scaling laws, Chinchilla, distillation.

  239. 48:45

    We'll cover days of diffusion, optimization, voice, fine-tuning. We basically have a paper club where every week we're gonna split up these core concepts into a few papers. We'll have a presentation of three to four papers.

  240. 49:00

    Everyone is welcome to join. We're gonna have a presentation on every core concept and then open discussion. This is not a course. Is... Courses are good, you know. You have active workshop, you build stuff, you, you actually, like, do active learning.

  241. 49:13

    This is still a Paper Club, you know? This is... If you wanna know the foundations of what's going on under the hood, these are the key papers to know.

  242. 49:21

    We'll invite a lot of speakers, we'll have people present, we'll have good discussions. But yeah, Test of Time Paper Club coming in June. Scan QR code. Let us know if you wanna be involved, if you wanna recommend a paper, share a paper.

  243. 49:35

    Um, we'll share curriculums soon. Join the Latent Space, uh, Discord. We already have a list of top 2025 papers. It's a paper every week that you can go through.

  244. 49:46

    We're gonna, we're gonna build off of that. And once again, final recap, what did we talk about today? Today, we talked about the new DeepSeek model. So two weeks ago, DeepSeek-R1 May 28 came out.

  245. 50:00

    Basically, DeepSeek took the last DeepSeek model that was as good as o1. We trained it to reason for twice as long. We got significantly better performance. DeepSeek-R1 May 28 can now do standard structured JSON output, native function calling, hallucinates less, reasons for twice as long, and is a much, much better per- jump in performance.

  246. 50:24

    From being o1-level, DeepSeek is now on par with OpenAI's o3 model and Gemini 2.5. Basically, across the board on all benchmarks, we are now as good as, um, Gemini 2.5 and OpenAI's o3.

  247. 50:39

    The other model that was released is, um, DeepSeek-R1-Qwen38B-Base. So we took Qwen38B, distilled it down into a reasoning model based on our longer traces, and we killed it. So it's a small 8B where we do post-training via distillation as SFT on reasoning traces, and model got really good.

  248. 51:02

    You take base Qwen38B and you take RQwen38B, it's very good. Uh, looking at the benchmarks here, you know, our open source, on-device runnable Qwen38B non-reasoning model is on par with Gemini 2.5 Flash Thinking, o3-mini-medium, better than Phi-4, significantly better than Qwen38B.

  249. 51:27

    Our 8B reasoner is better than Qwen32B. It's on par with Qwen's 235B reasoning model. Um, so you know, two major updates. New reasoning model from DeepSeek as good as, um, o3 and Gemini 2.5.

  250. 51:43

    New mini 8B model that is as good as 2.5 Thinking and o3-mini. Of course, all open source, run on your laptop, um, just as good, you know? And these are not even R2.

  251. 51:55

    This is not DeepSeek R2. This is our mini title refresh with a new date. Um, yeah, so high level, that's what we talked about. Um, you know, we see aha moments.

  252. 52:06

    Instead of, um, training on next token prediction, we scale this out to inference-time scaling. So we now train models to train and think for longer. We get aha moments, and that's kind of the update to our new, um, to our new DeepSeek models.

  253. 52:23

    So yeah, thanks everyone for coming. Thanks for listening. Join Paper Club. [audience applauding]

  254. 52:34

    Who is this? Oh yeah, um, so a lot of our regulars that help make Paper Club are here. Uh, Eugene, Era, RJ, Flo is here. Um, if anyone else is a regular in Paper Club, you know, come up.

  255. 52:47

    Uh, every week we have our weekly Paper Club. These are, these are the homies that make it possible. Uh, we're gonna have our second Paper Club. Test of Time will be there soon.

  256. 52:56

    But yeah, uh, you know, major shout-out. This is not, this is not me. This is not Swix. This is volunteers and more on Zoom. Every week on a weekday at noon, hundreds of you join in to discuss a paper.

  257. 53:09

    So you know, big shout-out to everyone here. Not possible without us. You know, all the authors as well that have come, all the authors that have been able to share.

  258. 53:17

    Let's just give it up for everyone that makes, uh, Paper Club possible. [audience applauding]

  259. 53:21

    Hey, give it up for Yan. Eugene Yan, who is over there running his track. Sam I know is speaking right now. Swix who is putting all this on. Um, yeah.

  260. 53:30

    Can you leave it? Of course. Um, once again, I'll leave our QR code here. If you're interested in Test of Time, volunteering for a paper, recommending a paper, fill it out.

  261. 53:41

    You know, help us out. This is our Paper Club selfie. But yeah, thanks for coming out, everyone. Enjoy the rest of the conference. [upbeat music]