GLM-5.2: Open Weights, Near-Frontier Intelligence — Zixuan Li, Z.ai

Zixuan Li· Z.ai13:46

Read the talk

GLM-5.2: Open Weights, Near-Frontier Intelligence

Zixuan Li introduces Z.ai’s latest model through long-horizon coding, thinking budgets and general-purpose capabilities, then explains how open weights support local deployment, specialization and collaboration. The closing announcement, Z Code, adds a coding harness around the model.

From a talk by Zixuan Li

At a glance

Ideas worth remembering

  • GLM-5.2’s reported non-thinking improvement over GLM-5.1 with thinking separates gains in model capability from gains obtained by spending more thinking tokens.

  • Open weights support both local deployment and domain fine-tuning; architecture and training-recipe visibility support a further goal of collaborative model development.

  • Z Code is a separate coding harness built for GLM-5.2, with bring-your-own-key access to other frontier models.

GLM kept the research name as the architecture changed

Zixuan Li joins remotely to introduce GLM-5.2, while Z.ai’s team meets developers at the World’s Fair. The opening clears up a small naming puzzle: why does a company called Z.ai make models called GLM? The name comes from a 2021 paper on general language model pre-training with autoregressive blank filling.

Selected presentation frame from GLM-5.2: Open Weights, Near-Frontier Intelligence — Zixuan Li, Z.ai at 219 seconds
GLM kept the research name as the architecture changed

GLM began as the name of a research approach and became the model family’s brand. Li says the current models no longer use that original architecture. That distinction matters when interpreting a release: continuity in the name does not mean continuity in the underlying design.

The research goal also broadens. Math and physics problems made reasoning models conspicuous, but solving those problems captures only part of the intelligence Z.ai wants to build. Across GLM-4.5 through GLM-4.7, the work includes reasoning, coding and agentic capabilities. This sets up the new release’s emphasis on tasks that require sustained work rather than a single answer.

0:230:47
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Harder coding tasks bring the thinking budget into view

After a brief slide-navigation mishap, the presentation reaches GLM-5.2’s coding and agentic results. Li places its performance between Claude Opus 4.7 and 4.8 on the displayed comparisons, citing difficult coding and terminal benchmarks and long-horizon tasks. He describes a significant improvement over GLM-5.1. These are Z.ai’s reported benchmark comparisons; the presentation does not give enough evaluation detail to turn that placement into a guarantee for a particular repository or workflow.

Selected presentation frame from GLM-5.2: Open Weights, Near-Frontier Intelligence — Zixuan Li, Z.ai at 324 seconds
Harder coding tasks bring the thinking budget into view

The release introduces a High thinking level. Its purpose follows from a practical pressure: harder tasks can consume more tokens, so the amount of thinking becomes a budget decision. A higher budget gives the model more room to reason, while token efficiency remains part of the design goal. The talk does not quantify the extra token use or its cost.

The most striking comparison changes both the model generation and the thinking setting: Li reports that GLM-5.2 in non-thinking mode beats GLM-5.1 with thinking enabled. This separates two ways performance can improve—building a more capable model and allowing it to spend more tokens on a task. In this comparison, the newer model improves enough to win without the older model’s thinking mode; the new High setting then provides another option for difficult work.

4:074:08
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:07 · section reference included

The coding interface is only one use of GLM

GLM often reaches users through Claude Code, Codex or OpenCode, which can make the model look like a coding specialist. Its training scope is broader. Li identifies several parallel areas of work:

  • Professional tasks: improvements on GDPval extend the discussion beyond coding benchmarks.
  • Mathematics: math remains a training target even as coding and agentic work receive attention.
  • Conversation: role play and general chat are also part of the model’s development.
Selected presentation frame from GLM-5.2: Open Weights, Near-Frontier Intelligence — Zixuan Li, Z.ai at 359 seconds
The coding interface is only one use of GLM

Li reports that GLM leads other open-weight models on the Artificial Analysis Intelligence Index and sits close to frontier models. The practical invitation is to try it on general chat and daily workflows as well as code. A model’s familiar interface can narrow how people test it; the capabilities described here call for a wider set of tasks.

5:596:29
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:59 · section reference included

Open weights let customers change where and how the model runs

Releasing the weights raises an obvious business question: other inference providers can serve the model too. Li acknowledges that concern and explains the decision through the overlap between customer needs and Z.ai’s needs. Customers want control, specialized capabilities and a role in future development; Z.ai wants trust and useful collaboration.

Consider the enterprise or government deployment example. The desired change is concrete: the organization wants the model running on its own premises. Z.ai releases the weights through Hugging Face, making them available for the organization to obtain and run on its own servers. That changes who operates the model and gives the customer more control over deployment. Li presents this as a way to build trust; the example describes a deployment option, not a completed installation or a demonstrated security outcome.

Where does access to the model lead? The diagram follows the release to two distinct uses: running the model locally and adapting its capabilities. Both depend on access to the weights, but they solve different customer problems.

Local deployment changes the place of execution. Fine-tuning changes the model for a domain. Li names law, finance and security as areas where customers want different capabilities, and describes Harvey as already fine-tuning a GLM model. A move to fine-tuning GLM-5.2 is a possibility he raises, rather than a completed adoption. The business opportunity is that application companies can develop specialized behavior instead of relying entirely on the same general-purpose model.

A third use goes beyond deployment and adaptation. Customers who want to help shape future models may need to understand the architecture and training recipe. Weight access enables the first two paths; architectural and training information supports this deeper collaboration. Li connects that visibility to making better bets with customers and the open-source community.

How it fits togetherTwo uses of released model weights

Weights become available through Hugging Face.

One release supports a change in deployment location and a separate change in domain capability.

7:027:31
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:02 · section reference included

The ecosystem includes the software around the weights

Open weights need supporting software and people who make them useful. Li credits Unsloth, NVIDIA, individual developers and application builders for GLM-5.2’s success. His definition of the ecosystem includes models, open-source software and other support. Those contributions help users work with the models and push Z.ai toward better ones.

The technical blog is presented as the practical starting point for investigating GLM-5.2. Its contents serve different purposes:

  • Model repository and access: the Hugging Face repository, chatbot agent and API offer ways to obtain or try the model.
  • Individual coding use: a coding plan supplies a subscription path for token use.
  • Training explanation: pipeline and recipe details help explain the technology and difficulties behind the model’s performance.
9:019:31
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:01 · section reference included

Z Code adds a harness built for GLM-5.2

The closing announcement is Z Code, Z.ai’s own coding harness. It is built for GLM-5.2, but Li says it also supports other frontier models and lets users bring their own API key. That distinguishes the model from the software through which a developer uses it: GLM-5.2 supplies the model capability, while Z Code provides a coding environment that can connect to different models.

Li compares its operation to Codex and mentions compaction techniques familiar from coding tools. He recommends it as a particularly good harness for GLM-5.2 and invites attendees to see the team’s booth demonstration. The announcement establishes the intended fit and model access options, but does not demonstrate the harness’s internal operation or establish a performance advantage over other harnesses.

The host’s wrap-up brings the discussion back to the people who make local models practical. The World’s Fair gathers model labs and developers so users can meet the teams behind their tools and ask detailed questions. NVIDIA, Unsloth, Ollama and fine-tuners are named in connection with the upcoming inference and local tracks. The ending gives the open-weight argument a practical setting: model releases, inference software, adaptation tools and application builders have to meet for the capabilities to reach users.

11:0111:14
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:01 · section reference included

Read the complete timestamped transcript
  1. 0:12

    Without further ado, we're gonna bring on Zixuan, uh, for, for his keynote, uh, if we're all set.

  2. 0:20

    Hey, how are you?

  3. 0:22

    Great, I'm here.

  4. 0:23

    Uh, yes, uh, to the magic of the internet. Uh, it's really good to see you again. Uh, Zixuan has, uh, you've been speaking with, uh, AI Singapore, uh, and obviously GLM is, uh, the talk of the town right now. Uh, there's been an amazing booth downstairs in the expo, and, uh, I'm, I think a lot of people are just very excited to hear from you on, uh, what all is going on at Z.ai. So if you wanna take it away.

  5. 0:47

    Yes. Sorry again for not showing up in person, but I have a team, whole team coming to the town, and we have a booth. So you have any questions, feel free to reach out to me on X, LinkedIn. Also, you can look for my team on our booth. And since I cannot see my slides, so I will need swyx to help me flip through all the slides for me. Yeah.

  6. 1:14

    You're good.

  7. 1:14

    Can you share, like, what, which slides we are in right now?

  8. 1:18

    Yeah. Uh, so we're, we're on the opening slide. Uh, w- we're talking about intelligence and Z.ai.

  9. 1:25

    Okay. So yeah, so you can, you can see that it's the first time for us to introduce GLM 4.2, uh, 5.2 to the world, and also we are gonna share something about Z.ai and GLM, because maybe people will think Z.ai and GLM, they're, they're irrelevant, right? Your, your company's not G.ai or your model is called Z1 or Z2. And you can find my X and our company's X account here, so you can just search

  10. 1:55

    for my name, Zixuan Li, and the Z.ai org, and you can follow them for the follow-ups. So yeah, we can go to the second slide.

  11. 2:06

    Yeah. Actually,

  12. 2:09

    I cannot see the slide, so I'll try to-

  13. 2:12

    So we, we-

  14. 2:12

    Yeah, try to make sure that i- it is correct.

  15. 2:15

    Yeah.

  16. 2:15

    So the company ac- actually is called zhi pu. Maybe some people have heard of it. And the mo- the models, uh, it's, it's called GLM. Actually, it's not a, a brand name, it's a generic term. So GLM actually represent general language model pre-training with autoregressive blank filling, and that paper was published back in 2021. So actually we were the, one of the first

  17. 2:45

    labs to do explorations on large language models, so at the same time with OpenAI and Anthropic and DeepMind. And even today, we, we no longer use GLM as the architecture. We still use the name GLM as our brand name. So we use J- GLM 5.1, 4.2, and it become, like, one of our, like, most prod, uh, proudest product and model. And the second thing that we look for is

  18. 3:15

    intelligence upper bound. So in terms of intelligence, we, we may feel that it's represent IQ, something, something like that. And when DeepSeek launched, when o1 launched, people are talking about the model's capability to solve math problems, physics problems. But what actually intelligence mean is not just, like, IQ or AME or other, other physics problem. So from GLM-4.5 to GLM-4.7, we are

  19. 3:45

    exploring, like, several things like reasoning, coding, agentic capabilities. So as you can see from the slides, so we, we add, like, uh... Yeah, the last slide. Yeah, we were half done. I haven't, like, finished that slide. Yeah, so okay. Yeah, need to go back to the,

  20. 4:07

    two, two slides.

  21. 4:08

    I don't have a back button.

  22. 4:09

    Back.

  23. 4:11

    I don't have a back button.

  24. 4:15

    Thank you. I would... I don't understand clickers-

  25. 4:18

    Thank you

  26. 4:18

    ... that don't have back buttons. Like, why? Okay, you know. Anyway, go ahead.

  27. 4:25

    Yeah, never mind. Yeah. Because people want to see the GLM 5.2. They, they don't want to see, like, GLM 5.1 or 5, but, like, GLM 5.2 actually specialize in coding and agentic task, as you can see from the graph. Because there are a lot of rumors whether your model is close to Mythos, Fable, but actually I want to share these slides to all of you. So you can see it's somewhere between Opus 4.7 and 4.8, and we use the hardest problems, like DeepSWE

  28. 4:54

    Terminal Bench 2.1, which was mentioned by the OpenAI team, uh, several minutes ago, and all the, like, long horizon task. And benchmark shows that the, the capability is on par with at least Opus 4.7, and it, it, it shows, uh, significant improvements over 5.1. Uh, also for GLM 4.2, we add a thinking, uh, level called High. So because we also

  29. 5:25

    notice as we move to the harder task, it may consume more tokens, and also we ca- we care a lot about the token efficiency. So it's the first time we add the high level for thinking budget. But even without thinking, the non-thinking model is better than the 5.1 thinking model. So I think it's a huge improvement for the open weight model. That's what really impressed the world and why people are talking about GLM 5.2

  30. 5:55

    lately. Okay, the next slide.

  31. 5:59

    And one, one thing that I want to mention is that GLM is more, more than coding model. Because people use it inside Claude Code, Codex, OpenCode. But actually, we have trained a lot of things outside coding. For example, we improve a lot in GDPval and also math problems. We also care about math problems, frankly speaking. And also we train a lot of thing, uh, related to role play, general chat. We want to improve

  32. 6:29

    every aspects of the model. So you can see from the Artificial Analysis Intelligence Index actually leads the other open weight model a lot and close to the frontier model. So I want you, if you, you haven't experienced GLM yet, you can use GLM to do general chat, use it to process your daily workflow, not just for coding, but you can explore the model, like, beyond the coding scope. Next slide.

  33. 7:02

    And GLM-5.2 is a open weight model. So people always ask me, "Why do you open weight?" So do you care about your business or, like, do you care about losing market to some inference providers? But actually, we open the weights for several things, because they are users' needs and they are our needs. If we can meet their needs, I think it's okay. It's definitely okay for us to, to open the model. For

  34. 7:31

    example, if our users want security and control and we want to build trust, we can open weight the model. For, for some cases, if an enterprise or government, especially in the Western world, want to use the model, we open weight, we upload to the Hugging Face so that they can use the model on premise. I think it's very beneficial for the whole ec- ecosystem to explore the model, especially when the capabilities is

  35. 8:01

    close to the frontier model. And second, if people want diversity. So in terms of diversity, I mean the capabilities in legal, finance, security. They can fine-tune the model, so we need to open the model to let them fine-tune. For example, Harvey is fine-tuning GLM-4.1, and maybe they're thinking about fine-tuning GLM-5.2 afterwards. And I've, I have, like,

  36. 8:31

    talked to a lot of other companies. They are also thinking about fine-tuning GLM as their, like, next step or next strategy to differentiate themselves from other application problem. And third, if our customer or an individual want to co-design and predict the future, sometimes they need to see the architecture of the model. They need to see the recipe of how you train the model.

  37. 9:01

    So we want to make the norm. We want to make right bets. So we want to co-shape the future with our customers, with the open source, uh, community. So I think our needs and our ecosystem pretty, like, fits into each other. And GLM-5.2 couldn't succeed without you. All the open source community players, like Unsloth, NVIDIA, and, and some, like, individual

  38. 9:31

    super developers, they are part of it. And thanks a lot to application builders like Peter, right? Because open source doesn't just include open source model, but also open source softwares and the open source other sorts of support. So all your support and what you are, you're doing right now pushes to, to make better models. I think, uh, you, you are the true

  39. 10:01

    hero. And the last, I want to share a, a great s- resource for you to go through GLM-5.2 is our tech blog. So actually, in that tech blog, we share several things like our Hugging Face re- uh, repo, how, how you can try the GLM-5.2. You can try the inside the chatbot agent, and also you can call the API. And also, we have a coding plan, like the Codex or Claude Code

  40. 10:31

    subscription for you to use your tokens as individuals. And also, we share something about our training pipeline, training recipe, which you can understand why it's a good model. So there are a lot of things behind the, behind the model, not just, um, a, a, a model that had great data. We also have fantastic technologies behind the model, so you can explore the model yourself. You can see what difficulties we have gone

  41. 11:01

    through. And the last slide actually is kind of a one more thing. So it's the first time we share ZCode to the whole community.

  42. 11:14

    Whoo.

  43. 11:14

    So that one more thing is we actually have our own harness. The ZCode actually is built for GLM-5.2, but also support all frontier models. You can bring your own key. You can connect to ZCode. Actually, the... I think the operation is, is similar to Codex. Actually, you, you can try some techniques like Gol or other, like, compact techni- techniques like what you did in,

  44. 11:45

    in Codex and Claude Code. And this harness, I think it's the perfect one for GLM-5.2. If you haven't experienced it, you can just search ZCode or you can go to our booth. We have our team members showing this harness to you, and welcome to, to our booth. And welcome to, like, talk to me in the future. And next time I'll definitely be in SF talking to everyone. Yeah, thanks.

  45. 12:13

    Thank you very much.

  46. 12:19

    Um, so, uh, trust me, the very first World's Fair, I wasn't allowed to come back into the country for, for my own conference, so I know exactly this feeling. Uh, but it, it, it's all good. Uh, the ZAI team, I really appreciate them, uh, making the effort. They really wanna meet you. They're here to meet you, and this is the whole point of the World's Fair, to bring all the top world's AI, uh, companies and labs, uh, all in one place so you can do business together, uh, meet the people behind the models that you use, ask the questions

  47. 12:49

    that he cannot answer in public, but you can ask in private. Um, I'm also very proud. Uh, he showed, uh, Zixuan showed that, that list of, uh, Hugging Face, uh, you know, top, uh, contributors. I think we are four for eight, uh, in that list, uh, present at World's Fair. Uh, we're working with NVIDIA and Unsloth and Ollama and, and, uh, all those sort of fine tuners as well to, uh, to make, uh, sort of the inference and the local tracks, uh, that you're gonna see over the next few days. Um, with that, thank you so much, Zixuan. Uh, I'm

  48. 13:18

    gonna invite on, uh, the next, uh, couple folks. I think, I, I think I might be doing Ali's job here. So, uh, we'll, we'll talk to Hugging Face, uh, next and, and Minimax. Thank you.