AI Engineer Code 2025
What Makes Open Models Fast in Production — Sujee Maniyam, Nebius
Read the talk
What Makes Open Models Fast in Production
Dylan Bristot and Sujee Maniyam explain how Nebius Token Factory connects production data to model improvement, then walk through the hardware, routing, caching and decoding choices behind fast inference.
From a talk by Dylan Bristot and Sujee Maniyam
At a glance
Ideas worth remembering
A production model needs a recurring path from inference logs through dataset preparation and post-training back to deployment.
Serving the same model on different hardware and engines can produce different performance and cost; engine selection belongs to the model-specific optimization work.
Cache-aware routing improves reuse by sending requests toward GPUs that already hold useful cached state.
Speculative decoding pairs a fast draft model with a large verifier; custom training uses application traffic to make the draft more useful.
Cache offloading preserves reusable work outside GPU memory, while prefill/decode separation assigns different resource demands to separate GPU groups.
Quantization requires experiments to balance serving efficiency against model quality loss.
The model is only the beginning of the production job
An easy model API can become a difficult production dependency. Dylan Bristot, who leads product marketing for Nebius Token Factory, opens with the choice facing teams that want to control model behavior and operating cost: use a closed API, or take responsibility for serving an open model themselves. Sujee Maniyam, a Nebius developer advocate, joins him to explain the engineering underneath that choice.
- Closed APIs: They make the first integration easy. Bristot’s concern is the ceiling afterward: limited ability to tune the model for a specific task, little visibility into shared infrastructure, and costs that grow with usage.
- Self-hosting: It gives the team control over the model and serving setup. That control also brings a substantial engineering project and an ongoing responsibility to keep it running.
- Managed open-model inference: Token Factory’s proposed middle path is to handle deployment and serving optimization while customers concentrate on their products and the behavior they need.
The infrastructure matters to this proposition. Nebius describes a stack extending from its own data centers and NVIDIA systems through bare-metal capacity, managed cloud and model serving. Owning those layers gives the inference team room to change how a model runs, rather than treating the hardware underneath an API as someone else’s fixed constraint.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn production completions into the next deployed model
Token Factory connects four activities that teams otherwise have to stitch together. At the time of the recording, Bristot describes more than 60 hosted models, including GLM, Kimi, DeepSeek and Qwen, with dedicated endpoints, structured outputs, function calling and batch APIs. Those endpoints start the cycle by producing real application traffic.
- Data Lab: Import inference logs and completions, then filter, version, export and batch datasets. This turns accumulated responses into material a team can inspect and use.
- Post-training: Use production logs or synthetic datasets for LoRA or full fine-tuning, distillation, custom speculative decoding, and model-specific quantization and calibration. The goal is to make the model better suited to the application’s expected behavior.
- Deployment: Push the resulting model onto the owned infrastructure so the improvement reaches the product and produces the next round of experience.
How does an application response become a model update? The cycle below shows the handoffs. Completions enter the data layer; selected data feeds post-training; the trained model returns to production. The useful relationship is the return path: deployment creates new experience to inspect, so serving a model becomes an ongoing improvement process.
Keeping that cycle working also requires less glamorous operations. Token Factory groups its responsibilities into engine-level optimization of latency, throughput and cost; operation of the underlying infrastructure; and enterprise requirements such as security, compliance and reliability. Customers still care where the model runs, which chips serve it and how much it has been quantized. A managed API hides operational work without making those choices irrelevant.
Run models and produce application completions.
The platform connects serving, dataset preparation, model improvement and deployment into a recurring cycle.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Match the model to the hardware and serving engine
Maniyam begins the technical half with model quality. An Artificial Analysis chart compares open and proprietary models, and he uses it to show that open models can be competitive, sometimes ahead of proprietary alternatives. That comparison supports considering open models for intelligence as well as cost; it does not establish parity for every application. Their portability also gives customers more choice about who serves them.
The hosted models have different roles. Large, capable models can act as teachers for smaller student models. That gives post-training a route to a smaller serving target: use the large model’s capabilities to help train a model that is less expensive to run. Later, a different pairing of small and large models will accelerate generation itself.
Serving performance then depends on matching several layers. Nebius optimizes kernels and runtimes for its hardware, including adoption of NVFP4, NVIDIA’s four-bit floating-point format. Above that, Token Factory uses open-source serving engines and an internal fork that it continually modifies. Models do not perform equally well on every engine, so engine selection happens per model rather than through one universal choice.
The customer sees an API, but the provider chooses the GPU, runtime and engine combination behind it. This helps explain Bristot’s earlier observation that the same open model can have different operating costs on different providers: identical model weights do not imply identical serving behavior.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Route requests toward the work already cached
A random load balancer misses something important about LLM workloads. Early requests might have had a short prompt and a short answer. Coding agents now send large codebases for refactoring, summarization or analysis, and a large input can lead to either a short or a long output. Request count alone does not describe the work being assigned.
Consider the coding workload Maniyam introduces. A request reaches a GPU that may already have useful cached context—or may have none of it. With random routing, requests scatter across GPUs without regard to that reusable work. The result is fragmented caches and missed opportunities to reuse them.
Cache-aware routing changes the destination decision. The router knows where cached material is stored and sends requests accordingly. In Maniyam’s comparison, scattered cache colors become grouped together: related work reaches a location where it can benefit from existing state. The observable change is better cache locality, which supports more cache hits and faster inference.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Let a small model draft and a large model verify
Speculative decoding targets the cost of generating tokens one after another with a large model. A smaller, faster model proposes tokens, and the large model verifies the proposed work. Maniyam compares it to a senior engineer handing work to a junior engineer and then checking the result.
The speedup depends on how much of that draft the large model can use. Accepted work avoids having the large model generate everything itself. Rejected work requires regeneration, reducing the benefit of drafting. The small model therefore helps most when its proposals suit the large model and the application’s traffic.
That creates a reason to train custom draft models. Maniyam reports up to 30% improvement from draft models trained with generic data and says application-specific data can improve the result further. The recording does not define the improvement metric or evaluation conditions, so the figure is a reported experimental result rather than a general speed guarantee. He describes a launching feature that captures production data and trains a draft model with nearly one-click setup, connecting the earlier data loop directly to serving optimization.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep reusable context, even when it leaves GPU memory
KV caching addresses repeated work during generation. Maniyam’s example starts with the prompt “The cat sat on” and then generates the next token. As the sequence grows, its existing context remains useful for later steps. The cache retains reusable state associated with that context; the next step looks up retained work rather than rebuilding everything already processed.
The example makes the distinction useful: caching supports the next generation step, rather than choosing all future words in advance. Maniyam reports speedups of five to ten times from caching, but does not specify the workload or baseline for that range. Long inputs and large context windows also make the retained cache consume substantial memory.
GPU memory is valuable, so retaining every cache there is costly. Token Factory automatically offloads cached state to regular memory and brings it back when needed. In the cat-prompt example, the important change is where the retained state lives: leaving GPU memory need not mean discarding the work used to build it.
The next optimization separates two phases of inference. Prefill processes the input context and calls for compute-efficient execution. Decode generates the continuation and places heavy demands on memory. Running both on the same GPUs makes them compete for resources; Token Factory assigns them to separate GPU sets and transfers the KV cache between them.
What connects the two GPU groups once their jobs are separated? The diagram shows the KV cache as the handoff. Prefill prepares context state; decode receives that state and continues generation. Separation lets each group suit its phase, while making cache transfer part of the serving path.
Process the input context with compute-efficient execution.
Prefill and decode use different GPU groups, with retained context state moving between them.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Find the precision that preserves useful quality
Quantization reduces numerical precision to make a model run more efficiently. Reducing it too far can degrade model quality. Token Factory experiments to find a sweet spot: an efficient representation whose quality loss remains acceptable. This is a model-serving decision that needs attention alongside engine choice, caching and decoding.
The closing recap brings the mechanisms together. Speculative decoding reduces the generation work assigned to the large model. KV caching retains useful work, cache-aware routing sends requests toward it, and separate prefill and decode groups divide competing resource demands. The API can conceal all of these operations, even as they determine how well very large models serve production traffic.
Maniyam ends by returning to coding as a practical use case. He describes using hosted open models for coding and incorporating them into Cursor or OpenCode, then invites the audience to try the platform and share feedback. The talk’s final move is from hidden serving machinery back to the application: these optimizations matter because they make capable models more useful in the tools people already use.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
Thank you for joining us today. Um, I'm Dylan. I'm here today with my colleague, Sujee. So I lead, uh, product marketing for Nebius Token Factory, and I'm here with Sujee who's as, uh, developer advocate, and we're gonna talk a bit about, uh, engineering, um, open LLMs for production. So before we jump right in, um, I wanna give a few words about what we're gonna cover today. We only have twenty minutes, so we'll try to keep this brief. Uh, but we're gonna give you a quick intro on who we are, what we do, uh, then talk about what it takes to build the infra for inference.
- 0:42
Um, a few words on model shaping optimizations, which are extremely important, uh, in what we do, and then we'll talk about model serving optimizations.
- 0:52
But before we jump right in, I just wanna give a few words, uh, about who we are and what we do. So we work for Nebius, uh, so specifically Nebius Token Factory. Uh, but Nebius is, like a full stack AI, um, cloud infrastructure. Um, so this means we actually not only exposing model APIs, we run the actual physical, uh, AI layer and, and infrastructure underneath it. Um, we have our own data centers, NVIDIA systems, uh, bare metal capacity, managed cloud, and optimizing the serving and the inference through Nebius Token Factory.
- 1:23
Um, top, top row here is really about scale and trust. So we're a publicly traded company on the Nasdaq, um, headquartered in Amsterdam. Uh, we're very close with NVIDIA, who made a two billion, uh, investment in Nebius a couple of month ago, and we're targeting a five gigawatts of NVIDIA systems by the end of, uh, twenty-thirty. Uh, bottom row is really more about the control of the stack. This is a quite a, a good edge we have. Uh, we operate own data centers across the EU and the US, uh, bare metal rather than shared cloud.
- 1:53
We're a very early adopter of NVIDIA Rubin, um, CP, like, uh, and, uh, BlueField storage. And we were amongst the first European clouds to run both Blackwell Ultra, HGX B300 and GB300s, uh, in production. Uh, we also have, um, s-s-like strong contracts from M-Microsoft and Meta agreements that kind of show that the largest AI builders are already building this kind of capacity from us. Um, and the last time is quite important. It's the, uh, to like read the important bridge to this talk is the vertical integration from
- 2:23
silicon to serving that we, um, rely on.
- 2:28
So today, we're talking about Nebius Token Factory. So the idea here is to, like, really, uh, engineer AI for products. So we're a managed inference company, uh, meaning we deploy these models and optimize them for our customers.
- 2:41
So we noticed a few thing today in the market of, uh, serving open source LLMs. Um, usually most AI teams are stuck between choosing two bad options. Um, closed APIs are really easy to start with, but you often hit a ceiling, uh, really fast, and you can't really tune the model to your specific use case. Uh, you share the infra with everyone else, so it kind of operates as a black bos-- black box, and the costs grow in a straight line with no way to really optimize things. The other option here is self-hosting. Obviously, it's quite fun, uh, gives you full control on the
- 3:11
model, but it's a massive engineering project. Um, so you really need a dedicated team to actually keep it running. Um, and you're usually a month away from production before you even start your actual product. So Token Factory is kind of the shift, uh, this third path we identified, uh, before you even start, like, on your, um, actual product. You get the control and the performance of self-hosting with the simplicity of a managed inference service. So we do the heavy lifting, the hard infrastructure work, and then customers stay focused on their product. This really unlocks, uh, three main
- 3:41
things, performance, cost, and behavior of the models. Um, as we've seen, the gap between proprietary and open source models has been closing a lot. And so this is where we kind of operate, and we feel comfortable, um, operating. It's really not only relying on the fact that this bridge and this gap is closing down, but also figuring out how you actually make it work in your own production, uh, use cases.
- 4:05
So this is really the core of Token Factory. Uh, we've built a full stack GenAI platform. Um, most platform we've seen on the markets either give you inference, training, uh, data tooling, but here we decided to connect everything under one loop. So we have a bunch of tools to-- that kind of unlock this. On the inference side, um, we allow you to run models in production. So we have sixty-plus models on the platform, GLM, Kimi, DeepSeek, Qwen, you name them. Um, we have dedicated endpoints. We have different flavors on the models, depending if you want to run them fast. Uh, we have different
- 4:34
tooling like structured outputs, function calling, and batch API. Second layer here is the data lab. So this allows you to really capture and structure the production logs. So we see a lot of customers actually running these models in production, importing the completions, and wanting a way to kind of slice and dice, uh, through these completions. And this is really what we allow with the data lab. So, uh, inference log imports, SQL data like, uh, dataset filtering, dataset versioning, um, exporting, and then batching, uh, these. Uh, the third layer is post-training. So we allow you to kind of fine-tune,
- 5:04
distill, optimize these models. So this is a natural follow-up with the data lab, is once you have your logs and your inference and you can generate synthetic datasets, you can bring it all to post-training and use that to kind of either through LoRA or full fine-tuning, make the models much more adapt to the behavior you're expecting from them. Uh, we also allow model distillation. We do custom spec decoding for customers who need really advanced, um, like cus-custom spec decoding, uh, tooling. Uh, we do like on-demand and like specific quantization and calibration of the model depending on
- 5:34
your preferences, and we do our RGPO. And the fourth layer here is the deployment. So once you, like run these models, once you like train these models, uh, you need a way to deploy them. So these-- this is what allows us to push the models into production on our own, um, owned infra. So for most teams, this is kind of usually a choice of different tools that you stitch together, uh, but that's creates kind of a lot of friction and need for iteration. It slows the need for like the, the ability to iterate here. And so what we allow is really
- 6:04
to make the data flow from, uh, the inference to the training, and the training flows directly into production. So that allows you to keep like a virtuous loop of your product. And that's really what I think is the main difference between, uh, running a model and running an actual production, uh, AI system.
- 6:22
Um, so here is this little slide about what we handle, so you don't have to. Um, so running models in production is actually quite complicated. It takes a lot of work. And so I think, uh, the way we approach it, the value here is in three main layers. The first one is the optimization. So we actively tune how models run, so the customers get much better speed and lower co- and lower costs without having to figure out, um, all of that themselves. Like, we really do the heavy lifting around the model deployment and optimization, and we do allow you to kind of, uh, come in and just play with them.
- 6:52
This is, like, all around optimizing latency, throughput, and cost at the actual engine level 'cause you could be running the same open source model on different providers and have completely different, um, results in terms of costs. So there's a lot of technical stuff that goes on under the hood, uh, to make this work, and Sujee will cover that in a few minutes. Uh, second layer is the infrastructure. So the fact that we actually own the hardware, uh, we keep it running, we own this, like, kind of vertical stack, um, gives us an edge that, uh, most providers don't have. So it's, it's, it's the control...
- 7:22
Like, we really control what the model does. Uh, we handle, like, everything underneath. And so we just give the, the customers, um, the full control of the be- behavior, and we take care of the infrastructure. Third layer is the enterprise readiness. Uh, I think this is really important in here 'cause, like, right now, open source seems kind of obvious, but you need a bunch of ways to, like, secure-- like, make it-- make sure it works in a secure environment. So there's a bunch of things here. Um, security, compliance, reliability, they're all requirements that all these large companies usually have, um, and they kind of
- 7:52
need those before they can sh- ship anything into production. And so a lot of AI providers make you figure out all of these three yourself, and, uh, our mission here is to really handle all of these so you don't have to.
- 8:08
Um, something that comes up a lot with AI teams is that they start with inference. Um, they just get a model up and running, and then they realize that pretty qu- quickly, quickly that running it is not just as easy as it seems, uh, and it's only the beginning of your path to running open source models at scale. So they really need to understand how it's actually performing in the real world, and this is kind of h- where we come in. They wanna improve it over time. Um, they need to redeploy updates without breaking things. Um, and, and that full cycle
- 8:37
from running to observing to capturing and to redeploying the model is really what separates team from that ship, like, great AI products from teams that ship demos, and I think that's re- really important to hear, especially when talking about open source. Um, and so Token Factory covers that full loop, uh, inference to run the models. So as I said, Kimi, DeepSeek, all the, all the main ones, uh, through real-time inference or batch inference, uh, through structured outputs and function calling, um, and on dedicated endpoints when it's necessary. Um, and then moving on to the data
- 9:07
lab to collect and analyze the real-world signal, uh, the completions, moving over to post-training to improve and customize the model, and finally moving to the last step, and probably the most important one, which is the deployment, to push, uh, updates with full control over the model. And you actually really... It's, it's really important for most our customers to know where, how it's running, and what chips it's running on, uh, the level of quantization and all that stuff, and that's really where we kind of operate, uh, and are comfortable. Most teams today, we figured, have to stitch different tools together to get this, and we actually
- 9:37
built, uh, one platform that tries and covers it, uh, end to end. So we really see this as a continuous improvement loop. Every cycle makes your AI more specific, uh, way faster and much cheaper.
- 9:50
And now we'll dive into scaling open models to production, and I'll invite Sujee on stage to take it over.
- 10:01
All right. Thank you, Dylan. All right, everyone. So, um, I'll zip through a few of these, but, um, I'm happy to, um, have a discussion afterwards, so come and find us if you have any questions. So we believe great inference start with great models, but that's not enough. You need great infrastructure behind the scene to get the best performance, and, uh, I believe that's what, uh, Nebius Token Factory provides. Uh, I get a lot of times I get asked like, "So this is great, you guys host open models, but
- 10:32
how good are the open models really?" And I like to show you this benchmark here. Uh, this is from Artificial Analysis. The black ones are proprietary models, and the blue ones are open models. And the cool thing you can see here is how, um, the open models are actually very competitive, sometimes even better than a lot of the proprietary models. So the gap between, uh, proprietary and open models is very narrow. Uh, so we don't really have to chase proprietary
- 11:01
models all the time for the intelligence. You can just, uh, just as well choose an open model that's just as smart and also ends up being much cheaper in most cases, which is pretty cool. So what does it mean for our customers? You get choice, right? You can run these open models in any provider, hopefully Token Factory. But, so there's no vendor lock-in, and also you end up saving a lot of money on tokenomics. And here's a, a few examples of models, uh, we host. Uh, the top, the top, uh, table is like the
- 11:31
large models. We call them sort of the teacher models, right? They are very capable, uh, very smart, but also kind of large. You can also use a large teacher models to train the smaller student models. So you sort of see that both large models and small models hosted. So I wanna sort of structure the talk a little bit, um, sort of from the hardware layer, the serving layer, and the optimizing layer. Uh, so Dylan t-talked a little bit about hardware. We are-- We work very closely with NVIDIA. We get the latest chipsets. And, uh,
- 12:01
not j- not just, um, the latest, we also optimize, uh, kernels and runtimes, so the models run really, uh, really well. For example, something like the NVIDIA Floating Point four standard, we embraced it. It gives us a really good model performance on latest NVIDIA chips. And if you think about model serving, that kind of takes us to the middle layer, and there are a few optimization I, I will talk about. Uh, starting with the engines, because you need an engine to serve the model. Uh, we utilize a lot of the open source
- 12:31
engines. Uh, we have our own internal fork, uh, that we continuously tweak and optimize, and we will deploy the model on the best engine possible. Uh, you know, not all the models run equally well on all the engines, so we figure out which engines run well and we'll deploy the models accordingly. So as a customer, you can just use an API not having to worry about, you know, which kind of GPU, which engine, because we take care of all that work behind the scene.
- 12:58
And, uh, when you have, like, hundreds of thousands of GPUs, uh, running models, load balancing and routing becomes an issue. So a lot of the time people say, "Well, that's easy, right? We just, you know, put a, put a load balancer in front, um, send traffic randomly to different GPUs." Uh, that's not how it works in LLMs because the workloads are slightly different. For example, just to kinda give you an idea, uh, in the early days of LLM, the inputs were small. I could just say, right, like a one-liner, and the output is also small. But now we are seeing your
- 13:28
input could be pretty large, like we are sending, like, large code bases, uh, as... using coding agents to do, like, refactor or summarize or analyze, right? So the output, as you can see, can vary from, like, you know, s- large to small, large to large, and routing them accordingly will give you, um, you know, it's not, it's not trivial. So what we employ is we employ, um, routers that actually are cache-aware. What I mean by that is, like, if you look at the left side here, you will see that my cache, like different colors, it's kind of
- 13:57
fragmented all over. So our cache hit rate isn't, isn't all that great because my request r- landing randomly at different GPUs that may or may not have the cache. But if you look at the, the one on the right, you can see that now the cache is much more coherent. See the colors are kind of together. And the router knows where the caches are stored, and it will route them accordingly. So when you actually do the inference, you get pretty good cache hit and very good speed.
- 14:23
Um, another thing, and this is something we are really excited about, um, it's called spec decoding. So the challenge is, uh, LLMs, they generate tokens one by one. One, two, three, four, right? And then large models take a long time to generate these tokens. So the idea for spec decoding is, how about we employ a smaller model, which is a lot of the time faster and cheaper to generate the tokens, but then have the large model verify at the end? So it's kind of like if you're like
- 14:53
a senior engineer, you're kind of farming out the work to a junior engineer, so they're kind of doing the heavy lifting, and then you are verifying the work. And this actually surprisingly scales really well. So the smaller models can generate tokens very quickly, and the large model can verify. And they-- a-and if it's good, we are done. If the large model doesn't like the, the answer, it'll actually regenerate the tokens for you. So that's kind of the worst-case scenario. But most of the time, you will get pretty good performance out of this by mixing small models and large models. And this
- 15:23
actually, uh, feature we actually baked into Token Factory, uh, that gives you pretty, pretty good optimization, um, speed. So one of the things we ex-experimented, we sort of trained draft models. Uh, you can train them using synthetic data. Even if you train them using like, you know, generic data, you get like, you know, we are seeing like up to thirty percent improvement. But it, it gets even better if you can train the smaller model using your own data, right? So we did a lot of experiments here, and, um, we are lo-launching a feature, feature you can actually, uh, train these draft models almost
- 15:53
like with one click. And, you know, we will capture your production data and train the draft models for you, which is pretty cool. Uh, another optimization, uh, behind the scene is KV cache. Um, this is probably one of the biggest R-ROI i-in doing inference because, uh, creating tokens are expensive, especially when you're doing one at a time. And the common sense is once you generate a token, it n- it doesn't change. So there's no need to keep, keep regenerating the same token over and over again. So the idea is we cache the
- 16:23
tokens. So just to kinda give you an example here, let's say this is my prompt, "The cat sat on," right? Let's say tho-those are my prompt, and I'm creating the next word. And as soon as we create the next token, we cache it. And the next time we need to look up, we don't have to regenerate it, regenerate it. We just look up the cache. Very easy. And this actually speeds up quite a bit. So just to kinda give you some ideas, w-we are seeing, uh, speedups anywhere from five to ten X. Not just five to ten percentage, actually ten X, right? Because caching really works well.
- 16:53
Um, the thing about caching is it takes a lot of memory, right? Because, you know, especially with our large inputs and the large token windows. And one thing we are, we, we've been implementing in, uh, Token Factory is when we do caching, we can actually, uh, because GPU memory is very precious, so we can offload caching out of GPU memory automatically, save it in a regular memory, and then bring it back. So we don't, like, throw away the cache we spent a lot of time building. We actually offload it and bring it back, and all of this is done automatically. So you don't
- 17:23
have to worry about managing cache. This is done for you behind the scenes. So, and th-this is one of the, um, um, one of the important optimization we have done to speed up inference. Uh, another one is called decoding. Um, so the way you think about this is like there are two stages to LLMs. So one of them is, like, sort of filling the context that is very CPU intensive, and then the other one is actually decoding that is very memory intensive. And a lot of the time, if you do them in the same GPU, they are kind of competing for resources. So what we have
- 17:53
done, we sort of separate these two phases out. So one GPU set does prefill. They are very, you know, uh, compute efficient, and the other one does, uh, decoding that's very memory efficient. And then they sort of transfer the KV cache in between. And we have seen this is actually working, working out pretty well.
- 18:12
And another thing we do, and, and this is already, you know, already baked into our, our platform, we also quantize the models. Uh, and when we do quantizing, you know, we are reducing the precision, so if-- but we don't wanna go too far, right? Because then your model starts, you know, quality starts getting degrading. So what we do is we do a lot of experiments to figure out like what's the kind of the, the sweet spot of quantization- So you can run the models efficiently, but not, not degrade the performance too much. So, so all of these th-things are already built
- 18:42
in, and, uh, yeah, and, and, and there's a, there's a few more, and in interest of time, also wrap this up here. And these are some of the things we kinda covered, uh, spec decoding, KV cache. And when you do caching, doing, um, routing, that is actually cache aware and also pre-fill, uh, you know, se- uh, separating pre-fill and decoding phases. And if you're using an API, you may not realize all these works going behind the scene, but they are, right? And th-this is what make up, you know, make a, a, a inference platform really great to serve, especially the large, large
- 19:12
models we are seeing, like the one trillion models, uh, we are seeing. And finally, I wanna quickly, um, give this out. So we, um, probably this is the most important slide of the presentation. So you can take the... Take a quick picture or scan the code. Uh, this will give you some credit on Token Factory platform. So we allow you for-- allow for you guys to try it out. And, and we have all the co-- uh, latest open source models, uh, GLM-5.2, Kimi, uh, K2.7. Uh, try them out. I've been using the-them a lot, a lot
- 19:42
for coding lately, and they work really, really well. So if you're using a Cursor or OpenCode, um, you can very easily in-incorporate our models to, um, uh, for coding work. Try them out, let us know. And there's a Discord channel, um, at the bottom. If you have any questions, you can post them there or you can, um, you know, um, message me or Dylan. We'll happy to, um, answer your questions. Uh, so I'll stop here. We have, like, ten seconds. So right on time. But, uh, we'll be around. Please come and talk to us if you have any questions. We'll happy to,
- 20:12
um, uh, chat more and then take any... take your questions. Thank you all.