AI Engineer World's Fair 2026
Is Speculative Decoding Worth It? Profiling vLLM on NVIDIA Blackwell — Akamai
Read the talk
Is Speculative Decoding Worth It? Profiling vLLM on NVIDIA Blackwell
A small model can draft several tokens before a large model checks them together. Sheilah Kirui’s side-by-side demo explains when that exchange saves time, when it wastes work, and why acceptance rate alone cannot decide.
From a talk by Sheilah Kirui
At a glance
Ideas worth remembering
Speculative decoding uses cheap sequential drafting and one target verification pass to advance several tokens when proposals are accepted.
A useful draft model balances speed with target agreement and requires memory for both its weights and additional KV caching.
The structured-output demonstration reported 1.6× faster generation with high acceptance; the creative case showed lower acceptance at high temperature.
Long inputs with short answers limit the overall gain because speculation accelerates decode while prefill still processes the context.
Spare VRAM, smaller batches and predictable continuations make speculation more promising; evaluate them together on the intended workload.
The expensive part is writing one token at a time
An application using a large language model can require hundreds of forward passes to produce an answer. Each generated token depends on the tokens before it, so the model keeps returning to the same expensive operation. Sheilah Kirui, Developer Advocate at Akamai, opens with this latency problem: when is it worth asking a smaller model to do some of the guesswork?
The request has two phases. Prefill processes the input tokens and builds a KV cache, which holds information the model uses during the rest of the request. Decode produces the answer, ordinarily one token at a time. This distinction matters because speculative decoding accelerates answer generation; it does not remove the initial work of reading the prompt.
The motivating example is a Llama 70 billion model. Repeatedly running a model of that size makes a long answer costly and slow. The opportunity is to reduce how often the large model must perform a separate generation step while keeping it responsible for which tokens become the answer.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Draft several tokens, then verify them together
A smaller draft model proposes a short continuation. It still generates autoregressively: each proposed token depends on the previous ones. Its advantage is that those sequential steps cost less. Kirui describes configuring three to five proposed tokens per cycle, then passing them to the larger target model for verification in one forward pass.
The target approves or rejects the proposals and supplies corrections for rejected tokens. Accepted proposals let one target pass advance the answer by several tokens. Rejections reduce that benefit because the draft model has spent time proposing text that cannot be kept. The target retains responsibility for the output; the recording explains this relationship without detailing the sampling and rejection procedure behind its claim that the output remains the same.
Where does the saved work come from? The flow below separates cheap sequential drafting from the target’s joint verification. Its useful relationship is that several draft steps feed one target pass. The benefit depends on how much of that proposed continuation survives verification.
Proposes three to five tokens sequentially.
Accepted proposals advance the answer; rejected proposals require correction by the target.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The draft must earn its memory and running time
“It’s not free.” Hosting a second model requires room for its weights and additional KV-cache storage for both models. Spare GPU memory therefore makes speculation worth considering. At high concurrency, the GPU may already be busy serving requests, leaving less opportunity for extra drafting work to help.
Draft selection balances two competing needs: the model must be cheap enough to run quickly and accurate enough to agree with the target. A very small model can make inexpensive guesses that are frequently rejected. A more capable draft can improve agreement while consuming more of the time and memory the optimization is supposed to save.
Kirui gives three practical selection criteria:
- Size: choose a model substantially smaller than the target; the stated rule of thumb is ten to fifty times smaller.
- Tokenizer: use the same tokenizer, ideally within the same model family, so the two models work with the same token representation. Different tokenizers introduce translation work.
- Cost and agreement: evaluate both how quickly the draft runs and how accurately it predicts what the target will accept.
The demo has a further constraint: both the baseline server and the speculative server run on one NVIDIA Blackwell GPU, with GPU utilization split between them. The stated weight footprints are about 16 GB for the baseline model and 2.5 GB for the draft model. Those are weight figures, rather than complete serving-memory requirements; the setup must also leave space for KV caching. The model choice makes this simultaneous comparison fit on the available hardware.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Structured output turns more guesses into useful tokens
Agreement depends on the task. Coding, JSON and SQL are the promising examples: their structure gives a small model more predictable continuations to propose. Poetry and brainstorming offer more possible directions, making agreement harder. This is a workload hypothesis to test, rather than a guarantee attached to a task label.
The demonstration turns that distinction into a side-by-side comparison. A small application has separate structured-output and creative-task tabs, and starts the baseline and speculative runs together. The first attempt stalls. After troubleshooting the vLLM servers, the results appear and Kirui reruns the structured case.
The structured-output run is reported as 1.6× faster with speculation, alongside a high acceptance rate and more generated tokens per second. Acceptance rate means the number of draft tokens accepted divided by the total number the draft proposed. This is a result for the demonstrated shared-GPU setup; the recording does not provide a complete benchmark configuration or numerical acceptance rate, so 1.6× should not be treated as a general Blackwell or vLLM speedup.
The causal chain explains why the structured example is useful. The draft proposes several tokens cheaply. High acceptance means much of that work survives the target’s check, allowing a verification pass to advance the answer by multiple tokens. The observed increase in tokens per second is the practical payoff. Acceptance rate helps explain that payoff, but generation speed still has to cover the cost of running the draft.
Switching to the creative case changes the agreement behavior: acceptance is much lower. Kirui attributes this to the wider variety of possible next tokens and the high temperature used for the task. More proposals are discarded, weakening the mechanism that made the structured run faster. The creative demonstration supplies a contrast in acceptance, rather than a quantified speed comparison.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A faster decode cannot remove a long prefill
Long-context applications introduce a different limit. Retrieval-augmented generation and document analysis may provide a large input before asking for a short answer. Much of the request then goes into processing the prompt and building the KV cache. Even a substantial improvement in answer generation changes only the smaller decode portion of the total wait.
The important comparison is the work spent reading versus writing. A short answer after a large document leaves relatively little generation to accelerate. A workload with more decoding work gives speculation more room to matter. Context length belongs in the decision alongside acceptance rate because it changes how much of the request the optimization can reach.
The closing checklist combines these independent constraints:
- Available VRAM: leave room for the additional model and caches.
- Task predictability: structured output is a stronger candidate than highly creative generation.
- Serving load: smaller batches are more promising; high concurrency can already occupy the GPU.
- Input and output balance: short-context, generation-heavy requests offer more opportunity than long inputs followed by brief answers.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Start with the serving engine, then explore other methods
The audience’s final question is practical: which tools support this? The demonstration uses vLLM, and Kirui recommends its documentation as a starting point. The separate draft-model setup is one relatively simple form of speculative decoding. The Q&A also names n-gram methods, Medusa and EAGLE as alternatives, without developing their mechanisms or comparing their performance.
For replication, Kirui points to Akamai Developers’ GitHub resources. Inference optimization is one part of that material; other contributors cover AI agents, managed Kubernetes and Akamai Functions. The useful next experiment follows the demo’s shape: compare a baseline with speculation on the task you actually serve, inspect acceptance and generation speed together, and account for the memory, context and concurrency conditions under which the result occurs.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
My name is Sheilah. I'm a DA at Akamai, and we're here today to discuss, um, the topic speculative decoding and how you can, um, figure out whether it's worth enabling for your workloads.
- 0:27
Okay, I'm a little anxious.
- 0:36
Okay, um.
- 0:39
So typically, when you ask LLM a question, there are usually two phases that happen under the hood. The first is the model would read your prompt, process all the input tokens, and then build what is known as a KV cache, um, which is essentially the model's working, um, memory, um, for the rest of the request, and this happens only once. Um, and then you have the decoding phase, which is, uh, when the model starts to actually write the answer. Um, so because every token depends on the previous tokens, um, this process usually happens one at a time.
- 1:09
Um, and so imagine you're deploying an application that uses a big model like Llama seventy billion model. Um, typically, like this would require like even up to hundreds, like forward passes on the model, and this is a very expensive process, right? Um, and it can cause a lot of latency for your user applications. Which, um, brings a question of like, what if you didn't have to use your bigger model to handle the token generation? What if you used a smaller
- 1:39
model to, um, do some of that guesswork to speculate on a couple tokens, um, which is still an auto-regressive process, but the whole point is because it's smaller, there's-- it's a faster process, so it would be a cheaper process. Um, the output is still the same because, um, the target model, the main model that has the accuracy still has to verify the tokens, and it does this in one forward pass. Um, so that's the idea behind speculative decoding. Um,
- 2:09
I'll just leave this here for a second. So as I said, you have a small model, um, which generates token. Typically, you can set it between three to five for each cycle. The target model would handle the predictions, um, like be able to scale, like basically approve or reject it. Um, and then for all the-- for the tokens that are rejected, the target model would,
- 2:33
um, recompute the, the token for that. Okay. So, um, this process, it's great. It can save you a lot in terms of inference generation, but it's not free. Um, because you're hosting now two models, you have to account for the memory required to host that second model. Not that much since it's a smaller model, but you also have to allocate extra space for, um, the KV cache for both models. Um, yeah, so typically, you would consider running
- 3:03
speculative decoding when you have, um, extra space in GPU. Um, for example, if your workload is high concurrency, your pro-- your GPUs are probably already busy trying to run all the requests. So it wouldn't make sense to implement speculative decoding. But if you have extra GPU space, it's worth considering, but I will talk about how to evaluate whether it's worth. Um, and the first question is, um, choosing the draft model. Um,
- 3:34
yeah. So typically, when you're ch-- de-deciding whether it's worth enabling, um, can the draft model predict as accurate as the target model? And then the other question is whether, um, it's as fast. So you would typically choose a model that's definitely like a fraction of the cost. Like, um, yeah, so these are the two main, um, decisions that you have to balance, right? Um,
- 3:58
I'm just gonna leave this here. Um, so for my situation, um, I will explain why I chose these two models. For my demo setup, when I was running this, I have access to one GPU. I was running on a Blackwell GPU, and I wanted to run both the baseline and the speculative decoding setup in one. Um, so to be able to do this, I had to choose a model that was small enough to be able to fit because I have to split the GPU utilization between the both servers. Um, so this is what that look like. For instance, you can-- the
- 4:28
baseline alone, um, takes about sixteen gigabytes of weight, and then the draft model is two point five, so that leaves a huge amount and space for KV caching. Um, and then when you're thinking about choosing the model, some of the things that you would usually consider is, uh, first of all, like I said, it has to be smaller, typically ten to fifty times smaller than your target model. Um, it has to have the same tokenizer,
- 4:59
um, ideally from the same model family. Otherwise, you know, you would have to like manually translate between the different tokenizers. Um, what's the other?
- 5:12
Yeah. Those are the main things. I think I'm forgetting one. Um, oh yeah, and like I said, the cost and accuracy. Uh.
- 5:24
Okay, so...
- 5:29
Oops.
- 5:36
So when I was running this, there's three questions. Um, speculative decoding, there's some workloads that would benefit from it versus others that wouldn't. Um, typically because the goal here is to get a smaller model that can predict almost as accurately as the target model would. It helps with some use cases when you have a workload that is more, like highly structured, like coding use cases. Um, what's another one? Like writing
- 6:06
JSON or SQL prompts, um, it would probably like benefit more versus if you were using it for a highly creative use case, like writing poetry or brainstorming or anything like that. Yeah, so.
- 6:21
I think I have...
- 6:26
This is a quick application. I'm not even sure if it's gonna work, but, um. So here I do have for the first demo, I just have, um, two tabs here. Um, like I said, this is just to focus on looking at how the acceptance rate performs depending on the type of task that you have. Um, so the first tab here is a structured output. Um, when I click this, it will run both at the same time.
- 6:53
Whoa.
- 6:57
Oh, God.
- 7:04
I'm just waiting.
- 7:24
Technical difficulties. I'm not sure why my demo is not running. Um, let's see.
- 7:34
Oh, do I have to reconnect? Um, just give me one second. I think I might have to restart the...
- 8:03
Fuck. Damn it. Looks like I need to restart the vLLM servers and my setup. Give me just a few minutes. Is this on? Um.
- 8:14
Um, LSOF minus I thousand. Okay, so it's running. Okay, that's running. Oh, okay. It's working now.
- 8:53
Oh, um.
- 9:00
Should I-- Yeah, I think I, I need... Oh, it's showing up. Okay. There we go. Um, so yeah, I was saying this is a demo. So when I run this... I'll redo it.
- 9:17
So you can see, um, as expected with the speculative on the right side, um, it was able to generate much, much faster, um, one point six times faster. Um, the key thing to note here is how high the acceptance rate is, which is the rate of the tokens that were accepted by the-- the number of tokens that were accepted over the total generated by the draft model. Um, another metric here, um, you can see with the
- 9:46
speculative enabled, um, you-- we are able to generate a lot more tokens per second, um, which improves the throughput. And then if I run it for the other case.
- 10:09
So here it's another use case.
- 10:14
Um, as expected, same thing. Um, but in this case, I guess the main metric to see here is the low acceptance rate. And like I said, this is because this is a more, um-- There's more variety in how the next tokens that would be generated because first of all, the temperature is set high. It's a creative use case. Um, yeah. Um, what else? I think that's pretty much for this, um,
- 10:43
case. And then the other situation where...
- 10:49
Oops.
- 10:52
Oh, yeah. This is what I was-- I should have put this up here, um, just to explain when it would benefit. And then the other use case would be context length. Um, so as I mentioned earlier, there's two parts of generating tokens. Um, with speculative decoding, it's supposed to accelerate the generating part, not the part where the model reads the prompt. So typically, if you have a use case where you need to provide the model with a lot of context, like for RAG applications or, um, analyzing lots of documents,
- 11:22
um, with those situations, um, the model will spend a lot more time with the portion of like building the KV caching. Um, especially if you're using it to do like a quick-- where the, the generated tokens are much less than the context, like the input. Um, yeah. So ba-- I think the main point here is just how to think through like what is your application doing. Um, is it highly structured? Do you have enough VRAM
- 11:53
space to even consider it? Are you doing more creative scenarios? Are you running a lot of concurrency, um, or not? Like smaller batch sizes would benefit from this. Um,
- 12:07
yeah.
- 12:11
Yeah, I'm gonna leave this slide here. This just basically summarizes, uh, what I just went through, and if anyone has questions, I'll take one or two questions now.
- 12:27
Ooh, I can't hear you. Oh, yeah. Um, the question was, are there particular tools that I recommend for speculative decoding? So I implemented this with vLLM, um, the serving engine. Um, they have really good documentation on how to get started with it. This is the simpler-- one of the simple versions of speculative
- 12:57
decoding. There's a couple others like the n-gram, Medusa, EAGLE that's more, like, the architecture beneath is, like, more advanced. Um, but for getting started, yeah, I would say you can't go wrong with just reading up. I-- When I was doing this, there's, there's not much content out there on this, but I did find a few, like, blogs. Uh, when you res-- when you type in speculative decoding, um, one of the first blogs that will come up by General Compute talks about more in detail. Um, there's also, like, two research
- 13:27
papers that were published, uh, by Google and what's the other company? Yeah. It-- But that's more, like, goes, like, more in-depth. Yeah. Um, but I'm also trying to maybe, like, make more content around this, like, maybe, like, videos. It, it was kinda hard to go through this. I w-- I'm a little all over the place, but I would definitely love to maybe, like, publish a blog or make snippet videos, um, yeah, on, like, YouTube or somewhere. But you can definitely find the resources.
- 14:03
Okay, yeah. All right, I should have covered that. Um, but I don't have the-- I should have put a QR code in.
- 14:12
Okay. Oh, um, as I said, I-I'm from Akamai, and if you're interested in, like, learning how to replicate this, you can find it on our website, our GitHub, GitHub, um, Akamai Developers.
- 14:29
Yeah. Um, yeah, so we have a, a lot of, like, different resources on here. Um, I'm focusing on infr-inference optimizations. Uh, we have other people that are focusing on building AI agents, um, and how to leverage our, like, managed Kubernetes service, how to use Akamai functions. So if you have time, please go check out our GitHub pages, and also we have a Discord channel that we're trying to build up now. Um, so it would be great to chat with some of you on the side. Thank you.