Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher
Read the talk
LLM Inference at Scale: From KV Memory to Serving Engines
Harshul Jain and Tanmay Sah explain why context, concurrency, and token generation strain GPUs, then connect each bottleneck to model and serving optimizations.
From a talk by Harshul Jain and Tanmay Sah
At a glance
Ideas worth remembering
A model fitting in GPU memory is only the starting point. KV state grows with context and concurrent sequences, while latency objectives can reduce usable concurrency below the memory ceiling.
Prefill and decode demand different optimizations: prefill performs substantial prompt computation, while decode repeatedly accesses weights and KV history for each new token. TTFT and inter-token latency must be assessed separately.
Weight quantization frees space for KV state; KV quantization reduces the state itself. Neither memory reduction alone proves higher throughput or preserved quality.
Paging reduces allocation waste, continuous batching reduces unused scheduling capacity, and prefix caching avoids repeated prompt computation. Each mechanism targets a different source of inefficiency.
The corrected latent-attention saving is 14-fold for the workshop example. The earlier 50–56-fold calculation omitted a layer multiplier, showing why comparisons must include the full architecture.
The presenters report similar vLLM and SGLang performance on standard requests and a three-to-four-times SGLang advantage in their agentic branching setup. Their recommendation is workload-specific, and they explicitly allow for different results elsewhere.
Choose the model and service requirements before selecting optimizations. Evaluate useful token throughput under those constraints, then pursue deeper cache or distributed-inference techniques only when they address a bottleneck the workload actually has.
Why inference becomes an operating-cost problem
Harshul Jain and Tanmay Sah open with a first-principles workshop for beginners and intermediate practitioners. The sequence follows a practical question: what makes inference expensive, what causes its bottlenecks, and which optimizations address those causes? They divide the interventions into changes to the model and changes to the system that serves it, before comparing inference engines.
The economic distinction is between training expenditure and recurring inference expenditure. Serving continues to consume resources as users arrive, sessions begin, and tokens pass through the system. Jain uses market estimates and a historical training-cost example to motivate the problem, but the useful engineering conclusion does not depend on those estimates: limited hardware and expensive computation make usage growth an operating-cost concern. Reducing token usage and making inference more efficient are the two responses he emphasizes.
The accompanying repository gathers slides, a benchmark report, and prepared Jupyter notebooks. Its purpose is to make scattered inference material easier to study and experiment with. The workshop's larger goal is similarly durable: understanding the underlying constraints should help practitioners evaluate subsequent optimizations instead of treating each new solution as an unrelated trick.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Three symptoms in a simple inference implementation
The first demonstration loads a roughly 15 GB model and examines GPU memory consumption. Jain initially walks through previously obtained results because of conference connectivity concerns. Loading the weights establishes a baseline, but inference consumes additional memory as the input grows. Contexts of 4,000, 16,000, or 32,000 tokens can therefore create an out-of-memory problem even when the model itself fits comfortably.
The second symptom is rising time to first token, abbreviated TTFT. A longer prompt makes the user wait longer before generation begins. Jain explicitly corrects his terminology here: the changing variable is context size, meaning the number of input tokens, rather than the size of an individual token. Memory growth and startup latency respond to the same increase in prompt length, although their mechanisms will differ.
The third symptom is poor throughput, measured as tokens or users served per second. In the simple implementation, five incoming requests are answered sequentially. Later requests wait for earlier work, so completion time increases with multiple users. These three observations—memory, TTFT, and throughput—establish the questions the rest of the workshop answers.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Attention explains why memory grows with context
The inference pipeline begins by converting text into tokens, mapping those tokens to embeddings, and passing them through transformer layers. The workshop temporarily treats one word as one token for simplicity; this is a teaching assumption. A generated token becomes part of the input for the next generation step, creating an autoregressive loop. The Mistral 7B example uses 32 transformer layers, while other models can have different layer counts.
Inside a transformer layer, Jain identifies normalization, attention, and feed-forward computation, then focuses on attention. Each token is projected into query, key, and value vectors. Attention determines how a token relates to preceding tokens, which requires access to their keys and values. Ten tokens need ten sets of these representations; a thousand tokens need a thousand. Keeping the keys and values therefore creates storage that grows with sequence length.
The KV calculation multiplies the two stored vectors, their 128-dimensional size, the 32 layers, the number of KV heads, and the storage precision. For the workshop's configuration, Jain gives approximately 131 KB per token. That becomes roughly half a GB at 4K context and 2.1 GB at 16K context. His spoken concurrency example associates about 42 GB with 4K context and 80 users. The supplied description instead associates 42 GB with 16,000 tokens and 80 users; those figures are inconsistent. The per-token estimate and spoken 4K example support the central mechanism: independent concurrent sequences multiply KV demand, and a 24 GB GPU cannot accommodate a 42 GB cache.
GPU capacity divides into model weights, runtime overhead, and the remaining space available for KV storage. The weights are fixed for a loaded model; overhead is treated as approximately fixed in this simplified accounting. The notebook gives about 14.6 GB for the model at 16-bit precision. Once those allocations are accounted for, longer contexts and more users compete for the same remaining memory. Supporting 160 users, for example, requires a shorter context than supporting fewer users under otherwise unchanged conditions.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Prefill spends compute; decode repeatedly moves data
Prefill processes the prompt, builds its keys and values, and computes attention relationships across its tokens. This is matrix-heavy work with substantial computation, so Jain describes it as compute bound. In the workshop's simplified timeline, prefill determines time to first token: more input tokens require more projections and attention work before the first output appears.
Decode generates subsequent tokens one after another. Each step computes attention for the new token while needing the preceding tokens' key and value representations. Compared with processing an entire prompt, this provides less computation per step. The time between successive outputs is inter-token latency, the fourth metric introduced in the workshop. It captures how quickly a response continues after its initial wait.
To explain the memory bottleneck, Jain distinguishes large GPU memory from smaller, faster shared memory. In his high-level account, computation loads chunks into shared memory, performs the math, and writes results back. Decode repeats this process for successive tokens. Its speed can therefore be limited by how quickly weights and KV data reach the computation, even when the GPU has additional arithmetic capacity.
Arithmetic intensity expresses this distinction as floating-point work per byte transferred. Decode moves weights and preceding KV vectors to compute an output for one new token, giving it relatively low arithmetic intensity. Prefill performs more computation over the transferred data and has higher intensity. The roofline discussion uses this relationship to distinguish bandwidth-limited work from compute-limited work.
The timing demonstration shows prefill duration increasing with input length. Decode timings remain nearer an average after excluding cold-start effects, but Jain cautions that they are not constant. A longer history still means more keys and values must be pulled from memory, so decode latency can also increase with context.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Capacity must satisfy a latency target
A memory-based concurrency estimate divides available KV memory by KV storage per user. With the model, GPU, and KV precision fixed, context length and concurrent users become the adjustable variables. Increasing concurrency requires reducing context at the memory limit, and that reduction may harm answer quality by limiting the information a request can carry.
Fitting the maximum number of users does not mean serving them acceptably. A business also has a latency service-level objective. More users and larger batches can increase decode time and affect TTFT, making useful capacity lower than the memory ceiling. Jain frames the resulting decision around quality, latency, and throughput. Premium chat favors quality and responsiveness, accepting fewer users per GPU; asynchronous agent work can favor quality and aggregate throughput because its tasks already run over longer periods.
The capacity calculator combines GPU memory, bandwidth, compute capacity, and hourly cost with model and workload choices. Jain selects a 7-billion-parameter model at FP16 and begins by fixing the requirement that matters most: latency for premium chat, or a minimum batch size for asynchronous work. His chat example considers a 10-millisecond latency target while preserving context for quality. These are illustrative calculator inputs, rather than a demonstrated service guarantee.
Hourly price alone can lead to the wrong GPU choice. A more expensive device may deliver a lower cost per million tokens if its useful throughput is sufficiently higher. The calculation must therefore follow the workload constraints: estimate the concurrency and context that remain viable at the required latency, then compare the resulting token economics. Jain emphasizes that inaccurate capacity estimates undermine this comparison.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Quantization reduces the weight allocation
Sah introduces two deliberately fictional teaching algorithms. The ostrich algorithm stands for waving through an assumption, such as treating compression as lossless. The world cup algorithm stands for dividing a large problem into smaller pieces and carrying useful results forward. Their role is to make the optimization decisions memorable; they are not numerical methods or evidence that an assumption is valid.
His weight-sizing example asks how a 120-billion-parameter model could fit on one 80 GB H100. At two bytes per parameter, the illustrative weight allocation is 240 GB. Reducing precision to eight bits gives approximately 120 GB, which still does not fit. Sah then describes a four-bit representation with an approximate 65 GB footprint. The example shows why the required compression ratio follows from the memory budget, although the exact model representation is presented tentatively.
For Mistral 7B, Sah discusses INT8, INT4, and NF4 alternatives to 16-bit weights. Fewer bits reduce storage, but preserved quality is an assumption that must be tested. He calls for external benchmark evaluation and distinguishes post-training quantization from incorporating quantization during training or fine-tuning. The decision is therefore both a sizing exercise and an assessment of the numerical changes the workload can tolerate.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Attention variants change how much KV state is needed
Sah uses a 4096-by-4096 matrix to illustrate splitting work. Dividing its columns into groups of 128 produces 32 blocks, providing an intuition for multiple attention heads and parallel computation. The example explains decomposition; the subsequent memory savings come from changing how many key and value representations those query heads use.
Multi-query attention represents the extreme of sharing one KV set across multiple query heads. Sah describes this informally as retaining one block instead of 32, assuming that queries can use the shared representation. Grouped-query attention takes a middle position: query heads form groups that share KV representations within each group. This reduces KV storage without taking sharing to the single-set extreme. The proposed quality preservation remains something to establish for the model and task.
Multi-head latent attention offers another route: compress key and value information into a latent vector and reconstruct the representations needed for attention. Sah notes that positional information complicates this approach, particularly its interaction with rotary positional encoding, or RoPE. The explanation identifies the need to preserve a positional component but does not derive the reconstruction or positional treatment in full.
Sparse attention changes which tokens receive attention. Instead of attending to the entire preceding sequence, the proposed strategy attends to selected important tokens. This targets the amount of attention work, rather than only the size of its stored representations. Sah presents it as an evolving direction; the workshop does not establish how those important tokens are selected or quantify the resulting quality tradeoff.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
FlashAttention uses tiles to reduce intermediate memory traffic
FlashAttention addresses repeated movement between large GPU memory and the computation. Sah contrasts a sequence of calculations and writes back to memory with processing smaller tiles of the query and key matrices in fast local memory. Keeping running state for online softmax allows attention to proceed tile by tile. The important mechanism is that breaking up the computation also changes where intermediate work stays, reducing the need for repeated large-memory transfers.
The storage comparison gives a concrete grouped-query example: reducing 32 KV heads to eight yields fourfold compression of that KV allocation. Latent-attention savings depend on the architecture, including layer count and latent dimensions. Sah gives a much larger preliminary savings estimate while expressing uncertainty about the dimensions, so it is not established here as a general compression ratio.
The attention scorecard remains qualitative. Sah describes grouped-query quality as close to multi-head attention while stressing that the use case matters, and describes single-set multi-query sharing as a stronger quality compromise. He also briefly introduces sliding windows, linear attention as summarizing information before lookup, and Mamba as a state-space model. These alternatives receive an orientation rather than a detailed implementation or measured comparison.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Memory savings do not establish throughput gains
Jain returns to the quantization notebook. After a weight-download delay, he reports approximately 15 GB at FP16 and 7.5 GB with INT8, illustrating a twofold reduction in weight memory. The freed capacity can support longer KV histories or additional concurrent users. The four-bit discussion gives varying approximate footprints, so the stable conclusion is the direction of the weight reduction rather than a single exact four-bit measurement.
The plot that follows contains theoretical throughput numbers rather than a throughput test performed in this demonstration. Jain also notes that some benchmarks they studied showed lower throughput with INT8 compression. Smaller storage therefore cannot be taken as proof of faster execution: it expands memory capacity, while actual serving speed remains an empirical question.
The attention notebook encounters a GPU-detection problem, limiting the live demonstration. Jain nevertheless makes a substantive correction to the earlier latent-attention calculation: the claimed 50–56-fold saving should be 14-fold for their example. The previous computation omitted the number-of-layers multiplier. This corrected figure is an example-specific storage comparison with multi-head attention, rather than a universal latent-attention ratio.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
KV caching, paging, batching, and prefix reuse
The serving discussion begins with the waste in recomputing earlier tokens' keys and values during an uncached generation loop. A KV cache retains those vectors and references them in later steps, trading memory for avoided computation. The improvement concerns reuse of projections already computed; the new token still needs access to the preceding KV state for attention.
Paged attention addresses allocation waste. Jain's toy example reserves 2 KB for a request that uses only 1 KB, wasting half the allocation. Requiring contiguous storage can also prevent useful free space from accommodating another request. The operating-system analogy separates logical order from physical placement: a sequence's KV state appears logically contiguous while mapping to different physical memory blocks. Blocks are allocated as new tokens require them, allowing memory to follow actual request growth.
Continuous batching addresses wasted scheduling capacity. In a fixed batch, new work waits until all its requests finish, even when some finish earlier. Continuous batching allows new requests to enter as capacity becomes available, keeping more useful work on the GPU. Paging concerns where state resides; batching concerns when requests can use the computation.
Prefix caching extends reuse across requests that share an initial token sequence, whereas the basic KV cache reuses state across generation steps within a request. KV quantization reduces the storage used by the cached vectors themselves. The resulting capacity can accommodate more tokens, longer contexts, or more requests; the workshop does not demonstrate that quantizing those vectors automatically improves answer quality. Jain presents these serving mechanisms as capabilities available through vLLM rather than features practitioners need to rebuild.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The serving benchmark separates speed from cache footprint
The benchmark keeps Mistral 7B as the model and sends a set of prompts through successive server configurations. Helper functions check server readiness, collect metrics, and measure KV usage. Jain says the full process takes around an hour because configurations require stopping servers, restarting them, and loading weights. These are reported benchmark results rather than a full live rerun.
On an H100, the plain Hugging Face baseline delivers around 51 tokens per second. Jain also reports first-token and inter-token latency values, but their units are not stated in the spoken passage. Moving to the default vLLM server, with paged attention, continuous batching, and KV caching, produces nearly 15 times the throughput in their comparison. He reports higher TTFT and lower inter-token latency, illustrating that a large throughput gain need not improve every latency measure.
Adding prefix caching raises throughput further and reduces TTFT while leaving inter-token latency approximately unchanged in the reported test. The description of its KV-usage changes is less definite, so it does not establish a precise memory benefit. Adding KV quantization leaves throughput and both latency measures approximately similar but lowers KV usage. This is a useful distinction: an optimization can improve the memory footprint without producing an immediate speedup on the tested workload.
The speculative-decoding configuration likewise produces approximately similar results in Jain's summary, with somewhat lower KV usage. Its inclusion does not establish a general decoding speedup. It sets up the next discussion of why draft generation and verification depend on how well their predictions agree.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Speculative decoding depends on accepted drafts
Sah explains speculative decoding as drafting several tokens with a smaller model, then asking the larger model to decide how many to accept. His example proposes four or five tokens per draft. The loop repeats, attempting to advance generation by more than one token at a time. Agreement between the draft and verifying model is central: draft work provides little benefit when few proposed tokens are accepted.
He suggests that predictable domains, including repeated code syntax, may suit this approach, but says his own testing did not find basic speculative decoding useful. That is a personal, setup-dependent result rather than a demonstration that the method never works. The workshop does not supply acceptance rates or a detailed cost breakdown that would explain the outcome quantitatively.
The related methods change how candidates are produced. Sah describes self-speculation using an auxiliary head within the main model, EAGLE using a trained predictor informed by the main model's internal features, and Medusa proposing tokens in parallel. He personally favors EAGLE over the alternatives discussed. These are brief mechanism sketches, without enough comparative measurements to establish a ranking across workloads.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Branching prefixes and hardware-specific engines
Sah contrasts hash-based prefix lookup with a radix-tree approach. Hash lookup reuses stored KV state when the relevant input matches; changing a word or letter can break that match. A radix tree represents shared prefixes and their branches, compressing paths that do not branch. Its value is reuse of the unchanged beginning of related sequences, rather than treating differently worded prompts as interchangeable.
Agent loops make shared prefixes especially relevant. Sah gives the example of a software-engineering role prompt repeated 200 times as a workflow loops. Shared instructions and related continuations create opportunities to retain reusable state while representing divergent branches separately. He identifies SGLang as using radix-tree prefix caching for this kind of reuse.
TensorRT-LLM enters as another inference engine. Sah distinguishes the broader TensorRT SDK from the LLM-serving engine and emphasizes its connection to NVIDIA hardware. Its approach includes optimizing layers and execution at the hardware level. This adds another engine-selection consideration: the fit between an execution stack and its target hardware.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Standard requests and agentic branching produce different comparisons
The presenters compare vLLM and SGLang on an H100 using questions from ShareGPT. For this standard request workload, they report no statistical difference between the engines, with similar requests per second, TTFT, and latency. The passage does not give the statistical procedure, sample size, or uncertainty intervals, so the finding applies to their reported comparison rather than proving general equivalence.
Their agentic test changes the interaction pattern. A first turn asks for a proposal to solve traffic congestion in a city. A second turn asks the model to review the proposal and rate it from 1 to 10, and this two-turn pattern repeats. Standardized prompts and context create the branching workload they want to test. Sah reports SGLang performing three to four times better in that setup, while explicitly warning that different setups may yield different results. The spoken summary does not specify one exact metric for that multiplier.
Jain's recommendation is to begin with vLLM for a standard API workload and consider SGLang when agentic behavior makes the default unsatisfactory. He also points to a separate 120-billion-parameter-model comparison that includes TensorRT-LLM, but does not report its numerical results here. Engine choice consequently remains tied to the request pattern, model, and hardware rather than a single winner for every deployment.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Choose a baseline, then deepen the bottleneck that matters
The closing engine survey mentions NVIDIA Dynamo in connection with agentic session routing, a simple Hugging Face path without a server, and research into multimodal serving. Jain then returns to the deployment sequence: establish a baseline, choose a model that meets the use case, reduce its memory footprint where appropriate, and use a serving engine to obtain the required throughput. He cautions against treating the small workshop model as an automatic production choice.
The next level of study concerns KV eviction, cache compression, and hybrid memory. These topics extend the question from how to retain reusable state to which state should remain resident and how it should be stored. Jain urges practitioners to connect each proposed solution to the problem it actually solves, then determine whether that problem exists in their own workload. The workshop names these directions without deriving their policies or implementation tradeoffs.
Distributed inference is left as a separate substantial subject, requiring its own treatment of internals and hands-on work. The presenters describe a proposed advanced workshop, invite feedback and expressions of interest, and close by offering to discuss questions offline. The recording therefore ends with a clear boundary: it builds the foundations of inference optimization, while distributed execution and deeper cache engineering remain further work.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:13
Uh so good afternoon everyone. Um my
- 0:16
name is Hershel Jan and he is Tanisha.
- 0:20
Uh and we would like to welcome you all
- 0:22
in this two hours workshop on the LLM
- 0:25
inference. Uh so the goal of this
- 0:28
workshop is to understand this domain
- 0:32
from the first principles uh dive deeper
- 0:35
into it and like understand what's going
- 0:38
on throughout the industry.
- 0:41
Uh a bit of background about us. So I am
- 0:45
a senior software engineer at Audible.
- 0:48
uh have been building MLA data platforms
- 0:51
for the past five years and on the sides
- 0:54
I have been writing this opensource
- 0:56
handbook on LLM inference
- 0:59
and Tanme he is the senior quantitative
- 1:03
modeler at XAN cup bank corporation he
- 1:06
recently completed his PhD and he has
- 1:09
been actively doing research in the
- 1:12
agent verifiers and the world models
- 1:18
Uh so a quick show of hands here. Uh Vu
- 1:22
here is like brand new to the LLM
- 1:24
inference.
- 1:27
Okay, great. And Vu here has like
- 1:29
deployed these models in production.
- 1:33
They have been tuning it. They have been
- 1:35
serving the production traffic.
- 1:38
Okay, great.
- 1:41
So this workshop is targeted towards the
- 1:44
beginner and the intermediate level. U
- 1:48
and all of the slides and exercises they
- 1:50
are in the repo. I will share that soon.
- 1:55
Here is the quick agenda for the
- 1:57
workshop. We will start with the problem
- 2:00
statement. We will try to understand few
- 2:02
of the pain points around LLM inference.
- 2:06
uh then we understand what causes those
- 2:08
pain points and build our foundations
- 2:10
from there.
- 2:12
Then we will dive into like two kind of
- 2:15
the optimizations that we do like the
- 2:17
model optimizations and the serving
- 2:19
optimizations.
- 2:21
uh and then we start learning about
- 2:23
different serving engines that are
- 2:25
available to deploy our LLM inference
- 2:29
solutions in production and we will
- 2:31
showcase some benchmarks and the
- 2:33
decision chart on like which engine to
- 2:36
use.
- 2:40
Cool. So to understand the pain points
- 2:43
first we need to know what is like LLM
- 2:46
inference. So, and probably a lot of us
- 2:49
already know this. Um, but yeah,
- 2:52
anything that you ask your AI to do like
- 2:55
whether it be generate a video, audio,
- 2:58
analyze any text, uh, analyze your
- 3:01
medical reports or like your tax bills,
- 3:04
all of that is like an LLM inference.
- 3:08
And this market is like approximately
- 3:11
$23 billion today.
- 3:15
uh semi analysis recently shared that if
- 3:18
you want to model like a Google search
- 3:20
queries with LLMs, you need like a
- 3:23
profit drain of like $36 billion
- 3:27
and query cost has to be less than 0.5
- 3:31
cents to keep your search business
- 3:33
profitable.
- 3:35
On the other hand, the business insider
- 3:38
mentioned like your AI has to be put on
- 3:40
diet
- 3:42
and everyone has to start auditing and
- 3:44
budgeting their token usage and all of
- 3:47
this is happening. Why? Because your
- 3:50
hardware is limited, compute is
- 3:52
expensive, your inference is expensive
- 3:55
and with the growing need of like more
- 3:58
and more AI usage, this inference cost
- 4:00
is rising more and more.
- 4:03
So
- 4:05
this stat, it's an old stat from the
- 4:08
open AI, but it's still it's still true.
- 4:11
So if you look at the like training cost
- 4:14
of the GPT3, it was like around $4.6
- 4:17
million. It was a one-time cost. But if
- 4:20
you see the inference cost that has been
- 4:23
like uh it's a recurring cost because
- 4:26
it's a operating cost that scales with
- 4:28
every user that comes in that every
- 4:30
token that comes in every session that
- 4:32
uh is being initiated on the like AI
- 4:39
and there are only two ways to basically
- 4:42
counter this. Uh one way is you reduce
- 4:45
your token usage.
- 4:47
um alternative is you should try to
- 4:50
optimize your inference solutions as a
- 4:53
inference service provider for your
- 4:56
customers and for yourself. And so this
- 4:59
um we have been seeing like lot and lot
- 5:02
of like new solutions coming out every
- 5:04
then and now. Um and so the idea would
- 5:08
be like okay we will try to build those
- 5:10
foundations that will help us understand
- 5:13
and evaluate like whatever ships next.
- 5:17
Um so yeah to get started like we will
- 5:21
do a quick demo like it's a short demo
- 5:24
of like what are the different pain
- 5:25
points around inference
- 5:28
u and so this is the repo uh I mean you
- 5:33
can pull it or you can also open it on
- 5:35
the GitHub
- 5:37
uh it's called LLM inference at scale a
- 5:39
bit of background here like four months
- 5:42
back when I didn't knew anything on the
- 5:43
LLM inference um I started learning it I
- 5:47
saw like lot of resources were
- 5:48
scattered. So we started putting it uh
- 5:51
like all together in one place uh so
- 5:53
that it could benefit people.
- 5:56
Um yes. So let me actually get out of
- 6:00
this slideshow mode and probably
- 6:04
go into
- 6:08
I will go to this extended mode.
- 6:13
Um
- 6:16
okay great.
- 6:19
Uh yeah so in this uh repository if you
- 6:23
see a readme file there is like a link
- 6:26
to the slides. Uh so this it will be
- 6:29
like this folder where you have like a
- 6:31
pptx and there is like a benchmark
- 6:33
report in there.
- 6:36
uh you can always like download it and
- 6:39
then for the demo purposes uh we have
- 6:42
couple of Jupyter notebooks. Uh we have
- 6:44
like collaborated with Moab who are the
- 6:47
like Google collab alternative
- 6:50
and what they basically provide you is
- 6:52
like a free RTX 6000 GPU. It's a 100 GB
- 6:56
V RAM GPU.
- 6:58
So
- 6:59
and we have like already set up these
- 7:01
notebooks so that it becomes easy to
- 7:04
like experiment with and like all of the
- 7:07
assets and everything are preset for
- 7:09
you. Uh
- 7:12
so uh we will start with like a simple
- 7:15
demo a
- 7:18
probably
- 7:21
let me just see
- 7:29
Okay.
- 7:31
Um yeah. So
- 7:35
when it comes to the inference, you need
- 7:37
to do an inference on a certain model,
- 7:40
right? Uh for the workshop purposes, we
- 7:42
are using a simple ML 7B model. Uh it's
- 7:46
a small model of around 15GB in size. So
- 7:49
we we are going to like load that into
- 7:51
the GPU. So,
- 7:55
and we would look like some of the GPU
- 7:57
stats as well. So, we see like okay, we
- 7:59
are working on the 6,000 Blackwell. Uh,
- 8:02
and you might be thinking I'm not
- 8:03
running the cells because I don't trust
- 8:05
the Wi-Fi at conferences.
- 8:08
So, yeah. So, I would probably be just
- 8:11
going over the results uh that we kind
- 8:13
of ran previously.
- 8:18
Uh, yeah. So, we have like a GPU which
- 8:20
is like 102GB.
- 8:23
Uh now the first thing that comes to my
- 8:26
mind is like what's my memory
- 8:27
consumption looks like when I do the LLM
- 8:29
inference. So I load this model
- 8:33
uh and I see like okay I have like a 15
- 8:36
GB here. So I have roughly like 87.5GB.
- 8:40
And now when I do the like inference
- 8:43
here
- 8:45
uh what I notice is like the more the
- 8:48
number of inputs I pass more is the
- 8:51
memory that I need.
- 8:53
uh and it's increasing slowly but it's
- 8:55
still increasing. [snorts] So imagine
- 8:57
like if you have a context length of
- 8:59
like around 4,000 or 16,000 or 32,000
- 9:03
uh tokens. Uh so this memory could like
- 9:06
really grow big and it you could
- 9:08
actually get like all of those out of
- 9:10
memory issues.
- 9:13
Uh so definitely this is like your
- 9:15
problem one like your memory increasing
- 9:17
with the increase in tokens. So in form
- 9:19
of like a simple visualization it looks
- 9:21
like this.
- 9:25
The second problem that you would see is
- 9:27
like
- 9:28
the time to your first token it's very
- 9:31
very slow. Uh we measure it by a metric
- 9:34
called TTFT. It's a short short form of
- 9:37
it. uh and when you try to like measure
- 9:41
the TTFT uh with the like input size you
- 9:46
would see like longer the context
- 9:49
you would see like this uh TTFT being
- 9:52
slow. So now there are two problems.
- 9:54
Your memory increases with the token
- 9:56
size. Your TTFT increases with the token
- 9:59
size.
- 10:00
Uh sorry not the token size, the context
- 10:02
size.
- 10:06
Uh and then the third is the like
- 10:08
throughput. The throughput is like how
- 10:11
many tokens can you serve per second and
- 10:14
then how many users can you serve per
- 10:16
second. So if you take a very very
- 10:18
vanilla implementation on your local
- 10:20
system
- 10:22
uh it would be like very sequential. So
- 10:24
if you send like five requests all those
- 10:27
five requests would be catered like
- 10:29
sequentially rather than parallelly.
- 10:32
Uh and so like your request basically
- 10:36
takes more time to complete if you have
- 10:38
like multiple users.
- 10:41
So these are the like three problems.
- 10:43
There is a fourth one. I haven't
- 10:44
described it here. probably we will
- 10:46
build that intuition as we move forward.
- 10:50
Uh but let's remember like these are the
- 10:52
three problems the memory TTFT and the
- 10:55
throughput.
- 10:58
Cool. Uh I will go back to the slides.
- 11:07
Okay, perfect.
- 11:15
Uh so it should be this.
- 11:20
>> Uh is it visible?
- 11:35
Oh, you you need the
- 11:38
Oh, okay. Workshop. Okay.
- 11:47
Yeah. So, within that repository, if you
- 11:50
see a workshop folder, you see that
- 11:53
readme and then the readme has all the
- 11:55
links, the slides and the demos.
- 12:04
Uh, does that work?
- 12:09
Okay, perfect.
- 12:16
Okay. uh so let's start working through
- 12:18
the foundations like let's start
- 12:20
understanding
- 12:22
uh what are the reasons behind those
- 12:24
pain points and for that like we have to
- 12:27
look at this inference pipeline
- 12:29
um
- 12:32
so we get like an input text uh that
- 12:35
text could have like any number of words
- 12:39
you convert those into the tokens so for
- 12:42
simplicity you can assume one word equal
- 12:45
to one token
- 12:47
uh then you kind of convert them into
- 12:49
like the embeddings and then you send it
- 12:51
to the like transformers
- 12:54
uh like there are 32 layers of
- 12:56
transformers but that's specific to the
- 12:57
Mistl 7B different models have different
- 13:00
kind number of layers uh and then you
- 13:03
generate a new token and that token
- 13:05
basically goes back to the input then
- 13:07
you generate another token and that
- 13:08
keeps on going now in this entire
- 13:11
pipeline you would see like 95% of your
- 13:15
compute is like taken by these
- 13:17
transformer layers. So it's worth
- 13:19
looking at like what goes within this
- 13:21
transformer layer.
- 13:25
Within this transformer layer you would
- 13:26
have like more layers. You have like a
- 13:29
normalization layer, you have an
- 13:31
attention layer, you have a feed forward
- 13:34
layer and all.
- 13:35
And attention layer is the one I think
- 13:38
that has been very very famous.
- 13:40
Attention is all you need paper. I think
- 13:42
that's very well known. So attention is
- 13:45
the most compute intensive layer and we
- 13:49
need to understand what goes within that
- 13:51
attention layer.
- 13:54
So what does attention do? Attention uh
- 13:58
so if you have an input text it needs to
- 14:00
find the attention scores of every token
- 14:03
with respect to all of the previous
- 14:04
tokens.
- 14:06
And to do that what it needs to do is
- 14:09
like it needs to project every token
- 14:11
into like a key query and the value
- 14:13
space. So in like in a simpler terms
- 14:18
just u understand this like if you have
- 14:22
10 tokens then it needs like the 10
- 14:24
different query key and the value
- 14:26
vectors. If there are 100 tokens you
- 14:29
would need 100 key and the value
- 14:30
vectors. If there are thousand tokens
- 14:32
you would need thousand key value
- 14:34
vectors. And so like your number of the
- 14:36
key and the value vectors they increase
- 14:38
as you increase the input size.
- 14:44
Uh and if you calculate the like KV size
- 14:48
per token
- 14:50
uh for a mist 7B it comes out to be 131
- 14:53
KV. Uh this is because like
- 14:57
uh you have two vectors K and V. You
- 15:00
have to multiply the size uh one vector
- 15:03
is like 128 dimensions. you have to
- 15:05
multi multiply it by 32 transformer
- 15:08
layers and then you have to multiply it
- 15:10
by the KV heads for ML 7B it's gate KV
- 15:15
heads it's not like 32 because it uses a
- 15:18
different kind of an attention mechanism
- 15:21
u which we will talk about for sure but
- 15:25
yeah so the KV size per token is like
- 15:29
your 131 KB now imagine if you have 4K
- 15:33
context uh So that size becomes like
- 15:35
half a GB. Uh if you do like 16k context
- 15:40
that size becomes 2.1 GB. Uh now
- 15:44
multiplay by the users like assume you
- 15:46
can serve multiple users together
- 15:50
at the same time within that GPU you
- 15:53
could have like 42GB with a 4K context
- 15:56
and 80 users. Um and if your GPU is only
- 16:00
like let's say 24GB, you are already
- 16:03
running out of the memory. So you cannot
- 16:05
serve that many users with that many
- 16:07
context.
- 16:12
To visualize this, look at a GPU memory.
- 16:15
So the GPU memory has like a model
- 16:17
weights which are pretty fixed. These
- 16:20
are pre-trained weights. Uh there is
- 16:23
like an overhead that is also fixed.
- 16:25
that also like that changes but it does
- 16:27
not change that much u overall you can
- 16:30
assume it's fixed and then there is like
- 16:33
a leftover memory so this leftover
- 16:36
memory is what being used by your KB me
- 16:39
like key and the value vectors so
- 16:44
assume like you have a one user you can
- 16:48
only serve that many key and the value
- 16:51
vectors or that many tokens which can
- 16:54
like fit in this entire 80GB uh like
- 16:58
memory that is left.
- 17:02
Uh
- 17:04
so we can show this with a simple demo
- 17:07
too. Uh
- 17:13
okay, let me
- 17:18
Okay, great.
- 17:23
Okay, great.
- 17:25
Uh, let me see if I can actually run
- 17:29
this.
- 17:31
What
- 17:36
probably?
- 17:38
Where the heck is this happening?
- 17:49
Okay, great.
- 17:52
Yeah. So you would see like the GPU is
- 17:55
attached.
- 18:07
So here we are just trying to confirm
- 18:09
the like memory based on the maths and
- 18:11
based on the intuition that we have
- 18:13
built. So the model memory is like let's
- 18:16
say if you have a 7 billion parameters
- 18:18
you are doing a 16 bit precision your
- 18:21
total memory comes out to be 14.6 GB
- 18:23
with you can basically verify that with
- 18:26
the maths. So if you do all that math
- 18:30
that comes out to be the 14.6GB.
- 18:33
Uh now comes the KVK and the KV size. So
- 18:38
this KV size is like your 131 KB per
- 18:40
token.
- 18:42
Uh and if you do that maths and you try
- 18:45
to like visualize this.
- 18:49
[sighs]
- 18:52
Oh, sure.
- 19:00
Wait.
- 19:03
Okay. And then let's just visualize
- 19:07
this.
- 19:09
Okay. Great.
- 19:11
Uh yeah so this is the like a memory
- 19:14
chart. So if you see like as your
- 19:17
context increases your memory keeps
- 19:19
increasing.
- 19:21
Then another thing to realize is like as
- 19:24
your users increase
- 19:26
uh then also your memory increases. So
- 19:29
if you want to serve like 160 users
- 19:33
uh on a GPU you can support like uh you
- 19:37
can only support like a lesser context
- 19:39
length. So there is always a tradeoff
- 19:42
between what context length you can
- 19:44
serve versus how much cost you can save
- 19:48
by like putting your multiple
- 19:51
uh users or the concurrent users into
- 19:53
like a single GPU. So you have to always
- 19:56
take that tradeoff and and we will go
- 19:58
through that like uh in couple of more
- 20:01
slides.
- 20:06
>> Uh can you repeat please?
- 20:14
Uh I'm sorry I'm cannot hear you.
- 20:24
>> Yeah.
- 20:33
Cool.
- 20:35
Uh okay. Great. So let me pull back.
- 20:41
So that was like memory. Uh we need to
- 20:44
understand uh why we had like uh slower
- 20:48
time to first token
- 20:51
uh when we increase the context length.
- 20:54
So for that like we need to understand
- 20:56
the two phases of inference and those
- 20:59
phases are like the prefill and the
- 21:00
decode phase. I think you would all seen
- 21:03
like a lot of articles but we just
- 21:05
wanted to explain it. Uh so when you
- 21:09
send like a lot of uh like when you send
- 21:11
these input tokens
- 21:14
what you want to do is uh you want to
- 21:16
build those key and the value vectors
- 21:18
that I mentioned for all the tokens.
- 21:21
Then you want to compute the attention
- 21:23
scores of every token with respect to
- 21:26
the previous token. All this operation
- 21:28
that you do it's a very very metricsh
- 21:31
heavy uh it's a very very compute heavy
- 21:34
operation and we all know like the GPUs
- 21:37
they are like very well suited for a
- 21:40
heavy compute workload so we call like a
- 21:44
prefill to be like a compute bound uh
- 21:47
and it does take some time to complete
- 21:50
so whatever time that this phase takes
- 21:52
to complete that's your time to the
- 21:55
first token
- 21:57
So if if you have like a more input
- 21:59
tokens, you have to generate more key
- 22:02
value vectors. You have to do lot more
- 22:04
attention math and because of that your
- 22:08
TTFT becomes more more slower.
- 22:12
Where is if once you generate one token
- 22:15
you need to keep doing this to generate
- 22:17
another tokens sequentially one after
- 22:20
another. But in that process every time
- 22:23
you have to build the key and the value
- 22:26
vectors of all the previous tokens
- 22:29
which is same as prefill like you were
- 22:31
building key value vectors there also
- 22:33
here also but in decode phase you are
- 22:36
only computing the attention math for
- 22:38
the new token
- 22:41
and that is why it's a very less it's
- 22:44
lesser compute oriented
- 22:46
and it's also called as memory bound. We
- 22:48
will see it shortly why it's called as
- 22:50
memory bound.
- 22:54
So in a classic timeline you would see
- 22:58
prefill and decode phase like this. So
- 23:01
time taken by prefill that's your time
- 23:03
to first token and then your time taken
- 23:06
by every decode step that's your
- 23:10
uh basically your inter token latency.
- 23:13
So that's like the fourth metric
- 23:16
uh that you need to worry about like
- 23:19
what's the time being taken by your
- 23:21
decode step.
- 23:25
Okay, cool. Uh
- 23:28
now why why does the like decode step or
- 23:32
why does decode takes time and why it's
- 23:36
being called as like a memory bound
- 23:38
operation? Let's try to understand that.
- 23:40
uh to understand that we need to look at
- 23:44
how the metrics map basically works on
- 23:46
the GPU on a high level. So GPU has two
- 23:49
kind of memories. You have a high
- 23:51
bandwidth memory. You have a shared
- 23:54
memory.
- 23:56
So the high bandwidth memory is a larger
- 23:58
size but a lower me like lower
- 24:00
bandwidth.
- 24:02
By lower bandwidth I mean like you can
- 24:04
transfer data out of it at a lower rate
- 24:07
compared to the shared memory. So the
- 24:09
shared memory is smaller in size but it
- 24:12
has a very very high bandwidth. Uh that
- 24:15
means you can transfer data in and out
- 24:17
of it with a very first thing. So when a
- 24:20
when you have to do a metric math so you
- 24:23
have to pick the data in chunks from the
- 24:26
high bandwidth memory you have to put it
- 24:28
into the shared memory.
- 24:31
Do that math write back the result into
- 24:33
the high bandwidth memory.
- 24:37
uh for the prefill phase when you have
- 24:42
to do this you have to do this matrix
- 24:45
math only once but for the decode phase
- 24:48
you have to do this metric math uh like
- 24:53
uh again and again because you're
- 24:55
generating each and every token
- 24:57
sequentially
- 24:59
and so like you it doesn't matter like
- 25:02
how fast is your decode
- 25:05
because now you can transfer your data
- 25:08
out of the high bandwidth memory into
- 25:11
the S shared memory at a certain speed
- 25:14
because you are limited by the high
- 25:16
bandwidth memory bandwidth speed
- 25:19
and so that governs your like token
- 25:22
sealing like at what rate can you
- 25:24
actually generate tokens out of the
- 25:27
decode step.
- 25:32
If you look at this in the roof line
- 25:35
plot
- 25:37
uh so there is a left section which is
- 25:41
called to be a memory bound.
- 25:43
Mathematically it's governed by the
- 25:45
arithmetic intensity.
- 25:47
Arithmetic intensity is the number of
- 25:50
flip-flop operations that you perform
- 25:52
per bite of data being transferred. So
- 25:55
for the decode step
- 25:58
uh decode step since you are
- 26:01
transferring lot of data like the key
- 26:04
and the value vectors of the all the
- 26:06
previous tokens the model weights but
- 26:08
you are doing the like less computation
- 26:10
because you're computing attention math
- 26:12
for only one token. Uh so it's
- 26:16
arithmetic intensity is very low but for
- 26:19
a prefill phase you are transferring the
- 26:21
data once but then like you are doing
- 26:25
this heavy computation and so it's
- 26:28
arithmetic intensity is very high. So
- 26:30
now you know like in terms of
- 26:32
mathematics like why the computer like
- 26:36
why the arithmetic intensity of prefill
- 26:38
is very high compared to your decor.
- 26:43
Uh [snorts]
- 26:44
okay
- 26:46
so this is like another small small
- 26:49
demo. Uh
- 26:52
every time I have to Okay.
- 26:57
Okay, great.
- 27:00
I hope this is already running. So yeah,
- 27:03
again we are loading the model.
- 27:07
Now this is the like the prefill cost.
- 27:10
So what we are basically doing is uh we
- 27:12
are getting the like um the input uh
- 27:15
text and then we are trying to generate
- 27:18
this um the prefill step the amount of
- 27:22
time it takes. We see like as we
- 27:24
increase the like size of the input
- 27:26
tokens this prefill is increasing. So
- 27:29
you and this is the reason why your TDF
- 27:32
increases
- 27:34
and then like your decode time. So the
- 27:37
decode time is like on average it stays
- 27:40
about the same. Uh and so if it is like
- 27:44
assuming like you ignore the like cold
- 27:46
start your decode time is like
- 27:48
approximately around the average line.
- 27:50
it it it it is still impacted by like u
- 27:56
the input size. It's not like it's a
- 27:58
constant uh time and it is because it
- 28:02
still needs to pull the key and the
- 28:03
value vectors from the memory for all
- 28:05
the previous tokens. So there is still
- 28:08
like uh that uh basically small increase
- 28:12
in time that you would see with the
- 28:14
decode step. And then this is the like
- 28:17
classic uh roof line plot.
- 28:21
Uh okay.
- 28:27
Presentation.
- 28:31
Okay. Five. Okay. Great.
- 28:35
Okay. So
- 28:38
now now let's try to understand like uh
- 28:41
the throughput dimension. You want to
- 28:44
understand how many users you can
- 28:46
actually serve and I think we saw like a
- 28:49
diagram of the GPU memory where we saw
- 28:52
okay there is some memory that is free
- 28:54
for the key and the value vectors to
- 28:55
grow.
- 28:57
So assume like you have just a single
- 29:00
user
- 29:02
uh
- 29:04
what's the total KV size that you have
- 29:07
you can basically support it's defined
- 29:09
by your context limit.
- 29:12
uh the max users that you can support is
- 29:15
like whatever is your GPU uh
- 29:18
availability like whatever is the memory
- 29:20
that is available in the GPU you divide
- 29:22
it by the key and the value size per
- 29:24
user uh and when you do that like it
- 29:28
comes out to be like your the concurrent
- 29:30
users.
- 29:33
Now assume like your GPU is fixed, your
- 29:38
model is fixed.
- 29:40
Uh
- 29:41
so
- 29:43
your KV size per token is fixed. There
- 29:46
are only two dimensions that are left
- 29:48
here which is context and your
- 29:51
concurrent users.
- 29:53
If you want to serve more concurrent
- 29:55
users, you have to reduce the context
- 29:57
length. If you reduce the context
- 29:59
length, you could impact your quality.
- 30:02
Uh so these are the two dimensions right
- 30:05
now that we are trading off.
- 30:09
Then if we but can you actually serve
- 30:12
the like max number of concurrent users?
- 30:17
Uh in an ideal world probably not
- 30:20
because
- 30:22
every business has like a latency SLO
- 30:25
that we have to meet.
- 30:28
So
- 30:30
if you remember like in the decode step
- 30:32
I said the time for the decode still
- 30:35
increases if you have more inputs.
- 30:38
It also increases if you have more
- 30:39
users.
- 30:41
So ultimately
- 30:44
uh your inter token latency also gets
- 30:47
impacted
- 30:49
if you have like a higher batch size and
- 30:51
your TTF also gets impacted. So now
- 30:54
there is a third dimension you have to
- 30:56
worry about which is like your latency.
- 30:58
So the three dimensions that you have is
- 31:01
like a quality latency and the
- 31:03
throughput. So it comes out to be like
- 31:05
this trade-off triangle where you have
- 31:08
to choose between the two. So for a
- 31:12
premium chat application
- 31:15
you would want to prioritize definitely
- 31:17
the quality and you want to prioritize
- 31:19
the like the latency. You would not want
- 31:22
your users to wait infinitely for the
- 31:25
like or like not infinitely but probably
- 31:28
for the larger latency.
- 31:31
You can always sacrifice the number of
- 31:33
users you can support on the GPU and
- 31:35
probably take that costed being more
- 31:38
customers
- 31:40
in form of like and and like if you
- 31:43
consider like an agent uh sorry the
- 31:46
async agent workload you would want to
- 31:49
like prioritize definitely quality and
- 31:50
the throughput
- 31:52
uh because these are the longunning
- 31:55
tasks
- 31:56
uh and you would want to like serve as
- 31:59
many as concurrent tasks. fast as
- 32:01
possible but with a very very higher
- 32:03
quality.
- 32:07
And often like we think like okay if the
- 32:11
GPU is like a very expensive GPU
- 32:15
uh that might not be a good fit for us.
- 32:19
Uh but it turns out that could actually
- 32:21
serve you the lowest cost per million uh
- 32:24
tokens.
- 32:27
Uh but you really have to trust your
- 32:30
kind of calculations on the max users
- 32:33
that you want and like uh you really
- 32:36
have to make those estimations uh
- 32:39
correctly.
- 32:43
Uh
- 32:49
so we do have like uh
- 32:55
let me just
- 32:59
Where is this?
- 33:02
Okay, great. So, for the capacity
- 33:05
calculator, uh there is like a link to
- 33:08
the collab because I was facing certain
- 33:11
issues with molab. I had to migrate out
- 33:13
the wall widget library and I didn't
- 33:16
have time. So, being lazy, I just picked
- 33:19
collab there. Uh apologies to Moab.
- 33:25
Uh
- 33:27
so my VR is connected.
- 33:35
Okay.
- 33:46
Wi-Fi probably.
- 33:51
Okay. Great.
- 34:02
So what we have done over here is we
- 34:04
have like shaded some like the GPUs with
- 34:07
their V RAMs, bandwidths, the flip-flops
- 34:10
and the cost per hours. Um
- 34:14
then we kind of like built this simple
- 34:17
uh like uh capacity calculator. This is
- 34:20
just a KV visualizer uh where you kind
- 34:23
of like when you increase the number of
- 34:25
tokens uh you see like your KV size it
- 34:28
increases and when you increase the
- 34:31
number of users your size is like
- 34:33
increasing at a much faster rate
- 34:37
and then
- 34:40
in this capacity calculator uh
- 34:46
let it run.
- 34:49
So we have like a model which we which
- 34:52
is like a 7 billion parameter model that
- 34:55
we selected.
- 34:58
We set the like precision to be FP16. Uh
- 35:01
now we decide the way we basically go by
- 35:04
the GPU decision is you have to decide
- 35:08
what's your like you have to fix one
- 35:11
dimension first which you care about the
- 35:13
most.
- 35:15
for premium chat I mentioned like
- 35:17
latency is definitely the one
- 35:21
uh and then like for the async workloads
- 35:23
the batch the minimum batch size that
- 35:25
you want to serve for from like a single
- 35:28
GPU that is the second dimension so you
- 35:31
want to fix these first so I will go
- 35:34
about like in a premium chat application
- 35:38
uh
- 35:39
so I can go ahead with like 10
- 35:41
milliseconds latency a minimum batch
- 35:45
size I don't care like I can so I'm okay
- 35:47
with like probably two
- 35:50
uh
- 35:53
okay so probably with the seven
- 35:56
concurrent users on a single GPU and
- 35:59
then like my context limit is very
- 36:01
important to me because I want to focus
- 36:03
on the quality as well
- 36:06
uh and so like I do see like some of the
- 36:09
GPUs so the H18GB
- 36:12
it's like a $8 per hour but like am I
- 36:18
300x is it? Yeah. So it's like around
- 36:21
$10 per hour but if you do all that
- 36:26
throughput math that we shared in the
- 36:28
mathematics before you could find like
- 36:31
your cost per million dollar tokens that
- 36:34
could be very very that could be like
- 36:36
lesser. So
- 36:39
you need to do such calculations by
- 36:41
fixing those dimensions and you need to
- 36:44
decide your GPU to like reduce your kind
- 36:47
of inference cost. This is at least the
- 36:50
first step that you can take towards
- 36:51
optimizing the inference.
- 36:56
Okay, cool.
- 37:01
So the next slide. So let me
- 37:08
Okay, great.
- 37:10
And so like now the next thing is about
- 37:13
the model optimization. So we are now
- 37:15
basically have built that foundation
- 37:17
where we understood some of the pain
- 37:19
points, reason behind those pain points,
- 37:21
why those were happening
- 37:23
um how we could like address that GPU
- 37:28
capacity thing. We need to understand
- 37:30
what can we do like what can we further
- 37:32
do about it. So it it is about the model
- 37:36
optimization and I think I would like to
- 37:38
invite Tan I he can talk more about
- 37:40
these model optimizations provided he
- 37:43
has worked uh on this like during his
- 37:46
research times
- 37:48
okay I can control
- 37:53
yeah here okay hi everyone uh mic check
- 37:58
am I audible at last yeah okay so hi I'm
- 38:02
Tesha I work as a senior quant modeler
- 38:05
and also I am an AI researcher. My work
- 38:08
focuses on a agent verification and
- 38:10
right now building world models. So for
- 38:13
this one model optimization
- 38:16
before we start model optimization so I
- 38:19
created a research template so that it
- 38:21
will be easy for us to understand all
- 38:23
these complex things. I so our template
- 38:27
is simple. First we will identify the
- 38:29
problem. Second step we will solve the
- 38:32
problem using two algorithms. These are
- 38:34
just fake algorithms. So first algorithm
- 38:36
is called ostrich algorithm. Whenever we
- 38:40
see uh just like ostrich whenever we see
- 38:42
a problem ostrich put their head into
- 38:45
the sand. So same thing we will do
- 38:47
whenever we face a problem we will just
- 38:50
ignore it. So this is an important
- 38:52
algorithm we should follow. Second one
- 38:55
is created it is called world cup
- 38:57
algorithm. For example, we don't know
- 39:01
who will win this FIFA World Cup. So,
- 39:03
what organizers did, they uh break the
- 39:08
48 teams into 12 groups, uh then round
- 39:12
32. So, round 32 right now is currently
- 39:14
going on. Uh then round 16, then
- 39:17
quarterfinals, uh then semi-finals and
- 39:20
finals. So what they are doing is that
- 39:24
uh they are breaking it into a smaller
- 39:26
problems and the useful results are
- 39:29
moving forward. So same analogy or same
- 39:32
algorithm we will use uh to understand
- 39:35
this model optimization all those
- 39:37
things. So yeah let's start. So I have
- 39:43
one H100 GPU.
- 39:46
I have to use this open-source model
- 39:49
what is called GPTOSS
- 39:51
120 billion parameter model. So right
- 39:55
now I think it's so they have trained it
- 39:57
on BF float 16 and weight is 240 GB.
- 40:02
What should I do?
- 40:04
This is the problem we have. So first
- 40:07
thing what we have to deal do is that
- 40:10
240 GB and 80 uh GB H100.
- 40:16
So and I have to fit only in one GPU or
- 40:19
not in multiple GPU. So what can we do?
- 40:22
I think simple step is that just
- 40:25
compress it. But how should we compress
- 40:29
it? Uh that's the another challenge. So
- 40:31
if we compress BF BF float 16 to FP8 uh
- 40:36
then it will be around 120 GB but our
- 40:39
GPU H100 is still 80 GB. So what I think
- 40:43
they did is that they compressed it into
- 40:45
further MX uh MX FP4 and I think size is
- 40:51
around 65 GB. So this is something we
- 40:55
can do uh compress but question so and
- 41:00
we will use over this ostrich algorithm
- 41:03
we are assuming that uh there is no loss
- 41:05
in compressing a bigger model into a a
- 41:09
smaller size. Second thing
- 41:12
in uh in this one okay yeah so in this
- 41:17
one in this slide we have used this
- 41:18
mistral 7B so 7 billion parameters so
- 41:22
it's a small model 7 billion parameters
- 41:24
so uh so if you multiply it by two bytes
- 41:27
so it so weight of it's around is 14
- 41:31
14.5 GB which can easily fit into H100
- 41:35
or even a a40 so
- 41:39
so Next uh what we can do is that like
- 41:43
mistral 7B instead of compressing it a
- 41:46
floating point 16 we can apply different
- 41:49
techniques like int 8 or int4 or nf4. So
- 41:53
basically we have to just use ostrich
- 41:55
algorithm and just believe that uh there
- 41:58
is no quality loss kind of things but
- 42:01
somehow we also have to mathematically
- 42:03
prove that by doing some kind of test
- 42:05
testing on some external benchmark that
- 42:07
whether it is working or not. So the and
- 42:12
this comes under post training
- 42:14
quantization kind of thing. One can also
- 42:16
do uh this one uh during finetuning one
- 42:20
can also do this kind of quantization.
- 42:22
This comes under a quant training kind
- 42:24
of thing. So uh let's move to our next
- 42:28
problem.
- 42:31
So
- 42:34
we have this huge matrices just just
- 42:38
imagine imagine uh 1,000 by 1,000 uh
- 42:43
dimension matrix A and another matrix
- 42:47
matrix um 1,000 by 1,000. So if we
- 42:51
multiply uh if we multiply by this two
- 42:54
matrices so number of operations will be
- 42:57
1,000 raised to the power q and this is
- 43:01
kind of a problem in terms of uh uh in
- 43:05
terms of computing. So we wondered our
- 43:07
matrix multiplication should be fast and
- 43:10
it should save memory. So what should we
- 43:14
do? We have a giant matrix. Okay, let's
- 43:18
take this one. Uh, Mr. 4096 by 4096.
- 43:22
What should we do uh to
- 43:25
solve our problem of speeding up the
- 43:28
things and saving the memory 4096 by
- 43:31
4096.
- 43:33
So first thing is that we will use just
- 43:35
our world cup algorithm. We can decide a
- 43:38
random number just break the block
- 43:41
vertically. It does not matter what you
- 43:43
are choosing it. So you have so let's
- 43:47
say uh we have 4096 uh columns we will
- 43:51
break it uh we will break it into a
- 43:54
group of 128 column each. So 128, 128, 128
- 44:00
128 uh vertical vert uh vertically so we
- 44:04
will get a 30 we will get this 32 blocks
- 44:07
if we divide this 4096
- 44:09
then
- 44:11
what will happen by doing this thing? So
- 44:13
if we just divide this one vertical
- 44:16
vertically then we can use a multiple
- 44:19
GPU to speed up the process. So this
- 44:22
kind of thing is called multi head
- 44:25
attention.
- 44:27
So what else can we do? We have a big
- 44:30
matrix like
- 44:33
uh as I have mentioned that ostrich
- 44:35
algorithm. So our main problem is
- 44:39
sizing. So what we what we can do is
- 44:42
that instead of having all those 32 uh
- 44:45
32 vertical blocks we will throw away uh
- 44:48
31 blocks and we will assume that one
- 44:51
block is sufficient enough that all the
- 44:54
queries uh can handle those blocks. Our
- 44:58
loss will be almost negligible and we
- 45:02
come up with this algorithm uh which uh
- 45:04
and this algorithm is called a
- 45:06
multiquery attention. So as we can see
- 45:10
right now we are at two spectrum. One is
- 45:12
multi head attention where we split it
- 45:16
into 32 blocks and use different uh
- 45:20
different uh GPUs or do some parallel
- 45:22
processing and at the same time we are
- 45:25
just throwing 31 blocks and uh we are
- 45:28
calling this is as a multi-query
- 45:31
attention. So
- 45:34
uh so at both extreme we should be come
- 45:36
up with a middle ground like something
- 45:39
we can say that
- 45:41
instead of throwing all the 31 uh maybe
- 45:44
we can group we can group we can group
- 45:46
some of the blocks together so that uh
- 45:50
uh and we can assume that uh similar
- 45:53
blocks will attend to a um similar kind
- 45:57
of uh queries. So this kind of technique
- 46:00
comes under grouped query attention
- 46:02
which is very popular right now. Uh even
- 46:04
in uh even in mistral or in other models
- 46:08
this grouped query attention works. So
- 46:11
right now we have understand that we
- 46:14
have a big matrix uh we can divide it
- 46:16
the way we want and doing some
- 46:18
mathematical calculation prove that loss
- 46:21
is almost negligible kind of thing. So
- 46:23
what else we can do?
- 46:26
So after that uh after this grouped
- 46:29
query attention
- 46:32
uh
- 46:34
see
- 46:36
uh we have a big matrix uh one is one is
- 46:40
key and one is value.
- 46:42
Let's compress that matrix into a latent
- 46:46
vector and then come up with some
- 46:49
algorithm to uh reconstruct from latent
- 46:52
vector uh to our original matrix. So
- 46:55
this kind of a strategy comes under this
- 46:58
one um multi head latent latent
- 47:01
attention but again it has some problems
- 47:04
with rope because rope is position
- 47:06
dependent and uh and it is position
- 47:08
independent kind of thing. So yeah one
- 47:11
needs to also include some uh index for
- 47:13
keys also so that one can map it. But
- 47:17
again main problem is that why why we
- 47:21
are why we are multiplying all those big
- 47:24
matrices. So because that's how this
- 47:27
attention mechanism works that
- 47:30
each token will pay attention to every
- 47:33
token. So how about let's don't pay
- 47:36
attention to all the previous token only
- 47:38
pay attention to the important tokens uh
- 47:41
which is important for us. So this is a
- 47:44
kind of uh this kind of field is uh
- 47:47
evolving. So this comes under sparse uh
- 47:49
deepseek sparse attention. So
- 47:53
uh yeah and yeah yeah yeah so okay next
- 48:00
yeah so next one is flash attention. So
- 48:04
uh so in flash attention so main pro so
- 48:07
main problem is that uh
- 48:11
uh so so currently so so currently not
- 48:14
currently so right now almost everyone
- 48:16
uses flash attention but way in 2022 or
- 48:19
2023 uh so that's how it works that's
- 48:23
how it works is that uh so
- 48:28
uh this
- 48:29
Q K query and A and key matrices they
- 48:33
were in HBM. Uh it loads uh it uh first
- 48:38
uh it loads into uh this one uh tensor
- 48:41
core and it do some uh it do some
- 48:44
calculation and then it will uh write it
- 48:46
back to uh HBM and then uh this process
- 48:49
goes on multiple times. So in flash
- 48:52
attention uh what they did is that
- 48:56
uh is that instead of multiplying the
- 48:58
whole matrices so they just divided it
- 49:01
into like our world cup algorithm
- 49:03
divided the bigger matrices into a small
- 49:06
tile and only put those small tiles uh
- 49:08
into a SBM so that uh it can process
- 49:12
multiplication fast and just uh keep uh
- 49:16
keeping track of this some three
- 49:17
variables so that they can calculate
- 49:19
this online softmax.
- 49:23
Yeah.
- 49:25
Next one. So, yeah. So, so this is just
- 49:28
mathematics. So, if we have a multi head
- 49:31
attention if it is 524
- 49:34
uh KV uh then it depends upon how much
- 49:38
how much grouping we want and so if
- 49:42
instead of 32 KV head we only want to
- 49:46
use uh 8 KV heads. So uh so so we can
- 49:50
get a compression of 4x times and this
- 49:53
multi head latent attention this formula
- 49:56
depends on the model to model how many
- 49:59
layers your model have. So in the
- 50:01
original deepseek paper uh I think they
- 50:03
have some 128 dimension
- 50:07
128 d 12 I don't remember the exact
- 50:10
dimension but according to that uh they
- 50:13
have used uh this one latent vector in
- 50:16
which they have used 512 as a dimension
- 50:19
and some 64
- 50:21
for for rope index. So and then they
- 50:25
show that it is a 50x 56x
- 50:29
uh more compressed than multi head
- 50:32
attention.
- 50:37
Okay.
- 50:39
Yeah. So uh so yeah so this is uh so
- 50:42
this is the uh this is the trade-off uh
- 50:45
trade-off diagram. So here I think we
- 50:46
have not talked about this linear
- 50:48
attention or mamba. So main problem is
- 50:52
just all this m Matrix multiplication.
- 50:55
Right now everyone is using attention.
- 50:57
Suppose in future
- 51:00
uh if we don't want to use attention or
- 51:03
rather than generating tokens
- 51:05
sequentially just use maybe diffusion
- 51:07
models where we can generate everything
- 51:09
simultaneously. So all these algorithms
- 51:12
will change also. But here I think they
- 51:15
have two more. One is linear attention
- 51:17
and one is mamba. So according to uh
- 51:20
this slide so
- 51:23
if we are not compressing anything so
- 51:25
MHA is just we are parallelizing the
- 51:28
process so there is no quality loss so
- 51:31
it's a good and then this uh grouped
- 51:34
query attention which is I think almost
- 51:36
uh every model is using uh just GQA and
- 51:40
DSA kind of thing or yeah
- 51:45
I think same thing we are providing in
- 51:46
the attention mechanism scorecard So uh
- 51:50
so I think uh this one mha quality is
- 51:53
good throughput is uh throughput is okay
- 51:56
and for grouped query attention it
- 51:59
depends upon your use case also though
- 52:02
yeah though
- 52:04
quality is almost similar to uh multi
- 52:07
head attention but use case also matters
- 52:09
a lot yeah multi-query attention is just
- 52:12
one extreme we are
- 52:15
I don't know why but we are just
- 52:17
assuming that we only need one block and
- 52:20
all the queries will attend to that
- 52:22
smaller smaller block. So, so quality is
- 52:26
not that great for M for MQA and this
- 52:30
multi head latent attention. So yeah if
- 52:34
you have tried some this deep seat
- 52:36
models so I think uh they are doing
- 52:38
great job yeah in in quality wise
- 52:42
besides that sliding window so all these
- 52:45
are sub techniques which
- 52:48
yeah yeah all these are some techniques
- 52:50
like I just slide the windows all those
- 52:53
things and instead of yeah instead of
- 52:56
multiplying everything so linear
- 52:59
attention is just saying that sum
- 53:00
summarize everything first uh and then
- 53:03
look up into it and then mamba this is
- 53:06
just a state space model. Yeah,
- 53:12
I can cover that. Okay.
- 53:15
Uh cool. Uh thank you T.
- 53:18
So for the model like optimizations we
- 53:22
also have like the two notebooks
- 53:25
here.
- 53:28
So there will be
- 53:32
I have to go to this.
- 53:42
Okay. Uh so for the quantization uh like
- 53:47
the demo
- 53:49
uh this is is this already run? No. Let
- 53:53
me just run this.
- 54:03
Okay. So we are loading the model which
- 54:05
is like uh ML 7B.
- 54:12
Uh so
- 54:17
this one is like with the FP16 baseline.
- 54:29
Wait.
- 54:31
Uh, did it run?
- 54:35
Okay. So, it's uh two millisecond run.
- 54:40
Did this run? Okay. So, yeah, this time
- 54:43
it's fetching that model with the FP16
- 54:47
precision.
- 54:57
the Wi-Fi.
- 55:03
It's going to take time.
- 55:07
Okay.
- 55:11
>> Yeah, it because it's downloading the
- 55:13
weights from the hugging face.
- 55:17
>> Huh.
- 55:20
Yeah. So, MOLAB is like running online.
- 55:23
>> Yes.
- 55:26
because it needs to make the network
- 55:27
call through to the hugging phase and
- 55:29
like it fetching
- 55:32
I don't know like but it's taking time
- 55:34
to download probably
- 55:52
Okay.
- 55:55
So good. Okay. So here we see like the
- 55:59
memory size is like 15 GB around
- 56:02
approximately with the FP16 precision.
- 56:05
We are trying to do the 2x compression
- 56:08
as Tmet talked about with the int8.
- 56:13
Let's download. Okay. So we do see like
- 56:16
your memory size is now like 7.5 GB.
- 56:21
What that means is now you have a more s
- 56:23
more memory for your KV to basically
- 56:26
grow. That means you can either serve
- 56:28
higher context limit or you can serve
- 56:30
the higher concurrent users there.
- 56:36
If you do the like in your basic you are
- 56:39
doing the 4x compression so that with
- 56:42
the 4x compression it would be more
- 56:45
lower. It would be I think around
- 56:48
3 to 4 GB.
- 56:51
Yeah. 4.5 GB
- 56:55
and Yep. So this is Wait.
- 57:02
So this is just a basic plot
- 57:05
of like
- 57:07
so these are the like theoretical
- 57:09
numbers. uh we are not doing the like
- 57:11
any throughput test here but uh usually
- 57:13
you would see like your memory increases
- 57:15
so pro you would also have like a bit of
- 57:18
higher uh throughput.
- 57:21
Uh from some of the benchmarks that we
- 57:23
studied we saw like the intate uh
- 57:26
compression it does have like a lower
- 57:28
throughput.
- 57:32
Okay. And then there is like a demo on
- 57:36
the like the attention mechanisms.
- 57:42
So for the attention okay I have to run
- 57:47
this.
- 57:58
Uh okay so it has run. Oh, wait. Why
- 58:04
does it say no GPU detected?
- 58:09
It should say the GPU should be
- 58:11
detected.
- 58:18
Oh, okay.
- 58:37
Wait,
- 58:50
but this is surprising.
- 58:57
Yeah, I guess it's not like able to
- 58:59
detect the GPU for some reason.
- 59:04
Uh we do have like a GPU here.
- 59:11
Uh okay, never mind.
- 59:14
Yeah. Yeah. So, but the like basic idea
- 59:16
here was more like
- 59:20
as you try to move towards like
- 59:22
compressing the computation like by
- 59:25
using different attention mechanisms
- 59:27
like moving from the multi head to the
- 59:30
grouped query attention and then to the
- 59:33
MLA you would start seeing some
- 59:35
optimizations.
- 59:38
Um I think yesterday night we were doing
- 59:41
some benchmarking. Uh I wanted to
- 59:44
correct this part. Uh so it wasn't like
- 59:47
50 56x it was 14x. Uh basically the demo
- 59:52
had a mistake of like a computation uh
- 59:55
where it did not multiply the number of
- 59:58
layers.
- 1:00:01
Uh yeah so apologies for that. Uh so
- 1:00:04
this MLA is like a 14x savings work in
- 1:00:08
comparison to like your multi head
- 1:00:10
attention.
- 1:00:14
Uh so now that we have understanding of
- 1:00:18
the pain points, the foundations, the
- 1:00:21
one side of the optimizations which is
- 1:00:23
the model optimizations,
- 1:00:25
we want to talk about what can you do on
- 1:00:29
the like the serving side.
- 1:00:33
So
- 1:00:36
the first thing is we saw like when you
- 1:00:38
perform like a simple decode step you
- 1:00:42
are pulling it you are basically pulling
- 1:00:44
the model weights and then you are
- 1:00:45
recomputing the key and the value
- 1:00:47
vectors for all the previous tokens even
- 1:00:51
though you already computed the those
- 1:00:53
vectors for the tokens.
- 1:00:56
So there is definitely like a lot of
- 1:00:59
compute wastage.
- 1:01:01
Uh and if you kind of analyze the time
- 1:01:04
complexity of it, it would come out to
- 1:01:06
be O of N². Uh and the way to resolve
- 1:01:10
that is like a classic trade-off against
- 1:01:11
the memory. You can maintain a memory of
- 1:01:15
those vectors against the tokens and you
- 1:01:18
can reference that memory. So that
- 1:01:20
memory was called as like KV cache.
- 1:01:24
uh and the like the flow looks something
- 1:01:26
like this
- 1:01:30
and then based on this KV cache there
- 1:01:32
were like four optimizations that were
- 1:01:34
really possible.
- 1:01:36
Um
- 1:01:38
the first one is about the page
- 1:01:41
detention. So what's the different
- 1:01:44
what's the problem today? So when you
- 1:01:46
send like multiple requests as the input
- 1:01:48
to the GPU
- 1:01:50
these requests are in a batch
- 1:01:54
uh every request is allocated like a
- 1:01:57
continuous memory storage let's say of
- 1:02:01
I'm just taking an example like let's
- 1:02:03
set uh 2 KB
- 1:02:06
however like your request needed only
- 1:02:08
let's say
- 1:02:10
uh 1 KB so there is like u 50% of that
- 1:02:16
memory fragmentation.
- 1:02:19
Uh and this fragmentation basically
- 1:02:22
leads to the memory wastage. That means
- 1:02:26
there was a space in the memory where
- 1:02:28
you could have served more requests but
- 1:02:31
you could not because you were looking
- 1:02:33
for that contigious block of the memory.
- 1:02:36
So an inspiration to was being taken
- 1:02:39
from like how the OS works like you
- 1:02:42
maintain a logical memory and you
- 1:02:44
basically have a physical memory.
- 1:02:47
So in the logical memory it would still
- 1:02:51
feel like
- 1:02:53
uh that the KV vector for the like every
- 1:02:57
token is like a contiguous
- 1:03:00
but it will be mapping to a different
- 1:03:02
physical address.
- 1:03:06
So that really helped like saving a lot
- 1:03:10
of memory. Uh and it was only possible
- 1:03:14
because you they considered like memory
- 1:03:16
as a set of blocks and you would be
- 1:03:19
dynamically allocating those blocks as
- 1:03:21
the request need as the like new tokens
- 1:03:25
comes in and they need that kind of
- 1:03:27
memory.
- 1:03:30
The another lever is like when you are
- 1:03:34
sending multiple requests
- 1:03:37
in the batch
- 1:03:39
GPU is like taking those requests
- 1:03:44
but it does not accepts the new batch
- 1:03:46
unless all the requests in that batch
- 1:03:48
gets completed. So the diagram looks
- 1:03:51
more like a page retention but here it
- 1:03:53
is more about like when is GPU available
- 1:03:57
to take the next batch. So there is a
- 1:04:01
time period where GPU is like sitting
- 1:04:03
really idle
- 1:04:05
and you want to like resolve for that
- 1:04:09
and for that like the idea was like okay
- 1:04:11
let's do that continuous batching.
- 1:04:17
So the continuous batching also really
- 1:04:19
helped with like throughput because now
- 1:04:21
you can ship more requests pretty
- 1:04:24
quickly. Keep making sure like GPU
- 1:04:26
always uh get is always like occupied
- 1:04:30
and it's not like uh sitting idle. So
- 1:04:33
you are saving on that compute.
- 1:04:36
The third is the like prefix caching. So
- 1:04:39
you remember like the KV cache helped
- 1:04:41
you save the computation for a single
- 1:04:44
request across the tokens.
- 1:04:47
But what if like you have the same
- 1:04:50
tokens across multiple requests? How do
- 1:04:53
you basically save against that? So the
- 1:04:56
prefix caching uh which was introduced
- 1:04:59
by VLM
- 1:05:03
exactly counters that
- 1:05:06
and then the third is like we talked
- 1:05:09
about the fourth actually. So we talked
- 1:05:12
about quantizing the model
- 1:05:15
but you could also you can also like
- 1:05:18
quantize the KV weights.
- 1:05:21
So that means now you you need like a
- 1:05:24
lesser space for your key and the value
- 1:05:27
vectors. That means you can serve more
- 1:05:29
key and the value vectors in the memory.
- 1:05:32
And that means like you can serve more
- 1:05:34
tokens. That means you can serve more
- 1:05:36
context context limit. And that means
- 1:05:38
like you can serve more model quality
- 1:05:45
and all of this is like uh already
- 1:05:49
present in the VLM.
- 1:05:51
You don't really need to reinvent that
- 1:05:54
wheel
- 1:05:56
uh and you can like deploy this VLM in
- 1:05:59
production and you could see that
- 1:06:02
basically growth.
- 1:06:05
So next we have like a benchmark that we
- 1:06:08
did. So this benchmark was
- 1:06:12
let me see if I have that
- 1:06:17
here
- 1:06:19
the demos.
- 1:06:23
So doing this benchmark takes like
- 1:06:26
around 1 hour because you have to
- 1:06:28
continuously stop and like restart the
- 1:06:30
VLM servers and you have to load the
- 1:06:33
models and all. So it does take a lot of
- 1:06:36
time in doing the testing but I can like
- 1:06:39
really tell you here what we are doing.
- 1:06:42
So we have kept the model as same like
- 1:06:45
the Mistful 7B.
- 1:06:48
Uh and then we have like the set of
- 1:06:51
input questions that we are sending. Uh
- 1:06:55
consider them as the prompts. Then we
- 1:06:57
have couple of helper functions here
- 1:06:59
like checking the server is up or not.
- 1:07:02
The server is the VLM server.
- 1:07:05
Then there are helper functions to get
- 1:07:08
the VLM metrics.
- 1:07:10
uh and I will talk about like what those
- 1:07:12
metrics are. Uh then there are like lot
- 1:07:15
of the benchmarks and all and then you
- 1:07:18
have to measure uh the KV usage and all.
- 1:07:23
So these are the like helper functions.
- 1:07:25
So the baseline is very simple like we
- 1:07:27
have a hugging phase baseline.
- 1:07:30
Uh this is the raw like sending the text
- 1:07:33
to the LLM getting back the response. We
- 1:07:37
see some results here. We saw like
- 1:07:40
hugging phase has a throughput of like
- 1:07:42
around 51 tokens per second. Time to
- 1:07:44
first token was like 54 and then the
- 1:07:46
inter token latency was 19.
- 1:07:50
Uh this bas uh this was all run on the
- 1:07:52
h100.
- 1:07:55
Uh and then we start like a very default
- 1:07:59
VLM server. So by default VLM provides
- 1:08:02
you the page detention, continuous
- 1:08:04
batching
- 1:08:06
and the KV caching.
- 1:08:08
So three things are present by default
- 1:08:12
and when you try to compare those
- 1:08:15
benchmarks you see your throughput is
- 1:08:18
like almost 15x you are able to serve
- 1:08:22
more tokens per second
- 1:08:25
then
- 1:08:26
your time to the first token
- 1:08:30
uh that also rises
- 1:08:34
and then your v the inter token latency
- 1:08:37
kind goes down and then your KV versus
- 1:08:41
users and the versus context rate
- 1:08:42
increases for sure.
- 1:08:47
Now when you apply the prefix caching to
- 1:08:50
it
- 1:08:52
so with the prefix caching you see like
- 1:08:54
your throughput increases
- 1:08:57
more your TDF decreases your inter token
- 1:09:01
latency is approximately same uh and
- 1:09:04
then your KV cache usage versus the
- 1:09:07
users it's kind of going down
- 1:09:10
the vers context it's not going down
- 1:09:12
it's approximately same I think this is
- 1:09:15
also approximately same it's like not
- 1:09:18
that uh big of a deal
- 1:09:21
when you apply the like KV quantization
- 1:09:24
on top of it.
- 1:09:27
So it becomes like so so you see like
- 1:09:31
your throughput is like almost similar.
- 1:09:34
Your time to first token is similar.
- 1:09:36
Your token latency is similar but then
- 1:09:40
your KV usage actually goes down. And
- 1:09:42
this is because like you have quantized
- 1:09:45
your key value space.
- 1:09:48
Uh and then there is a concept of
- 1:09:51
speculative decoding that TME will talk
- 1:09:54
about. Uh
- 1:09:56
so when you try to benchmark those so
- 1:10:00
you also see like there is a uh like a
- 1:10:03
bit of like the less KV usage there
- 1:10:06
although like the results are
- 1:10:07
approximately same.
- 1:10:15
So yeah, I mean overall like these are
- 1:10:19
the like the metrics across probably I
- 1:10:22
should
- 1:10:25
zoom out. Okay, it's not zoom out. It's
- 1:10:27
not working.
- 1:10:30
Great. So yeah, this is the like VLM
- 1:10:34
benchmarks. Um it's your production
- 1:10:36
default by the way. uh we will also
- 1:10:39
share that decision tree uh when we try
- 1:10:43
to talk about like the other engines.
- 1:10:49
So
- 1:10:51
yeah, so we should talk about like what
- 1:10:54
are some of the other inference
- 1:10:56
optimizations we can do on top of it and
- 1:10:59
what were some of the other solutions
- 1:11:01
that came out.
- 1:11:05
Uh so I would like to again invite
- 1:11:07
Tanme. He's going to talk about like
- 1:11:10
some of these optimizations.
- 1:11:18
Oh, sorry. Uh I'm so sorry. Uh I didn't
- 1:11:21
enable the slides.
- 1:11:25
Uh what was the Okay, great.
- 1:11:30
Perfect.
- 1:11:30
>> Which one?
- 1:11:31
>> The speculative.
- 1:11:32
>> Yeah. Thank you, Hersel. Yeah.
- 1:11:36
So,
- 1:11:37
so all these are like speculative
- 1:11:39
decoding all these are the uh so so what
- 1:11:43
we say uh different flavors of same kind
- 1:11:46
of soda. So this uh this technique comes
- 1:11:50
under decoding accelerator. So first one
- 1:11:53
so we are only talking about this
- 1:11:55
speculative decoding but there are other
- 1:11:58
variants like self speculative eagle
- 1:12:01
medusa
- 1:12:02
I only like I think uh this one eagle
- 1:12:05
algorithm
- 1:12:07
personally I don't think
- 1:12:10
speculative decoding works because main
- 1:12:12
problem is alignment okay so let's start
- 1:12:15
with what is uh speculative decoding
- 1:12:18
main problem is that in transformer
- 1:12:20
architecture All these tokens are
- 1:12:22
generated sequentially one by one by
- 1:12:26
one. How about just use a smaller model
- 1:12:30
and let a smaller model to generate
- 1:12:34
maybe let's say four or five tokens and
- 1:12:38
this teacher model or we can say
- 1:12:40
according to our world cup algorithm we
- 1:12:42
can say referee. So referee will decide
- 1:12:45
how many uh tokens it accept and this
- 1:12:49
loop keeps on going on and our
- 1:12:53
assumption is that there are certain
- 1:12:56
domain where this kind of things will
- 1:12:59
work like maybe in decode maybe in
- 1:13:01
coding or where almost there is no
- 1:13:04
creativity uh each uh code or syntax is
- 1:13:08
almost similar. So maybe it can help it.
- 1:13:11
But uh based on personal testing, I
- 1:13:14
didn't find this speculative decoding
- 1:13:17
useful at all. But other techniques like
- 1:13:21
uh self speculative decoding where
- 1:13:24
teacher model also have one head
- 1:13:27
auxiliary head and it will do same
- 1:13:29
similar kind of things what this base
- 1:13:32
model or small model is doing it. But
- 1:13:36
then this eagle came Eagle 1 2 3 I don't
- 1:13:39
know how many version versions are but
- 1:13:42
it is just saying that instead of
- 1:13:45
creating instead of generating tokens uh
- 1:13:47
let's uh train a small model inside
- 1:13:50
train a small model and just take a
- 1:13:53
features from one of its uh one of main
- 1:13:57
models layer so that instead of
- 1:13:59
generating token uh it will generate uh
- 1:14:01
this features so so uh so eagle is uh
- 1:14:05
Eagle is better compared to this other
- 1:14:09
kind of technologies and then another
- 1:14:11
one is Medusa which is just saying that
- 1:14:14
just generate all the tokens parallelly.
- 1:14:18
Uh okay. So here so here in this slide
- 1:14:23
>> yeah the next slide.
- 1:14:25
>> Okay.
- 1:14:29
Okay. Yeah. Okay. Now we come to uh now
- 1:14:33
we will come to this one prefix caching.
- 1:14:35
So I don't know whether people are using
- 1:14:37
this one static prefix caching or not
- 1:14:39
but thing is that main problem with
- 1:14:42
prefix caching is that sometimes we type
- 1:14:46
and make a small kind of mistake and
- 1:14:48
this standard static prefix caching is
- 1:14:51
basically it takes a prompt do some
- 1:14:53
hashing and then next time when user
- 1:14:55
asks similar kind of question it will
- 1:14:57
try to match the hash. So if hash is uh
- 1:15:01
if hash is equal then then it will
- 1:15:04
instead of recomputing all those K and B
- 1:15:06
it will just uh take it from from the
- 1:15:09
storage but you know that sometimes we
- 1:15:11
make a mistake or maybe we can just
- 1:15:13
change a word or letter something like
- 1:15:15
that then we have a very higher uh cache
- 1:15:19
uh hit cache miss hit rate so that's why
- 1:15:23
uh this one uh radics tree so radics
- 1:15:26
tree is becoming very popular and also
- 1:15:29
also because of agent. So I think almost
- 1:15:32
everyone is doing agent and most of the
- 1:15:34
computation is going during TT during
- 1:15:37
test time inference kind of thing where
- 1:15:39
we keep on asking same kind of questions
- 1:15:42
and prompt for example you are an expert
- 1:15:45
software engineer multiply by 200 times.
- 1:15:49
This kind of loop keeps on going inside
- 1:15:52
this uh agentic agentic kind of things
- 1:15:55
where it is al necessary to keep uh or
- 1:15:59
store similar kind of things in a radics
- 1:16:03
tree. So radics tree is just so so radic
- 1:16:07
tree is just advanced version of this
- 1:16:09
prefix tree where where we will just
- 1:16:12
where we will just collapse a node if it
- 1:16:14
does not have a does not have any branch
- 1:16:18
and for this kind of work where keep on
- 1:16:23
repeating same thing this uh red x tree
- 1:16:28
helps a lot and st lang uh use this kind
- 1:16:32
of algorithm
- 1:16:33
for prefix caching.
- 1:16:37
Okay. Yeah. Then there is another thing.
- 1:16:39
One is tensor RT LLM. This is very
- 1:16:42
confusing. When I first started, I was I
- 1:16:46
was just confused. What is tensor RTLM?
- 1:16:49
So yeah. So tensor RT is just a uh it's
- 1:16:53
just a standard uh SDK kind of thing.
- 1:16:56
Tensor RTLM is just an inference engine
- 1:16:59
just like VLM, SG lang. But problem is
- 1:17:03
that it is related to Nvidia. They
- 1:17:07
optimized each and every layer and every
- 1:17:10
problem as I mentioned in our world cup
- 1:17:13
algorithm. They just break everything
- 1:17:15
and optimized everything at hardware
- 1:17:18
level also. So uh yeah. So okay next.
- 1:17:27
Yeah. So for this workshop we also uh
- 1:17:30
did some benchmarking like which is best
- 1:17:34
uh so our setup was something similar
- 1:17:37
was so so we did two kind of testing.
- 1:17:39
First one is without uh without agentic
- 1:17:43
testing where we just so
- 1:17:46
we use this shared GPT uh this one data
- 1:17:48
set and uh just ask those questions uh
- 1:17:53
using VLM and SG lang.
- 1:17:57
Okay.
- 1:18:06
Yeah. Okay.
- 1:18:09
And let me just zoom it up. Okay, great.
- 1:18:14
Okay. Yeah. So, yeah, for this workshop,
- 1:18:16
we used H100 and of our first testing
- 1:18:20
was that uh we just uh we just asked uh
- 1:18:24
we take questions from shared GPT and
- 1:18:26
put it into VLM, SG lang and we found
- 1:18:30
that actually there's no statistical
- 1:18:32
difference between which one is better.
- 1:18:35
So both have almost similar kind. So
- 1:18:38
both are fulfilling similar kind of
- 1:18:40
request per second uh TTFT and latency.
- 1:18:44
So but only difference we have seen
- 1:18:47
during agentic uh agentic branching. So
- 1:18:52
uh what we did was that we asked that
- 1:18:54
similar kind of question that you are
- 1:18:56
the best this one software engineer in
- 1:18:58
the world just solve the problem of
- 1:19:02
traffic congestion in this city kind of
- 1:19:04
thing. Then we put this into LLM. LLM
- 1:19:09
generates some output. Then we did
- 1:19:11
another uh round two also. So once this
- 1:19:15
LLM generates this output, then in round
- 1:19:19
two we have specially mentioned that uh
- 1:19:23
provide uh review the proposal and give
- 1:19:26
ratings from 1 to 10. So this uh two
- 1:19:30
turns we did uh and this loop keeps on
- 1:19:34
uh repeating it. Uh what we found is
- 1:19:36
that for this kind of
- 1:19:40
uh workflow where everything is standard
- 1:19:43
all those prompts and context
- 1:19:45
engineering comes into the picture. If
- 1:19:47
we do proper this agentic branching then
- 1:19:50
I think uh this HG lang is three to four
- 1:19:53
times better. But again this depends
- 1:19:56
upon the different setup maybe uh if you
- 1:19:59
do it uh you may get different results.
- 1:20:03
Okay. Yeah. So I think uh did we
- 1:20:07
uploaded it on GitHub? Okay.
- 1:20:10
>> Yeah. So the PDF is like also in the
- 1:20:14
drive. Uh it's the same link as the
- 1:20:17
slides.
- 1:20:19
So a quick summary here.
- 1:20:23
So on a standard API workload throughput
- 1:20:26
you would see like a VLM and the SG lang
- 1:20:30
would behave same. So if you don't have
- 1:20:33
if you have like a standard workload
- 1:20:35
definitely go with VLM. It's the
- 1:20:37
production default anyways. But what
- 1:20:39
Tanme was also saying is when you try to
- 1:20:43
like make it like agentic workloads that
- 1:20:46
is where like your SG link really shines
- 1:20:51
uh and uh it kind of like provides you
- 1:20:53
all the benefits.
- 1:20:57
So yeah keep like VLM as a default but
- 1:21:00
if you have agentic workloads probably
- 1:21:03
try to move as the towards the SG lang.
- 1:21:05
if you're not happy with DB LLM. Uh but
- 1:21:12
uh okay. Uh let me
- 1:21:21
Okay.
- 1:21:26
And then like there is like the like a
- 1:21:29
comparison that is done at the 120
- 1:21:32
billion like for the GPTO OSS 120
- 1:21:35
billion. Um this is a benchmark that was
- 1:21:39
prepared by clarify. So there is like a
- 1:21:44
blog link here. Oh nice.
- 1:21:48
Okay. Yeah. So they did the similar
- 1:21:52
benchmark and they included like a
- 1:21:54
tensor RT LLM in it.
- 1:21:58
Definitely you can always go through
- 1:22:00
these benchmarks and try to understand
- 1:22:02
which basically suits your use case. As
- 1:22:06
we mentioned like Tensor RT they try to
- 1:22:08
optimize the hardware side as well
- 1:22:10
having the peak hardware performance.
- 1:22:17
Wait uh this is
- 1:22:22
okay.
- 1:22:26
Yeah.
- 1:22:27
And then like in terms of when you want
- 1:22:30
to dep pick like your engines
- 1:22:34
once you figure out like between VLM, SG
- 1:22:37
lang tenserati so that there are some
- 1:22:39
new engines that are popping up Nvidia
- 1:22:42
Dynamo for sure. Uh so they are also for
- 1:22:46
the agentic uh session routing.
- 1:22:49
Uh hugging phase is always there. It's a
- 1:22:52
simple no server. Then there is like an
- 1:22:55
MSAR
- 1:22:57
engine that was recently proposed by
- 1:22:59
Stanford. They are for like the
- 1:23:02
multimodel.
- 1:23:05
Uh so definitely you could explore those
- 1:23:08
and when you try to basically just to
- 1:23:11
like give a quick summary uh you we
- 1:23:14
start with like a baseline
- 1:23:17
we try to find what model could fit our
- 1:23:20
use cases.
- 1:23:22
Um,
- 1:23:24
so you could pick like uh Deep Seek, you
- 1:23:27
could pick like don't pick like a
- 1:23:30
Mistral 7B. I mean, it's not good. Uh,
- 1:23:34
but yeah, so you pick your model and you
- 1:23:38
want to like have a smaller memory and
- 1:23:41
you want to try to fit that bigger model
- 1:23:43
into smaller memory so that you could
- 1:23:45
save cost on the GPU cost. So you can do
- 1:23:49
like all those com quantization
- 1:23:52
then you can apply all those serving
- 1:23:54
optimizations by using the right serving
- 1:23:56
engine under the hood. So that can
- 1:23:59
really provide you that
- 1:24:02
throughut that you really want.
- 1:24:07
And
- 1:24:10
now something that you can do uh after
- 1:24:13
going back home pro because we cannot
- 1:24:16
like actually go over all the material
- 1:24:18
here uh is definitely reading about some
- 1:24:22
of the source informations like
- 1:24:24
different attention mechanisms different
- 1:24:26
like these engines like try to just read
- 1:24:29
the different benchmarks which are
- 1:24:31
present online as well
- 1:24:34
and then there are a lot of like
- 1:24:37
in-depth guides or the next phases of it
- 1:24:40
which is like learning about some KV
- 1:24:43
eviction strategies. So world is moving
- 1:24:46
towards having a separate KV cache
- 1:24:48
engineering domain. So you want to
- 1:24:51
understand what's going on in there. So
- 1:24:53
KV cache KV eviction cache compressions
- 1:24:56
hybrid memories. So there are like lot
- 1:24:59
of solutions that are happening around
- 1:25:00
there. So always try to stick to those
- 1:25:04
foundations or like the fundamentals or
- 1:25:07
the first principles and try to see like
- 1:25:10
which solution basically solves what
- 1:25:12
problem and whether you actually need
- 1:25:14
that problem to be solved for your use
- 1:25:17
case
- 1:25:18
and then there is like distributed LLM
- 1:25:21
inference which is like a different
- 1:25:24
painoint altogether. Uh you would
- 1:25:27
probably need like a two-hour workshop
- 1:25:29
there as well.
- 1:25:31
uh to like go over like all the
- 1:25:34
internals do all the hands-on.
- 1:25:41
Yes. And this is something we are trying
- 1:25:43
to propose for the AI engineer New York
- 1:25:45
session uh which is to like dive deeper
- 1:25:48
into the advanced sections of the LLM
- 1:25:50
inference. So this workshop was more for
- 1:25:53
the like beginner and the intermediate
- 1:25:55
level. Um so in this form we do have
- 1:25:58
like a feedback as well plus also the
- 1:26:01
interest. Um if uh you think like we
- 1:26:05
need certain improvements on certain
- 1:26:07
sections definitely give that feedback
- 1:26:09
as well and if you want to see this
- 1:26:11
workshop in like New York uh fair you I
- 1:26:17
mean definitely feel free to enroll your
- 1:26:19
interest.
- 1:26:22
Uhhuh.
- 1:26:24
How is it possible?
- 1:26:28
Well,
- 1:26:34
let me just check.
- 1:26:46
>> Huh?
- 1:26:48
>> URL works, right? Not the QR code. Okay.
- 1:26:52
Probably I forgot to link those two
- 1:26:54
together.
- 1:27:00
Z
- 1:27:02
G A five.
- 1:27:11
Okay, cool. Yes. So if you can give that
- 1:27:14
feedback let me just okay
- 1:27:20
that will be fine um and yeah I think we
- 1:27:24
would like to wrap this workshop then
- 1:27:28
I'm sure like lot of you would be having
- 1:27:30
a lot of questions so we can take all
- 1:27:32
those like offline uh we can meet uh and
- 1:27:36
we can uh like talk about those
- 1:27:37
questions.
- 1:27:39
>> Yeah sure. Uh thank you everyone. Thanks
- 1:27:42
for joining. Uh I think it was really
- 1:27:46
meaningful and all of you like came
- 1:27:49
here. Uh thanks a lot.
- 1:27:51
>> Yeah. Thanks.