
1:28:12
Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher
This workshop explains LLM inference bottlenecks through attention, GPU memory, KV caching, and the prefill and decode phases. It covers model quantization and attention variants, then serving optimizations and workload-specific comparisons of vLLM and SGLang.
Harshul Jain · Tanmay Sah