AI Engineer
Upcoming EventsPast EventsTalksData/MCP/SkillsAbout
TalksSpeakersTopicsOrganizationsData/MCP/Skills
← All speakers
HJ

Bio, Work & Ideas

Harshul Jain

Conference affiliation: Senior Software Engineer · Audible

BiographyWatch talks

Harshul Jain is credited as a senior software engineer at Audible in the recorded workshop Deep dive on LLM Inference at Scale. His portion explores serving techniques including paged attention, continuous batching, prefix caching and KV cache quantization.

harshuljain.substack.com ↗@harshuljain13 ↗hjain1393 ↗@hj1393 ↗

1 conference talk

1:28:12

Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher

This workshop explains LLM inference bottlenecks through attention, GPU memory, KV caching, and the prefill and decode phases. It covers model quantization and attention variants, then serving optimizations and workload-specific comparisons of vLLM and SGLang.

Harshul Jain · Tanmay Sah

References

  • ai.engineer/talks/y2W4FNAuPEA-deep-dive-on-llm-inference-at-scale ↗

Conferences

AI Engineer Miami 2027AI Engineer Code 2026AI Engineer Europe 2026AI Engineer Miami 2026AI Engineer NYC 2026AI Engineer Paris 2026AI Engineer Shanghai 2026AI Engineer Singapore 2026AI Engineer World's Fair 2026AI Engineer Code 2025AI Engineer Paris 2025AI Engineer Summit 2025AI Engineer World's Fair 2025AI Engineer World's Fair 2024AI Engineer Summit 2023
AI Engineer
GitHubTwitterLinkedInYouTubeCode of ConductPrivacyTerms of Service

© 2026 AI Engineer