Introduction to LLM serving with SGLang
AI Engineer World's Fair 2025 · 43:42
AI inference and training infrastructure
Baseten provides infrastructure for developers and enterprises to train, deploy, and scale AI models in production. Its Dedicated Inference serves custom and fine-tuned models, while Model APIs provide pre-optimized models for testing and production use. Training runs on the same infrastructure stack as inference. Teams can deploy in Baseten’s managed cloud, their own VPCs, or a hybrid configuration, supporting applications such as image generation, transcription, and voice agents. Customers include Cursor, Notion, Abridge, and Harvey.
Founded in 2019 to address the difficulty of deploying and scaling machine learning systems, Baseten’s founders are CEO Tuhin Srivastava, CTO Amir Haghighat, Chief Scientist Phil Howes, and Pankaj Gupta. Its engineering contributions include Truss, an open-source Python package for packaging, serving, and deploying models, and Chains, a Python framework for workflows combining models and business logic. Chains gives each component its own hardware and autoscaling, allowing developers to separate CPU processing from GPU inference while retaining local testing and typed interfaces.
In June 2026, Baseten raised a $1.5 billion Series F at a $13 billion valuation. Its annualized revenue run rate reached an estimated $600 million as of March 2026. Baseten charges for usage, through API consumption or minutes and hours of GPU capacity.
AI Engineer World's Fair 2025 · 43:42
AI Engineer World's Fair 2025 · 15:13
AI Engineer World's Fair 2024 · 1:40:01
Affiliations reflect their AIE appearances, not necessarily current employment.
Start with Philip Kiely and Pankaj Gupta's workshop to follow the path from supported-model selection through deployment with Baseten and Truss to benchmarking time to first token.
Philip Kiely · Pankaj GuptaAI Engineer World's Fair 2024
Learn how GPU audio decoding, torch.compile, PyTorch inference mode, and client-side latency affect a production text-to-speech pipeline.
Philip KielyAI Engineer World's Fair 2025
Use Amir Haghighat's talk to understand the transition from buying vertical AI products and testing hosted models to operating differentiated enterprise AI systems.
Amir HaghighatAI Engineer World's Fair 2025
Baseten presenters Philip Kiely and Yineng Zhang introduce SGLang as an open-source serving framework for language and multimodal models, guide attendees through workshop setup and model deployment, and discuss GPU-based inference optimization, quantization, hardware-specific configuration, CUDA kernels, and cache-aware routing.
Philip Kiely · Yineng ZhangAI Engineer World's Fair 2025
The workshops connect model deployment to hardware-specific optimization, from engine building and FP8 quantization in TensorRT-LLM to CUDA kernels and cache-aware routing in SGLang.
Philip Kiely's Orpheus TTS talk distinguishes meeting the real-time generation threshold from maximizing token throughput. Once that threshold is met, time to first byte, concurrent streams, and GPU efficiency become priorities.
Amir Haghighat connects proprietary enterprise data and provider interoperability with the constraints of running open models, including inference quality and dedicated cloud deployments.
Affiliations reflect each recorded session, not necessarily current employment.