← All organizations

AI inference and training infrastructure

Baseten

Baseten provides infrastructure for developers and enterprises to train, deploy, and scale AI models in production. Its Dedicated Inference serves custom and fine-tuned models, while Model APIs provide pre-optimized models for testing and production use. Training runs on the same infrastructure stack as inference. Teams can deploy in Baseten’s managed cloud, their own VPCs, or a hybrid configuration, supporting applications such as image generation, transcription, and voice agents. Customers include Cursor, Notion, Abridge, and Harvey.

Founded in 2019 to address the difficulty of deploying and scaling machine learning systems, Baseten’s founders are CEO Tuhin Srivastava, CTO Amir Haghighat, Chief Scientist Phil Howes, and Pankaj Gupta. Its engineering contributions include Truss, an open-source Python package for packaging, serving, and deploying models, and Chains, a Python framework for workflows combining models and business logic. Chains gives each component its own hardware and autoscaling, allowing developers to separate CPU processing from GPU inference while retaining local testing and typed interfaces.

In June 2026, Baseten raised a $1.5 billion Series F at a $13 billion valuation. Its annualized revenue run rate reached an estimated $600 million as of March 2026. Baseten charges for usage, through API consumption or minutes and hours of GPU capacity.

www.baseten.co

Start here

  1. From model weights to API endpoint with TensorRT-LLM

    Start with Philip Kiely and Pankaj Gupta's workshop to follow the path from supported-model selection through deployment with Baseten and Truss to benchmarking time to first token.

    Philip Kiely · Pankaj GuptaAI Engineer World's Fair 2024

  2. Optimizing inference for voice models in production

    Learn how GPU audio decoding, torch.compile, PyTorch inference mode, and client-side latency affect a production text-to-speech pipeline.

    Philip KielyAI Engineer World's Fair 2025

  3. The Rise of Open Models in the Enterprise

    Use Amir Haghighat's talk to understand the transition from buying vertical AI products and testing hosted models to operating differentiated enterprise AI systems.

    Amir HaghighatAI Engineer World's Fair 2025

  4. Introduction to LLM serving with SGLang

    Baseten presenters Philip Kiely and Yineng Zhang introduce SGLang as an open-source serving framework for language and multimodal models, guide attendees through workshop setup and model deployment, and discuss GPU-based inference optimization, quantization, hardware-specific configuration, CUDA kernels, and cache-aware routing.

    Philip Kiely · Yineng ZhangAI Engineer World's Fair 2025

Messages from the stage

Serving frameworks meet GPU configuration

The workshops connect model deployment to hardware-specific optimization, from engine building and FP8 quantization in TensorRT-LLM to CUDA kernels and cache-aware routing in SGLang.

Voice changes the performance target

Philip Kiely's Orpheus TTS talk distinguishes meeting the real-time generation threshold from maximizing token throughput. Once that threshold is met, time to first byte, concurrent streams, and GPU efficiency become priorities.

Enterprise differentiation has operational demands

Amir Haghighat connects proprietary enterprise data and provider interoperability with the constraints of running open models, including inference quality and dedicated cloud deployments.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-27